EDBT 2026 Demo / reviewers in the wild / expert
Shengjin Wang
dblp:38/2763
· DBLP profile ↗
167ranked-venue papers
0as first author
79since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 117 · 51 since 2021Artificial intelligence and machine learning · 94 · 45 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TCoT: Trajectory Chain-of-Thoughts for Robotic Manipulation with Failure Recovery in Vision-Language-Action ModelabstractRecent advances in vision-language-action (VLA) models have demonstrated impressive generalization for robotic manipulation. However, these models often operate by directly mapping visual and linguistic inputs to subsequent actions, lacking intermediate task planning, along with failure detection and recovery ability. These limitations prevent them from effectively decomposing complex tasks, recognizing problems, and correcting erroneous actions, ultimately resulting in complete task failure. This significantly hinders their ability to perform long-horizon tasks and generalization ability. To this end, we introduce TCoT: Trajectory Chain-of-Thought, a unified VLA framework that enhances this direct mapping with trajectory planning as well as failure detection and recovery. TCoT leverages hierarchy trajectories as a precise and compact representation of CoT reasoning for manipulation: global planning provides a high-level, goal-oriented trajectory to guide the robot toward its task objective, while local planning focuses on real-time adjustments to address dynamic changes. Moreover, we designed the Global-Local Switching Recovery algorithm that detects and effectively recovers from failures. Experimental results reveal that TCoT surpasses the state-of-the-art methods across both real and simulated scenarios and exhibits superior generalization capabilities. Huaqiang Wang, Shengjin Wang |
AAAI | 5 |
| 2026 | Preserving Topological and Geometric Embeddings for Point Cloud RecoveryabstractRecovering point clouds involves the sequential process of sampling and restoration, yet existing methods struggle to effectively leverage both topological and geometric attributes. To address this, we propose an end-to-end architecture named TopGeoFormer, which maintains these critical properties throughout the sampling and restoration phases. First, we revisit traditional feature extraction techniques to yield topological embedding using a continuous mapping of relative relationships between neighboring points, and integrate it in both phases for preserving the structure of the original space. Second, we propose the InterTwining Attention to fully merge topological and geometric embeddings, which queries shape with local awareness in both phases to form a learnable 3D shape context facilitated with point-wise, point-shape-wise, and intra-shape features. Third, we introduce a full geometry loss and a topological constraint loss to optimize the embeddings in both Euclidean and topological spaces. The geometry loss uses inconsistent matching between coarse-to-fine generations and targets for reconstructing better geometric details, and the constraint loss limits embedding variances for better approximation of the topological space. In experiments, we comprehensively analyze the circumstances using the conventional and learning-based sampling/upsampling/recovery algorithms. The quantitative and qualitative results demonstrate that our method significantly outperforms existing sampling and recovery methods. Kaiyue Zhou, Zelong Tan, Ya-li Li, Shengjin Wang |
AAAI | 5 |
| 2026 | Human-Background Decoupling for Domain-Generalizable Person Re-identification
Yali Li 0001, Shengjin Wang |
ICIC (12) | 3 |
| 2026 | Unsupervised Temporal Correspondence Learning for Unified Video Object RemovalabstractVideo object removal aims at erasing a target object in the entire video and filling holes with plausible contents, given an object mask in the first frame as input. Existing solutions mostly break down the task into (supervised) mask tracking and (self-supervised) video completion, and then separately tackle them with tailored designs. In this paper, we introduce a new setup, coined as unified video object removal, where mask tracking and completion are addressed within a unified framework. Despite introducing more challenges, the setup is promising for future practical usage. We embrace the observation that these two sub-tasks have strong inherent connections in terms of pixel-level temporal correspondence. Making full use of the connections could be beneficial considering the complexity of both algorithm and deployment. We propose a single network linking the two sub-tasks by inferring temporal correspondences across multiple frames, i.e., correspondences between valid-valid (V-V) pixel pairs for mask tracking and correspondences between valid-hole (V-H) pixel pairs for video completion. Thanks to the unified setup, the network can be learned end-to-end in a totally unsupervised fashion without any annotations. We demonstrate that our method can generate visually pleasing results and perform favorably against existing separate solutions in realistic test cases. Zhongdao Wang, Jinglu Wang, Xiao Li 0030, Yali Li 0001, Yan Lu 0001, Shengjin Wang |
IEEE Trans. Image Process. | 6 |
| 2025 | LIBA: Language Instructed Multi-granularity Bridge Assistant for 3D Visual Groundingabstract3D Vision Grounding (3D-VG) seeks to unravel referential language and identify targets in 3D physical world. Prevailing methods align with the 2D-VG's pipeline to pinpoint the referred object in a categorical multi-modal reasoning manner. However, the geometric complexities of 3D scenes and the nuanced syntactic structures of language, exacerbates the \textbf{granularity inconsistency} of point cloud and text features, hindering the development of 3D-VG systems in complex scenarios. Towards this issue, we propose LIBA, a Language-Instructed multi-granularity Bridge Assistant tailored for 3D-VG task. LIBA tackles this issue as follows. (1) \textit{How to establish a multi-granularity 3D vision-text feature alignment in a unified model}? We advance a bilateral Dynamic Bridge Adapter (DBA) build multi-granularity interaction of 3D vision and language backnones during feature extraction. We further develop the Language-aware Cross-scale Object Modulation (LCOM) module to integrate multi-scale point cloud features modulated by language information. (2) After aligning multi-modal features, \textit{how to fully harness language model's knowledge to bolster vision concepts understanding}? A LLM-guided Hierarchical Query Selection (LLM-HQS) module incorporates world knowledge of Large Language Model~(LLM) to ground the target referral via an Attribute-then-Relation reasoning process. In this manner, our LIBA inherits reasoning prowess and world knowledge of LLM to bridge point clouds and texts at multiple granularities. Experiments on ScanRefer and Nr3D/Sr3D benchmarks substantiate the superiority of our LIBA, trumping state-of-the-arts by a considerable margin. Yali Li 0001, Eastman Z. Y. Wu, Shengjin Wang |
AAAI | 4 |
| 2025 | HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene InteractionabstractWhile flourishing developments have been witnessed in text-to-motion generation, synthesizing physically realistic, controllable, language-conditioned Human Scene Interactions (HSI) remains a relatively underexplored landscape. Current HSI methods naively rely on conditional Variational AutoEncoder (cVAE) and diffusion models. They are typically associated with limited modalities of control signals and task-specific frameworks design, leading to inflexible adaptation across various interaction scenarios and descriptive-unfaithful motions in diverse 3D physical environments. In this paper, we propose HSI-GPT, a General-Purpose Large Scene-Motion-Language Model that applies "next-token prediction" paradigm of Large Language Models to the HSI domain. HSI-GPT not only exhibits remarkable flexibility to accommodate diverse control signals (3D scenes, textual commands, key-frame poses, as well as scene affordances), but it seamlessly supports various HSI-related tasks (e.g., multi-modal controlled HSI generation, HSI understanding, and general motion completion in 3D scenes). First, HSI-GPT quantizes textual descriptions and human motions into discrete, LLM-interpretable tokens with multi-modal tokenizers. Inspired by multi-modal learning, we develop a recipe for aligning mixed-modality tokens into the shared embedding space of LLMs. These interaction tokens are then organized into unified instruction following prompts, allowing HSI-GPT to fine-tune on question-and-answer tasks. Extensive experiments and visualizations validate that our general-purpose HSI-GPT model delivers exceptional performance across multiple HSI-related tasks. Yali Li 0001, Shengjin Wang |
CVPR | 4 |
| 2025 | Attention Augmented Structure-centric Bias Mitigation with Feature DisentanglementabstractImage classification models often rely on superficial visual features, such as textures or colors, leading to undesired bias. This can compromise the robustness and reliability of deep models, particularly their performance on out-of-distribution (o.o.d.) datasets. Existing approaches, focusing on data-centric aspects, typically predefine specific bias types to mitigate the impact of these superficial features. However, such data-centric methods may lack extensibility due to their focus on predefined biases. In this paper, we propose an attention augmented structure-centric bias mitigation method, considering network architecture can be flexibly manipulated to address a variety of visual features. This method captures global semantic representations by integrating the strengths of both self-attention and convolution, introducing a global receptive field to Convolutional Neural Networks. By incorporating feature disentanglement and augmentation, our concise network demonstrates improved performance as feature diversity increases in the latent space. Our method achieves state-of-the-art results on synthetic datasets (Colored MNIST and Corrupted CIFAR10) and shows impressive performance on challenging real-world datasets (ImageNet and BFFHQ), with improvements of about 2%-5% across different subsets. Xuege Hou, Yali Li 0001, Shengjin Wang |
ICASSP | 3 |
| 2025 | Find Details in Long Videos: Tower-of-Thoughts and Self-Retrieval Augmented Generation for Video UnderstandingabstractThe Large Vision-Language Model (LVLM) has achieved impressive performance in the field of visual-language understanding. However, its ability to understand longer videos is still limited due to the length and information diversity of multi-modal videos. Moreover, accurately matching detailed content within videos remains an open research problem. We design a new framework for LVLM inference, "Tower of Thoughts" (ToT), which extends the "Chain-of-Thought" (CoT) approach to the visual domain and constructs the high-dimensional semantics of the complete videos from the bottom up. Meanwhile, to achieve question-answering for video details within the constraints of the restricted context window, we propose a method of self-retrieval augmented generation (SRAG), which makes it possible to obtain details from long videos by storing and accessing video text as dense vectors in non-parametric memory. The solution of combining the ToT with SRAG enables our model to have cross-modal high-density semantic fusion and comprehensive and accurate generation capabilities, thereby achieving rationalized video answers. Experiments on public benchmarks demonstrate the effectiveness of our proposed method. In addition, we also conducted experiments on multi-modal long videos in the open world and achieved remarkable outcomes. These results provide new perspectives and technical routes for the future development of visual language models. Tong Yue, Mingrui Xiao, Dafeng Zhang, Yali Li 0001, Shengjin Wang |
ICASSP | 6 |
| 2025 | Dynamic Object Queries for Transformer-based Incremental Object DetectionabstractIncremental object detection (IOD) aims to sequentially learn new classes, while maintaining the capability to locate and identify old ones. Prior methodologies mainly tackle catastrophic forgetting through knowledge distillation and exemplar replay, ignoring the conflict between limited model capacity and increasing knowledge. In this paper, we propose the Dynamic object Query-based DEtection TRansformer (DyQ-DETR), which incrementally expands the model representation ability to achieve stability-plasticity tradeoff. First, a new set of learnable object queries are fed into the decoder to represent new classes. Second, we propose the isolated bipartite matching for object queries in different phases, based on disentangled self-attention. Thanks to the separate supervision and computation over object queries, we further present the risk-balanced partial calibration for effective exemplar replay. Extensive experiments demonstrate that DyQ-DETR significantly surpasses the state-of-the-art methods, with limited parameter overhead. The code is available at https://github.com/THUzhangjic/DyQ-DETR. Jichuan Zhang, Wei Li 0110, Shuang Cheng, Yali Li 0001, Shengjin Wang |
ICASSP | 5 |
| 2025 | BookBot: A Robotic Manipulation Benchmark for Voice-Driven Book Recognition and Grasping in Cluttered EnvironmentsabstractBooks, as enduring repositories of cultural heritage as well as knowledge, play a fundamental role in human development. Although advances in embodied AI and robotics revolutionize automation in domains, e.g., manufacturing and logistics, robotic book manipulation remains an underexplored frontier. Two primary bottlenecks impede progress: (1) scarcity of fine-grained annotated datasets for benchmarking robotic book manipulation, and (2) lack of unified perception-action frameworks capable of dynamically coupling multi-modal sensing and manipulation in real-world scenarios. To these issues, we present THU-Book, the first open-access benchmark featuring 643 3D scene captures, encompassing 11,298 high-fidelity book instances with rich annotations to support tasks from book recognition and localization to grasping and repositioning. Building upon this foundation, we develop BookBot, a novel voice-interactive book manipulation pipeline to support cross-environmental, multilingual, and multi-categorical book manipulation. First, we utilize Large Language Models (LLMs) to parse and comprehend ambiguity in user instructions. We further propose an instance segmentation module combined with OCR tool to link language to visual instances. Finally, we introduce a PCA-based manipulation policy to refine the robotic grasp pose, utilizing the principal components of the books’ geometry, improving the precision and efficiency of grasping. Experiments conducted on the THU-Book benchmark validate the effectiveness of our BookBot. The dataset is available at https://github.com/wanghq-public/BookBot. Huaqiang Wang, Shengjin Wang |
IROS | 5 |
| 2025 | UPL-Net: Uncertainty-aware prompt learning network for semi-supervised action recognition
Shu Yang 0007, Yali Li 0001, Shengjin Wang |
Neurocomputing | 3 |
| 2025 | Deep representation learning for license plate recognition in low quality video images
Kemeng Zhao, Liangrui Peng, Pei Tang, Shengjin Wang |
Mach. Vis. Appl. | 6 |
| 2025 | UniDetector: Towards Universal Object Detection With Heterogeneous SupervisionabstractIn this paper, we formally address universal object detection, which aims to detect every category in every scene. The dependence on human annotations, the limited visual information, and the novel categories in open world severely restrict the universality of detectors. We propose UniDetector, a universal object detector that recognizes enormous categories in the open world. The critical points for UniDetector are: 1) it leverages images of multiple sources and heterogeneous label spaces in training through image-text alignment, which guarantees sufficient information for universal representations. 2) it involves heterogeneous supervision training, which alleviates the dependence on the limited fully-labeled images. 3) it generalizes to open world easily while keeping the balance between seen and unseen classes. 4) it further promotes generalizing to novel categories through our proposed decoupling training manner and probability calibration. These contributions allow UniDetector to detect over 7 k categories, the largest measurable size so far, with only about 500 classes participating in training. Our UniDetector behaves the strong zero-shot ability on large-vocabulary datasets - it surpasses supervised baselines by more than 5% without seeing any corresponding images. On 13 detection datasets with various scenes, UniDetector also achieves state-of-the-art performance with only a 3% amount of training data. Zhenyu Wang 0005, Yali Li 0001, Xi Chen 0119, Ser-Nam Lim, Antonio Torralba 0001, Hengshuang Zhao, Shengjin Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | SnapCFL: A Pre-Clustering-Based Clustered Federated Learning Framework for Data and System HeterogeneitiesabstractFederated Learning (FL) has emerged as a promising framework to address data privacy concerns associated with mobile devices, in contrast to conventional Machine Learning (ML). However, traditional FL encounters significant challenges due to the heterogeneities among different clients. Clustered Federated Learning (CFL) has demonstrated effectiveness in mitigating the data heterogeneity challenge, which significantly limits a broader application of FL. Nevertheless, existing CFL approaches often tightly couple the clustering process with the main FL process, affecting the flexibility and performance of CFL. In this paper, we propose a pre-clustering-based CFL approach, named SnapCFL, which decouples the CFL process into pre-clustering and main FL stages, considering both the impact of heterogeneity on CFL accuracy and the framework's flexibility. The pre-clustering stage models the measurement of data similarity as a two-sample hypothesis testing problem to more accurately group clients and alleviate data heterogeneity. In the main FL stage, a constraint-based client selection method is employed to address the system heterogeneity problem. We conduct extensive experiments using popular datasets with various heterogeneity settings. The results demonstrate that SnapCFL achieves excellent performance in terms of accuracy and efficiency. Compared to five other state-of-the-art approaches, SnapCFL can improve model accuracy by 0.7%$\sim$36.4%, and achieve the same level of accuracy with at least 0.08× the convergence time. Yujun Cheng, Weiting Zhang, Jiawen Kang 0001, Shengjin Wang, Dusit Niyato |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | Contrastive Unsupervised Representation Learning With Optimize-Selected Training SamplesabstractContrastive unsupervised representation learning (CURL) is a technique that seeks to learn feature sets from unlabeled data. It has found widespread and successful application in unsupervised feature learning, with the design of positive and negative pairs serving as the type of data samples. While CURL has seen empirical successes in recent years, there is still room for improvement in terms of the pair data generation process. This includes tasks such as combining and re-filtering samples, or implementing transformations among positive/negative pairs. We refer to this as the sample selection process. In this article, we introduce an optimized pair-data sample selection method for CURL. This method efficiently ensures that the two types of sampled data (similar pair and dissimilar pair) do not belong to the same class. We provide a theoretical analysis to demonstrate why our proposed method enhances learning performance by analyzing its error probability. Furthermore, we extend our proof into PAC-Bayes generalization to illustrate how our method tightens the bounds provided in previous literature. Our numerical experiments on text/image datasets show that our method achieves competitive accuracy with good generalization bounds. Yujun Cheng, Xuejing Li, Shengjin Wang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Probably Approximately Correct Bayes Meta-Learning With Parameterized-Bounded GuaranteesabstractIn meta-learning, the learner extracts knowledge from the observed tasks and quickly adapts to unseen future tasks. We provide a novel and rigorous-analyzed probably approximately correct Bayes (PAC-Bayes) meta-learning method with parameterized bounds, which learns a posterior distribution from given priors and the data samples. The proposed method is designed to improve generalization stabilities with tighter bound guarantees. We prove that the proposed PAC-Bayes bound of the meta-learner is tighter than previous work under a given condition in a rigorous theoretical way. An explicit theoretical analysis of the generalization errors is also given based on the proposed meta-learning method. Using the proposed bound in our work, we deduce an optimal objective function of the meta-learner that should be minimized during the meta-training process. We validate our theoretical hypothesis by conducting synthetic and real-world environments for meta-learning. Both rigorous proofs and experimental results reveal that our method yields state-of-the-art performances under a variety of meta-learning tasks in terms of accuracy and uncertainty robustness. Yujun Cheng, Junyu Shen, Xuejing Li, Shengjin Wang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | G3-LQ: Marrying Hyperbolic Alignment with Explicit Semantic-Geometric Modeling for 3D Visual GroundingabstractGrounding referred objects in 3D scenes is a burgeoning vision-language task pivotal for propelling Embodied AI, as it endeavors to connect the 3D physical world with free-form descriptions. Compared to the 2D counterparts, challenges posed by the variability of 3D visual grounding remain relatively unsolved in existing studies: 1) the underlying geometric and complex spatial relationships in 3D scene. 2) the inherent complexity of 3D grounded language. 3) the inconsistencies between text and geometric features. To tackle these issues, we propose G3-LQ, a DEtection TRansformer-based model tailored for 3D visual grounding task. G3-LQ explicitly models Geometric-aware visual rep-resentations and Generates fine-Grained Language-guided object Queries in an overarching framework, which com-prises two dedicated modules. Specifically, the Position Adaptive Geometric Exploring (PAGE) unearths underlying information of 3D objects in the geometric details and spatial relationships perspectives. The Fine-grained Language-guided Query Selection (Flan-QS) delves into syntactic structure of texts and generates object queries that exhibit higher relevance towards fine-grained text features. Finally, a pioneering Poincaré Semantic Alignment (PSA) loss establishes semantic-geometry consistencies by modeling non-linear vision-text feature mappings and aligning them on a hyperbolic prototype-Poincaré ball. Extensive experiments verify the superiority of our G3-LQ method, trumping the state-of-the-arts by a considerable margin. Yali Li 0001, Shengjin Wang |
CVPR | 3 |
| 2024 | Exploring Pose-Aware Human-Object Interaction via Hybrid LearningabstractHuman-Object Interaction (HOI) detection plays a crucial role in visual scene comprehension. In recent advancements, two-stage detectors have taken a prominent position. However, they are encumbered by two primary challenges. First, the misalignment between feature representation and relation reasoning gives rise to a deficiency in discrimi-native features crucial for interaction detection. Second, due to sparse annotation, the second-stage interaction head generates numerous candidatepairs, with only a small fraction receiving supervision. Towards these issues, we propose a hybrid learning method based on pose-aware HOI feature refinement. Specifically, we de-vise pose-aware feature refinement that encodes spatial fea-tures by considering human body pose characteristics. It can direct attention towards key regions, ultimately offering a wealth of fine-grained features imperative for HOI de-tection. Further, we introduce a hybrid learning method that combines HOI triplets with probabilistic soft labels supervision, which is regenerated from decoupled verb-object pairs. This method explores the implicit connections between the interactions, enhancing model generalization without requiring additional data. Our method establishes state-of-the-art performance on HICO-DET benchmark and excels notably in detecting rare HOIs. Eastman Z. Y. Wu, Yali Li 0001, Shengjin Wang |
CVPR | 4 |
| 2024 | Risk-Aware Self-consistent Imitation Learning for Trajectory Planning in Autonomous Driving
Yixuan Fan, Yali Li 0001, Shengjin Wang |
ECCV (13) | 3 |
| 2024 | OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation
Zhenyu Wang 0005, Yali Li 0001, Taichi Liu, Hengshuang Zhao, Shengjin Wang |
ECCV (47) | 5 |
| 2024 | RCIF: Towards Robust Distributed DNN Collaborative Inference Under Highly Lossy NetworksabstractCollaborative Inference is a prospective paradigm for accelerating Deep Neural Network (DNN) inference by harnessing the computational resources of multiple devices. However, in highly lossy network environments, such as those encountered in wireless communication systems, the transmission loss of intermediate feature maps between devices can result in significant degradation of co-inference accuracy. In this paper, we first conduct a comprehensive investigation into the impact of intermediate feature map loss in real-world wireless scenarios and provide an in-depth analysis of loss patterns under UDP transmission. Motivated by these observations, we introduce Robust Co-inference Framework (RCIF), a novel framework that employs a hierarchical mask strategy to selectively drop activations at two different scales of feature maps. This approach enhances the robustness of DNN co-inference in the presence of network losses. Our evaluation on a variety of datasets and network architectures demonstrates that RCIF significantly enhances the accuracy and robustness of distributed DNN co-inference under highly lossy network conditions. Specifically, our results show that RCIF can achieve up to a 659% increase in accuracy compared to the original model under particularly poor network conditions. Yujun Cheng, Shengjin Wang |
ICASSP | 3 |
| 2024 | FED-SDS: Adaptive Structured Dynamic Sparsity for Federated Learning Under Heterogeneous ClientsabstractFederated Learning (FL) is a widely utilized distributed learning methodology that facilitates real-time continuous learning while preserving client privacy. In most FL implementations, it is assumed that all edge clients possess sufficient computational capabilities to participate in the training of a Deep Neural Network (DNN) model. However, in practical applications, some clients may have limited resources and can only train a significantly smaller local model. To address system heterogeneity, this paper introduces Fed-SDS, an approach that adaptively tailors sparsity strategies for local models. In comparison to existing sparse FL schemes, Fed-SDS improves convergence and enhances model accuracy through a novel channel-wise sparsity metric, namely Mean Weight Magnitude with Gradient (MWMG). In our experiments, we compared Fed-SDS with other sparse FL methods. Empirical results demonstrate that while other sparse methods can significantly impact convergence, Fed-SDS can achieve the highest task accuracies and convergence speed in various system and data heterogeneity scenarios. Yujun Cheng, Shengjin Wang |
ICASSP | 3 |
| 2024 | Learning Generalizable Visual Representations via Self-Supervised Information BottleneckabstractNumerous approaches have recently emerged in the realm of self-supervised visual representation learning. While these methods have demonstrated empirical success, a theoretical foundation that understands and unifies these diverse techniques remains to be established. In this work, we draw inspiration from the principles underlying brain-based learning and propose a new method named self-supervised information bottleneck. Our method aims to maximize the mutual information between representations of views derived from the same image, while maintaining a minimal mutual information between the view and its corresponding representation at the same time. The brain-inspired method provides a unified information-theoretic perspective on various self-supervised approaches. This unified framework also empowers the model to learn generalizable visual representations for diverse downstream tasks and data distributions, achieving state-of-the-art performance across a wide variety of image and video tasks. Yali Li 0001, Shengjin Wang |
ICASSP | 3 |
| 2024 | Alice Benchmarks: Connecting Real World Re-Identification with the SyntheticabstractFor object re-identification (re-ID), learning from synthetic data has become a promising strategy to cheaply acquire large-scale annotated datasets and effective models, with few privacy concerns. Many interesting research problems arise from this strategy, e.g., how to reduce the domain gap between synthetic source and real-world target. To facilitate developing more new approaches in learning from synthetic data, we introduce the Alice benchmarks, large-scale datasets providing benchmarks as well as evaluation protocols to the research community. Within the Alice benchmarks, two object re-ID tasks are offered: person and vehicle re-ID. We collected and annotated two challenging real-world target datasets: AlicePerson and AliceVehicle, captured under various illuminations, image resolutions, etc. As an important feature of our real target, the clusterability of its training set is not manually guaranteed to make it closer to a real domain adaptation test scenario. Correspondingly, we reuse existing PersonX and VehicleX as synthetic source domains. The primary goal is to train models from synthetic data that can work effectively in the real world. In this paper, we detail the settings of Alice benchmarks, provide an analysis of existing commonly-used domain adaptation methods, and discuss some interesting future directions. An online server has been set up for the community to evaluate methods conveniently and fairly. Datasets and the online server details are available at https://sites.google.com/view/alice-benchmarks. Xiaoxiao Sun 0002, Yue Yao 0001, Shengjin Wang, Hongdong Li, Liang Zheng 0001 |
ICLR | 3 |
| 2024 | Representation Distillation for Efficient Self-Supervised LearningabstractSiamese self-supervised learning has shown significant progress recently, which relies on Siamese networks with identical encoders in the two branches. However, due to this inherent design of Siamese networks, the overall model capacity is primarily constrained by the encoder of interest, resulting in the representation bottleneck problem during pre-training. To address this limitation, we propose a new Distill Your Own Latent (DYOL) method that can perform self-supervised learning between branches with different architectures. So a larger target network can be employed to provide stronger self-supervision. We first decouple the update process of the target network from the online network to prevent shortcut learning. Then we distill the representation directly from the target network into the online network by enforcing the view consistency between networks. Extensive experiments on various downstream tasks validate the effectiveness of our method. Importantly, the results demonstrate that strong target networks are efficient self-supervised distillers, which enable small online networks to attain similar results to large target networks (parameter efficiency) and achieve superior performance with a much smaller number of pretraining epochs and samples (time and data efficiency). Yali Li 0001, Shengjin Wang |
ICME | 3 |
| 2024 | A Dataset with Multi-Modal Information and Multi-Granularity Descriptions for Video CaptioningabstractVideo captioning aims to generate natural language descriptions automatically from videos. While datasets like MSVD and MSR-VTT have driven research in recent years, they predominantly focus on visual features and describe simple actions, ignoring audio, text, and other modal information. Which, however, is limited, because multi-modal information plays an important role in generating accurate captions. In this study, we introduce a dataset, News-11k, which includes over 150,000 captions with multi-modal information from more than 11,000 selected news video clips. We annotate multi-granularity captions from three perspectives: coarse-grained, medium-grained, and fine-grained captions. Due to the characteristics of news videos, generating accurate captions on our dataset requires multi-modal understanding ability. Therefore, we propose a baseline model for multi-modal video captioning. To address the challenge of multi-modal information fusion, we devise the concatenating modal embedding strategy. Experiments indicate that multi-modal information significantly enhances the understanding of the deeper semantics in videos. Data will be made available on https://github.com/David-Zeng-Zijian/News-11k. Mingrui Xiao, Shu Yang 0007, Yali Li 0001, Shengjin Wang |
ICME | 6 |
| 2024 | Cascaded Network with Hierarchical Self-Distillation for Sparse Point Cloud ClassificationabstractIncomplete point clouds that are generally scanned with flaws or large gaps can cause to inaccurate predictions by learning-based classification methods. To address this issue, in this paper, we propose an end-to-end architecture that compensates for and identifies partial point clouds on the fly. First, we propose a cascaded solution that integrates both the upstream and downstream networks simultaneously, allowing the task-oriented downstream to identify the points generated by the completion-oriented upstream. These two streams complement each other, resulting in improved performance for both completion and downstream-dependent tasks. Second, to explicitly understand the predicted points’ pattern, we introduce hierarchical self-distillation (HSD), which can be applied to arbitrary hierarchy-based point cloud methods. On the classification task, our proposed method performs competitively on the synthetic dataset and achieves superior results on the challenging real-world benchmark when compared to the state-of-the-art models. Kaiyue Zhou, Ming Dong 0001, Peiyuan Zhi, Shengjin Wang |
ICME | 4 |
| 2024 | Cross Teaching between Single-Spectral and Multi-Spectral Detection Transformers for Remote Sensing Object DetectionabstractIn recent years, to enhance all-weather observation capabilities, remote sensing platforms have been increasingly equipped with thermal infrared (TIR) sensors in addition to visible spectrum (RGB) sensors. Consequently, remote sensing object detection methods started utilizing images from both modalities to improve detection accuracy. However, the inconsistency of target visibility across the two spectrums would introduce confusion of the multi-spectral model, leading to lower detection performance compared to the single-spectral TIR model. In this study, a fine-grained cross-teaching method is proposed to mitigate confusion caused by visibility inconsistency. In detail, leveraging the one-to-one matching mechanism of Detection Transformer (DETR), if a target is visible in both images, knowledge is distilled from a multi-spectral DETR to a TIR-only DETR. Conversely, if an object is only visible in the TIR image, reverse distillation is conducted. Experiments show that cross teaching aids both the single-spectral and multi-spectral models, achieving the state-of-the-art performance on the DroneVehicle dataset. Jiahe Zhu, Kaiyue Zhou, Shengjin Wang, Hongbing Ma |
IGARSS | 4 |
| 2024 | Transformer-Based Few-Shot Object Detection with Enhanced Fine-Tuning StabilityabstractFew-shot object detection (FSOD) aims at detecting unseen classes with limited annotated novel examples. while the fine-tuning paradigm has been proven effective for FSOD, most existing approaches are based on Faster R-CNN framework, neglecting exploration of advanced frameworks such as DETR. In this paper, we propose a transformer-based few-shot object detection method. To address the instability issue in few-shot fine-tuning of DETR, we propose a simple yet effective method to enhance the fine-tuning stability from two perspectives: 1) an enhanced initialization method for classification layer, hollow initialization. 2) Category-Level Dropout (CLD), to mitigate effect of missing annotations by controlling gradient backpropagation of classifier. Comprehensive experimental results on MS COCO show that our method significantly improves the baseline and achieves very promising performance. Zuyu Chen, Yali Li 0001, Shengjin Wang |
IJCNN | 3 |
| 2024 | Adaptively Building a Video-language Model for Video Captioning and Retrieval without Massive Video Pretraining
Zihao Liu 0022, Shengjin Wang, Jiayao Qian |
ACM Multimedia | 3 |
| 2024 | One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object DetectionabstractThe current trend in computer vision is to utilize one universal model to address all various tasks. Achieving such a universal model inevitably requires incorporating multi-domain data for joint training to learn across multiple problem scenarios. In point cloud based 3D object detection, however, such multi-domain joint training is highly challenging, because large domain gaps among point clouds from different datasets lead to the severe domain-interference problem. In this paper, we propose OneDet3D, a universal one-for-all model that addresses 3D detection across different domains, including diverse indoor and outdoor scenes, within the same framework and only one set of parameters. We propose the domain-aware partitioning in scatter and context, guided by a routing mechanism, to address the data interference issue, and further incorporate the text modality for a language-guided classification to unify the multi-dataset label spaces and mitigate the category interference issue. The fully sparse structure and anchor-free head further accommodate point clouds with significant scale disparities. Extensive experiments demonstrate the strong universal ability of OneDet3D to utilize only one trained model for addressing almost all 3D object detection tasks (Fig. 1). We will open-source the code for future research and applications. Zhenyu Wang 0005, Yali Li 0001, Hengshuang Zhao, Shengjin Wang |
NeurIPS | 4 |
| 2024 | RCIF: Toward Robust Distributed DNN Collaborative Inference Under Highly Lossy IoT NetworksabstractWith the rapid growth of the number of devices generating and collecting data, there has been a surge in the large-scale emergence of artificial intelligence (AI) applications predicated on Internet of Things (IoT) networks and terminals. Collaborative Inference is a prospective paradigm for accelerating Deep Neural Network (DNN) inference by harnessing the computational resources of multiple IoT devices. However, in highly lossy network environments, such as those encountered in wireless communication systems, the transmission loss of intermediate feature maps between devices can result in significant degradation of co-inference accuracy. In this paper, we first conduct a comprehensive investigation into the impact of intermediate feature map loss in real-world wireless scenarios and provide an in-depth analysis of loss patterns under UDP transmission. Motivated by these observations, we introduce Robust Co-inference Framework (RCIF), a novel framework that employs a hierarchical mask strategy to selectively drop activations at two different scales of feature maps. This approach enhances the robustness of DNN co-inference in the presence of network losses. Our evaluation on a variety of datasets and network architectures demonstrates that RCIF significantly enhances the accuracy and robustness of distributed DNN co-inference under highly lossy network conditions. Specifically, our results show that RCIF can achieve up to a 659% increase in accuracy compared to the original model under particularly poor network conditions. Yujun Cheng, Shengjin Wang |
IEEE Internet Things J. | 3 |
| 2024 | Violent Video Recognition Based on Global-Local Visual and Audio Contrastive LearningabstractThe aim of the violent recognition task is to determine whether a video contains violent behaviors. Given that violent behavior often comes with visual and audio anomalies, multimodal approaches have always played an important role in this field. However, existing methods have been limited by the insufficient utilization of audio-visual self-supervised semantic cues and correlation, resulting in a restricted representational capacity of the network and low generalization due to the scarcity of available violent video datasets. To address this issue, we propose a violent action recognition model based on global-local visual and audio contrastive learning. Our model introduces global and local contrastive objectives to achieve audio-visual multi-grained semantic alignment and leverage the correlation for violent video recognition. Experimental results demonstrate that our proposed model improves state-of-the-art by 2.31% on the VSD dataset, 0.71% on the Violent-Flows dataset, and 1.43% on the VCD dataset. Zihao Liu 0022, Shengjin Wang, Yimeng Shang |
IEEE Signal Process. Lett. | 3 |
| 2024 | Lightweight Whole-Body Human Pose Estimation With Two-Stage Refinement Training StrategyabstractHuman whole-body pose estimation is a challenging task since the model needs to learn more keypoints than the body-only case. To meet the needs of real-time performance while maintaining accuracy is also a hard issue in whole-body pose estimation due to the learning capability of lightweight networks. In order to solve the above problems to a large extent, we propose a light whole-body pose estimation method with an optimized training strategy. The model is designed based on bottom-up architecture as a base network followed by a refinement network. We propose a two-stage training process, which learns rough features in the first stage and then improves estimation precision in the second stage. An online data augmentation procedure is proposed in the second stage to improve refinement performance. We also introduce a separate learning refinement structure that fine-tunes for body, foot, and hand part independently. Experimental results show that our method improves over 8%–10% average precision compared with other lightweight state-of-the-art approaches in the whole-body pose estimation task, with nearly a quarter (25%) size of model parameters saved. Mingen Liu, Junyu Shen, Yujun Cheng, Shengjin Wang |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2024 | Learning Prompt-Enhanced Context Features for Weakly-Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection aims to locate abnormal activities in untrimmed videos without the need for frame-level supervision. Prior work has utilized graph convolution networks or self-attention mechanisms alongside multiple instance learning (MIL)-based classification loss to model temporal relations and learn discriminative features. However, these approaches are limited in two aspects: 1) Multi-branch parallel architectures, while capturing multi-scale temporal dependencies, inevitably lead to increased parameter and computational costs. 2) The binarized MIL constraint only ensures the interclass separability while neglecting the fine-grained discriminability within anomalous classes. To this end, we introduce a novel WS-VAD framework that focuses on efficient temporal modeling and anomaly innerclass discriminability. We first construct a Temporal Context Aggregation (TCA) module that simultaneously captures local-global dependencies by reusing an attention matrix along with adaptive context fusion. In addition, we propose a Prompt-Enhanced Learning (PEL) module that incorporates semantic priors using knowledge-based prompts to boost the discrimination of visual features while ensuring separability across anomaly subclasses. The proposed components have been validated through extensive experiments, which demonstrate superior performance on three challenging datasets, UCF-Crime, XD-Violence and ShanghaiTech, with fewer parameters and reduced computational effort. Notably, our method can significantly improve the detection accuracy for certain anomaly subclasses and reduced the false alarm rate. Our code is available at: https://github.com/yujiangpu20/PEL4VAD. Yujiang Pu, Shengjin Wang |
IEEE Trans. Image Process. | 4 |
| 2023 | Detecting Everything in the Open World: Towards Universal Object DetectionabstractIn this paper, we formally address universal object detection, which aims to detect every scene and predict every category. The dependence on human annotations, the limited visual information, and the novel categories in the open world severely restrict the universality of traditional detectors. We propose UniDetector, a universal object detector that has the ability to recognize enormous categories in the open world. The critical points for the universality of UniDetector are: 1) it leverages images of multiple sources and heterogeneous label spaces for training through the alignment of image and text spaces, which guarantees sufficient information for universal representations. 2) it generalizes to the open world easily while keeping the balance between seen and unseen classes, thanks to abundant information from both vision and language modalities. 3) it further promotes the generalization ability to novel categories through our proposed decoupling training manner and probability calibration. These contributions allow UniDetector to detect over 7k categories, the largest measurable category size so far, with only about 500 classes participating in training. Our UniDetector behaves the strong zero-shot generalization ability on largevocabulary datasets - it surpasses the traditional supervised baselines by more than 4% on average without seeing any corresponding images. On 13 public detection datasets with various scenes, UniDetector also achieves state-of-the-art performance with only a 3% amount of training data.11Codes are available at https://github.com/zhenyuw16/UniDetector. Zhenyu Wang 0005, Yali Li 0001, Xi Chen 0119, Ser-Nam Lim, Antonio Torralba 0001, Hengshuang Zhao, Shengjin Wang |
CVPR | 7 |
| 2023 | Divcon: Learning Concept Sequences for Semantically Diverse Image CaptioningabstractHuman generated image captions contain diverse semantic concepts, while this is still a difficult task for machines. The frequency distribution of semantic concepts in datasets is usually extremely imbalanced, leading to models repeatedly describe frequently occurring semantic concepts, resulting in a decline in the semantic diversity. In this paper, we propose a novel two-step method for diverse image captioning, generating descriptions with more diverse semantic concepts (Di-vCon). Firstly, we developed a concept sequence generator to auto-regressively generate concept sequences. This benefits the model by decoding sequences in a small searching space. Then a sentence generator takes as input the concept sequences and generates descriptions for each sequence. Experiments show that DivCon can generate captions containing diverse semantic concepts and pay more attention to the less occurring concepts. In the diverse image captioning task, Div-Con achieves the state-of-the-art results on MSCOCO dataset with oracle CIDEr and SPICE scores of 1.684 and 0.302. Yali Li 0001, Shengjin Wang |
ICASSP | 3 |
| 2023 | Identity-Seeking Self-Supervised Representation Learning for Generalizable Person Re-identificationabstractThis paper aims to learn a domain-generalizable (DG) person re-identification (ReID) representation from large-scale videos without any annotation. Prior DG ReID methods employ limited labeled data for training due to the high cost of annotation, which restricts further advances. To overcome the barriers of data and annotation, we propose to utilize large-scale unsupervised data for training. The key issue lies in how to mine identity information. To this end, we propose an Identity-seeking Self-supervised Representation learning (ISR) method. ISR constructs positive pairs from inter-frame images by modeling the instance association as a maximum-weight bipartite matching problem. A reliability-guided contrastive loss is further presented to suppress the adverse impact of noisy positive pairs, ensuring that reliable positive pairs dominate the learning process. The training cost of ISR scales approximately linearly with the data size, making it feasible to utilize large-scale data for training. The learned representation exhibits superior generalization ability. Without human annotation and fine-tuning, ISR achieves 87.0% Rank-1 on Market-1501 and 56.4% Rank-1 on MSMT17, outperforming the best supervised domain-generalizable method by 5.0% and 19.5%, respectively. In the pre-training→fine-tuning scenario, ISR achieves state-of-the-art performance, with 88.4% Rank-1 on MSMT17. The code is at https://github.com/dcp15/ISR_ICCV2023_Oral. Zhaopeng Dou, Zhongdao Wang, Yali Li 0001, Shengjin Wang |
ICCV | 4 |
| 2023 | RoICLIP: Text-Enhanced UAV-Based Video Object Detection
Yali Li 0001, Shengjin Wang |
ICIG (4) | 3 |
| 2023 | Feature Decoupling and Uncertainty Estimation for 3D Object DetectionabstractIn the real scene of 3D object detection, the point cloud collected for a single object is incomplete, resulting in misalignment of classification and regression features and uncertainty of object boundaries. Existing works pay little attention to the uncertainty caused by the incompleteness of point clouds. To address this issue, we propose a Feature Decoupling and Uncertainty Estimation single-stage 3D object detector named FDUE-Net. First, we design a Classification-Regression Attention Decoupling module to extract shared low-level and high-level features in a decoupling paradigm, generating the feature maps that are more helpful for classification or regression by layer-level and space-level attention modules. Furthermore, we propose an Uncertainty Estimation head (UE-head), which improves the quality of the predicted bounding boxes by modeling the regression values as general distributions, and uses the predicted distributions to correct the classification scores. Experiments on the KITTI dataset show that the proposed method achieves significant improvement, with 3D detection performance improved by 1.46% on the moderate set compared to the baseline, and competitive performance compared to the state-of-the-arts. Peiyuan Zhi, Kaiyue Zhou, Yali Li 0001, Shengjin Wang |
ICME | 4 |
| 2023 | Look Before You Drive: Boosting Trajectory Forecasting via Imagining FutureabstractPredicting the future trajectories of other agents in the scene fast and effectively is crucial for autonomous driving systems. We note that high-quality predictions require us to take into account the subjective initiative of the target agents, which is reflected by the fact that they themselves make decisions based on their own predictions about the future, just like our ego vehicle's prediction-planning system. However, this characteristic has been neglected in previous studies. We introduce Look Before You Drive (LBYD), a two-stage approach that explicitly incorporates both past observations and future estimates to make predictions. To get a preliminary estimate of the future, we propose a neat and effective baseline capable of making predictions for multiple agents simultaneously. We use only the most basic structures, mainly Transformer, to ensure sufficient inference speed and room for expansion. On this basis, we cooperatively train two networks to enable the coarse estimates to boost final forecasting. Our experiments demonstrate that LBYD can significantly surpass the baseline performance. Moreover, while state-of-the-art methods rely on considering heterogeneity and artificially designed inductive biases for attention modeling, LBYD performs on par with SOTA without them on both the Argoverse 1 and the large scale Argoverse 2 datasets, and can run at 67 FPS on an RTX 3090 GPU. Yixuan Fan, Yali Li 0001, Shengjin Wang |
IROS | 4 |
| 2023 | VL-Grasp: a 6-Dof Interactive Grasp Policy for Language-Oriented Objects in Cluttered Indoor ScenesabstractRobotic grasping faces new challenges in human-robot-interaction scenarios. We consider the task that the robot grasps a target object designated by human's language directives. The robot not only needs to locate a target based on vision-and-language information, but also needs to predict the reasonable grasp pose candidate at various views and postures. In this work, we propose a novel interactive grasp policy, named Visual-Lingual-Grasp (VL-Grasp), to grasp the target specified by human language. First, we build a new challenging visual grounding dataset to provide functional training data for robotic interactive perception in indoor environments. Second, we propose a 6- Dof interactive grasp policy combined with visual grounding and 6- Dof grasp pose detection to extend the universality of interactive grasping. Third, we design a grasp pose filter module to enhance the performance of the policy. Experiments demonstrate the effectiveness and extendibility of the VL-Grasp in real world. The VL-Grasp achieves a success rate of 72.5 % in different indoor scenes. The code and dataset is available at https://github.com/luyh20/VL-Grasp. Yuhao Lu, Yixuan Fan, Beixing Deng, Fangfu Liu, Yali Li 0001, Shengjin Wang |
IROS | 6 |
| 2023 | Uni3DETR: Unified 3D Detection TransformerabstractExisting point cloud based 3D detectors are designed for the particular scene, either indoor or outdoor ones. Because of the substantial differences in object distribution and point density within point clouds collected from various environments, coupled with the intricate nature of 3D metrics, there is still a lack of a unified network architecture that can accommodate diverse scenes. In this paper, we propose Uni3DETR, a unified 3D detector that addresses indoor and outdoor 3D detection within the same framework. Specifically, we employ the detection transformer with point-voxel interaction for object prediction, which leverages voxel features and points for cross-attention and behaves resistant to the discrepancies from data. We then propose the mixture of query points, which sufficiently exploits global information for dense small-range indoor scenes and local information for large-range sparse outdoor ones. Furthermore, our proposed decoupled IoU provides an easy-to-optimize training target for localization by disentangling the $xy$ and $z$ space. Extensive experiments validate that Uni3DETR exhibits excellent performance consistently on both indoor and outdoor 3D detection. In contrast to previous specialized detectors, which may perform well on some particular datasets but suffer a substantial degradation on different scenes, Uni3DETR demonstrates the strong generalization ability under heterogeneous conditions (Fig. 1). Zhenyu Wang 0005, Yali Li 0001, Xi Chen 0119, Hengshuang Zhao, Shengjin Wang |
NeurIPS | 5 |
| 2023 | A prompt tuning method for few-shot action recognitionabstractVision-language pre-training models learn visual concepts from image-text or video-text pairs, which can be adopted for visual-textual tasks. In this paper, we adopt these concepts as prior knowledge to solve the unreliable problem of minimizing the loss of limited training samples in few-shot action recognition tasks. In particular, a two-stage framework of vision-language pre-training and prompt tuning is designed. In the pre-training stage, multi-modal encoding models are jointly trained on video-text pairs to learn the semantic correspondence between video and text. In the prompt tuning stage, a prompt module with instance-level bias is trained on a few video samples to utilize the pre-trained concepts for the classification task. The experimental results show that the proposed method is superior to the baseline and state-of-the-art few-shot action recognition methods on two public video benchmarks. Shu Yang 0007, Yali Li 0001, Shengjin Wang |
VCIP | 3 |
| 2023 | Transformer Based Remote Sensing Object Detection With Enhanced Multispectral Feature ExtractionabstractAs a convention, satellites and drones are equipped with sensors of both the visible light spectrum and the infrared (IR) spectrum. However, existing remote sensing object detection methods mostly use RGB images captured by the visible light camera while ignoring IR images. Even for algorithms that take RGB-IR image pairs as input, they may fail to extract all potential features in both spectrums. This letter proposes Multispectral DETR, a remote sensing object detector based on the deformable attention mechanism. To enhance multispectral feature extraction and attention, DropSpectrum and SwitchSpectrum methods are further proposed. DropSpectrum facilitates the extraction of multispectral features by requiring the model to detection some of the targets with only one spectrum. SwitchSpectrum eliminates the level bias caused by the fixed order of RGB-IR feature maps and enhances attention on multispectral features. Experiments on the VEDAI dataset show the state-of-the-art performance of Multispectral DETR and the effectiveness of both DropSpectrum and SwitchSpectrum. Jiahe Zhu, Huan Zhang 0013, Zelong Tan, Shengjin Wang, Hongbing Ma |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | AdaZoom: Towards Scale-Aware Large Scene Object DetectionabstractDetection in large scenes is a challenging issue due to small objects and extreme scale variation. It is difficult for the deep-learning-based detector to extract features of small objects with only a few pixels. Most existing methods employ image pyramid and feature pyramid for multi-scale inference to alleviate this issue. However, they lack scale awareness to adapt to objects with different scales. In this paper, we propose a novel Adaptive Zoom (AdaZoom) network for scale-aware large scene object detection. There are three main contributions. First, an Adaptive Zoom network is proposed to actively focus and adaptively zoom the focused regions for high-performance object detection in large scenes. Second, to tackle the problem of missing annotations for focused regions, we train AdaZoom with the reward which measures the quality of generated regions, based on the paradigm of deep reinforcement learning. At last, we propose collaborative training to iteratively promote the joint performance of AdaZoom and the detector. To validate the effectiveness, we conduct extensive experiments on VisDrone2019, UAVDT and DOTA datasets. The experiments show AdaZoom brings consistent and significant improvement over different detection networks, achieving state-of-the-art performance on these datasets, especially outperforming the existing methods by AP of 4.64% on VisDrone2019. Jingtao Xu, Yali Li 0001, Shengjin Wang |
IEEE Trans. Multim. | 3 |
| 2022 | Delving into Probabilistic Uncertainty for Unsupervised Domain Adaptive Person Re-identificationabstractClustering-based unsupervised domain adaptive (UDA) person re-identification (ReID) reduces exhaustive annotations. However, owing to unsatisfactory feature embedding and imperfect clustering, pseudo labels for target domain data inherently contain an unknown proportion of wrong ones, which would mislead feature learning. In this paper, we propose an approach named probabilistic uncertainty guided progressive label refinery (P2LR) for domain adaptive person re-identification. First, we propose to model the labeling uncertainty with the probabilistic distance along with ideal single-peak distributions. A quantitative criterion is established to measure the uncertainty of pseudo labels and facilitate the network training. Second, we explore a progressive strategy for refining pseudo labels. With the uncertainty-guided alternative optimization, we balance between the exploration of target domain data and the negative effects of noisy labeling. On top of a strong baseline, we obtain significant improvements and achieve the state-of-the-art performance on four UDA ReID benchmarks. Specifically, our method outperforms the baseline by 6.5% mAP on the Duke2Market task, while surpassing the state-of-the-art method by 2.5% mAP on the Market2MSMT task. Code is available at: https://github.com/JeyesHan/P2LR. Yali Li 0001, Shengjin Wang |
AAAI | 3 |
| 2022 | Disentangling based Environment-Robust Feature Learning for Person ReID
Yali Li 0001, Shengjin Wang |
BMVC | 3 |
| 2022 | Polishing Network for Decoding of Higher-Quality Diverse Image Captions
Yali Li 0001, Shengjin Wang |
BMVC | 3 |
| 2022 | R(Det)2: Randomized Decision Routing for Object DetectionabstractIn the paradigm of object detection, the decision head is an important part, which affects detection performance significantly. Yet how to design a high-performance decision head remains to be an open issue. In this paper, we propose a novel approach to combine decision trees and deep neural networks in an end-to-end learning manner for object detection. First, we disentangle the decision choices and prediction values by plugging soft decision trees into neural networks. To facilitate effective learning, we propose randomized decision routing with node selective and associative losses, which can boost the feature representative learning and network decision simultaneously. Second, we develop the decision head for object detection with narrow branches to generate the routing probabilities and masks, for the purpose of obtaining divergent decisions from different nodes. We name this approach as the randomized decision routing for object detection, abbreviated as R(Det)2. Experiments on MS-COCO dataset demonstrate that R(Det)2 is effective to improve the detection performance. Equipped with existing detectors, it achieves 1.4 ~ 3.6% AP improvement. Yali Li 0001, Shengjin Wang |
CVPR | 2 |
| 2022 | OSKDet: Orientation-sensitive Keypoint Localization for Rotated Object DetectionabstractRotated object detection is a challenging issue in computer vision field. Inadequate rotated representation and the confusion of parametric regression have been the bottleneck for high performance rotated detection. In this paper, we propose an orientation-sensitive keypoint based rotated detector OSKDet. First, we adopt a set of keypoints to represent the target and predict the keypoint heatmap on ROI to get the rotated box. By proposing the orientation-sensitive heatmap, OSKDet could learn the shape and direction of rotated target implicitly and has stronger modeling capabilities for rotated representation, which improves the localization accuracy and acquires high quality detection results. Second, we explore a new unordered keypoint representation paradigm, which could avoid the confusion of keypoint regression caused by rule based ordering. Further-more, we propose a localization quality uncertainty module to better predict the classification score by the distribution uncertainty of keypoints heatmap. Experimental results on several public benchmarks show the state-of-the-art performance of OSKDet. Specifically, we achieve an AP of 80.91% on DOTA, 89.98% on HRSC2016, 97.27% on UCAS-AOD, and a F-measure of 92.18% on ICDAR2015, 81.43% on ICDAR2017, respectively. Dongchen Lu, Yali Li 0001, Shengjin Wang |
CVPR | 4 |
| 2022 | Noisy Boundaries: Lemon or Lemonade for Semi-supervised Instance Segmentation?abstractCurrent instance segmentation methods rely heavily on pixel-level annotated images. The huge cost to obtain such fully-annotated images restricts the dataset scale and limits the performance. In this paper, we formally address semi-supervised instance segmentation, where unlabeled images are employed to boost the performance. We construct a framework for semi-supervised instance segmentation by assigning pixel-level pseudo labels. Under this framework, we point out that noisy boundaries associated with pseudo labels are double-edged. We propose to exploit and resist them in a unified manner simultaneously: 1) To combat the negative effects of noisy boundaries, we propose a noise-tolerant mask head by leveraging low-resolution features. 2) To enhance the positive impacts, we introduce a boundary-preserving map for learning detailed information within boundary-relevant regions. We evaluate our approach by extensive experiments. It behaves extraordinarily, outperforming the supervised baseline by a large margin, more than 6% on Cityscapes, 7% on COCO and 4.5% on BDD100k. On Cityscapes, our method achieves comparable performance by utilizing only 30% labeled images. Zhenyu Wang 0005, Yali Li 0001, Shengjin Wang |
CVPR | 3 |
| 2022 | Reliability-Aware Prediction via Uncertainty Learning for Person Image Retrieval
Zhaopeng Dou, Zhongdao Wang, Yali Li 0001, Shengjin Wang |
ECCV (14) | 5 |
| 2022 | GraphCSPN: Geometry-Aware Depth Completion via Dynamic GCNs
Xiaofei Shao, Yali Li 0001, Shengjin Wang |
ECCV (33) | 5 |
| 2022 | Progressive-Granularity Retrieval Via Hierarchical Feature Alignment for Person Re-IdentificationabstractPerson re-identification (re-ID) aims to match pedestrian images from non-overlapping cameras. It is a challenging task because of the feature misalignment problem caused by occlusion. In this paper, inspired by the coarse-to-fine nature of human perception, we propose a novel Progressive-Granularity Retrieval (PGR) method to tackle this issue. Specifically, (i) we define instance-level, part-level and pixel-level features for an image. PGR learns these features by a single feature extractor to capture hierarchical clues in the image. (ii) These features are inherently related but different in perceptual granularity, and they can provide complementary information. For each type of feature, we propose a corresponding similarity metric to achieve hierarchical feature alignment. (iii) In training, we learn the model end-to-end. In inference, a progressive retrieval strategy is introduced to efficiently aggregate the complementary information provided by these features. Extensive experiments on three bench-marks of both occluded and holistic-body re-ID tasks show the effectiveness of the proposed method. Especially, our method significantly outperforms state-of-the-art by 4.5% Rank-1 score on the challenging Occluded-Duke dataset. Zhaopeng Dou, Zhongdao Wang, Yali Li 0001, Shengjin Wang |
ICASSP | 4 |
| 2022 | Few-Shot Object Detection with Local Correspondence RPN and Attentive HeadabstractExisting object detection methods rely heavily on a large number of annotated bounding boxes, which is expensive to collect. In this paper, we propose a novel few-shot object detection method named GCN-FSOD. Intending to find informal local correspondence to fully explore cues of novel classes, we propose the local correspondence region proposal network (lcRPN) and the attentive detection head for few-shot detection. Taking features from the support-query image pair as inputs, lcRPN generates region proposals by mining fine-grained local correspondence with the help of GCNs. Then the proposed attentive head performs precise detection. We conduct extensive experiments on the wildly adopted MS-COCO benchmark. The proposed GCN-FSOD brings significant performance gains and outperforms the state-of-the-art by a large margin (1.7% mAP for 10-shot). Yali Li 0001, Shengjin Wang |
ICASSP | 3 |
| 2022 | CRPN: Distinguish Novel Categories Via Class-Relevant Region Proposal Network for Few-Shot Object DetectionabstractFew-shot object detection (FSOD) has attracted more attention in computer vision, where only very few training examples are presented during model learning process. A commonly-overlooked issue in FSOD is that novel classes are usually classified as background clutters in the pre-training process. Another difficulty of FSOD is that the detection performance degrades especially under higher IoU thresholds since previous deep metric learning (DML) requires frozen region proposals without class-relevant box regression. In this work, we propose a Class-relevant Region Proposal Network (CRPN). The CRPN can derive network parameters for novel classes from pre-trained convolution kernels according to their feature similarity, which is used to eliminate the above mentioned adverse effects and improve the performance of few-shot object detection. The proposed CPRN is able to kill two birds with one stone and has two main contributions: (1) transfer a region proposal network pre-trained on base classes to novel classes; (2) perform class-dependent bounding-box regression which previous DML classifier lacks. For experimental testing, we achieve 12.7% AP75 in MS COCO dataset and 28.6% AP75 in ImageNet2015 dataset under the few-shot setting introduced by previous works, which exceeds the state-of-the-art by a certain margin. Yali Li 0001, Shengjin Wang |
ICASSP | 3 |
| 2022 | HID 2022: The 3rd International Competition on Human Identification at a DistanceabstractThe paper provides a summary of the Competition on Human Identification at a Distance 2022 (HID 2022), which is the third one in a series of competitions. HID 2022 is for promoting the research in human identification at a distance by providing a benchmark to evaluate different methods. The competition attracted 112 valid registered teams. 71 teams and 51 teams submitted their results in the first phase and the second phase, respectively. Very encouraging results have been achieved, and the accuracies of the top teams are much higher than those achieved in the previous two competitions. In this paper, we introduce the competition including the dataset, experimental settings, competition organization, results from the top teams and their analysis. The methods used by the top teams are also presented in the paper. The progress of this competition can give us an optimistic view on gait recognition. Shiqi Yu 0001, Yongzhen Huang, Liang Wang 0001, Yasushi Makihara, Shengjin Wang, Md. Atiqur Rahman Ahad, Mark S. Nixon |
IJCB | 5 |
| 2022 | Hybrid Physical Metric For 6-DoF Grasp Pose Detectionabstract6-DoF grasp pose detection of multi-grasp and multi-object is a challenge task in the field of intelligent robot. To imitate human reasoning ability for grasping objects, data driven methods are widely studied. With the introduction of large-scale datasets, we discover that a single physical metric usually generates several discrete levels of grasp confidence scores, which cannot finely distinguish millions of grasp poses and leads to inaccurate prediction results. In this paper, we propose a hybrid physical metric to solve this evaluation insufficiency. First, we define a novel metric is based on the force-closure metric, supplemented by the measurement of the object flatness, gravity and collision. Second, we leverage this hybrid physical metric to generate elaborate confidence scores. Third, to learn the new confidence scores effectively, we design a multi-resolution network called Flatness Gravity Collision GraspNet (FGC-GraspNet). FGC-GraspNet proposes a multi-resolution features learning architecture for multiple tasks and introduces a new joint loss function that enhances the average precision of the grasp detection. The network evaluation and adequate real robot experiments demonstrate the effectiveness of our hybrid physical metric and FGC-GraspNet. Our method achieves 90.5% success rate in real-world cluttered scenes. Our code is available at https://github.com/luyh20IFGC-GraspNet. Yuhao Lu, Beixing Deng, Zhenyu Wang 0005, Peiyuan Zhi, Yali Li 0001, Shengjin Wang |
ICRA | 6 |
| 2022 | Self-Supervised Learning via Maximum Entropy CodingabstractA mainstream type of current self-supervised learning methods pursues a general-purpose representation that can be well transferred to downstream tasks, typically by optimizing on a given pretext task such as instance discrimination. In this work, we argue that existing pretext tasks inevitably introduce biases into the learned representation, which in turn leads to biased transfer performance on various downstream tasks. To cope with this issue, we propose Maximum Entropy Coding (MEC), a more principled objective that explicitly optimizes on the structure of the representation, so that the learned representation is less biased and thus generalizes better to unseen downstream tasks. Inspired by the principle of maximum entropy in information theory, we hypothesize that a generalizable representation should be the one that admits the maximum entropy among all plausible representations. To make the objective end-to-end trainable, we propose to leverage the minimal coding length in lossy data coding as a computationally tractable surrogate for the entropy, and further derive a scalable reformulation of the objective that allows fast computation. Extensive experiments demonstrate that MEC learns a more generalizable representation than previous methods based on specific pretext tasks. It achieves state-of-the-art performance consistently on various downstream tasks, including not only ImageNet linear probe, but also semi-supervised classification, object detection, instance segmentation, and object tracking. Interestingly, we show that existing batch-wise and feature-wise self-supervised objectives could be seen equivalent to low-order approximations of MEC. Code and pre-trained models are available at https://github.com/xinliu20/MEC. Zhongdao Wang, Yali Li 0001, Shengjin Wang |
NeurIPS | 4 |
| 2022 | Semantic multimodal violence detection based on local-to-global embedding
Yujiang Pu, Shengjin Wang, Zihao Liu 0022, Chaonan Gu |
Neurocomputing | 3 |
| 2022 | Robust 3D face modeling and tracking from RGB-D images
Changwei Luo, Juyong Zhang, Changcun Bao, Yali Li 0001, Shengjin Wang |
Multim. Syst. | 6 |
| 2022 | SurRF: Unsupervised Multi-View Stereopsis by Learning Surface Radiance FieldabstractThe recent success in supervised multi-view stereopsis (MVS) relies on the onerously collected real-world 3D data. While the latest differentiable rendering techniques enable unsupervised MVS, they are restricted to discretized (e.g., point cloud) or implicit geometric representation, suffering from either low integrity for a textureless region or less geometric details for complex scenes. In this paper, we propose SurRF, an unsupervised MVS pipeline by learning Surface Radiance Field, i.e., a radiance field defined on a continuous and explicit 2D surface. Our key insight is that, in a local region, the explicit surface can be gradually deformed from a continuous initialization along view-dependent camera rays by differentiable rendering. That enables us to define the radiance field only on a 2D deformable surface rather than in a dense volume of 3D space, leading to compact representation while maintaining complete shape and realistic texture for large-scale complex scenes. We experimentally demonstrate that the proposed SurRF produces competitive results over the-state-of-the-art on various real-world challenging scenes, without any 3D supervision. Moreover, SurRF shows great potential in owning the joint advantages of mesh (scene manipulation), continuous surface (high geometric resolution), and radiance field (realistic rendering). Jinzhi Zhang, Mengqi Ji, Zhiwei Xu 0002, Shengjin Wang, Lu Fang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Adaptive Affinity for Associations in Multi-Target Multi-Camera TrackingabstractData associations in multi-target multi-camera tracking (MTMCT) usually estimate affinity directly from re-identification (re-ID) feature distances. However, we argue that it might not be the best choice given the difference in matching scopes between re-ID and MTMCT problems. Re-ID systems focus on global matching, which retrieves targets from all cameras and all times. In contrast, data association in tracking is a local matching problem, since its candidates only come from neighboring locations and time frames. In this paper, we design experiments to verify such misfit between global re-ID feature distances and local matching in tracking, and propose a simple yet effective approach to adapt affinity estimations to corresponding matching scopes in MTMCT. Instead of trying to deal with all appearance changes, we tailor the affinity metric to specialize in ones that might emerge during data associations. To this end, we introduce a new data sampling scheme with temporal windows originally used for data associations in tracking. Minimizing the mismatch, the adaptive affinity module brings significant improvements over global re-ID distance, and produces competitive performance on CityFlow and DukeMTMC datasets. Yunzhong Hou, Zhongdao Wang, Shengjin Wang, Liang Zheng 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | BooDet: Gradient Boosting Object Detection With Additive Learning-Based Prediction AggregationabstractIn recent years, the community of object detection has witnessed remarkable progress with the development of deep neural networks. But the detection performance still suffers from the dilemma between complex networks and single-vector predictions. In this paper, we propose a novel approach to boost the object detection performance based on aggregating predictions. First, we propose a unified module with adjustable hyper-structure to generate multiple predictions from a single detection network. Second, we formulate the additive learning for aggregating predictions, which reduces the classification and regression losses by progressively adding the prediction values. Based on the gradient Boosting strategy, the optimization of the additional predictions is further modeled as weighted regression problems to fit the Newton-descent directions. By aggregating multiple predictions from a single network, we propose the BooDet approach which can Bootstrap the classification and bounding box regression for high-performance object Detection. In particular, we plug the BooDet into Cascade R-CNN for object detection. Extensive experiments show that the proposed approach is quite effective to improve object detection. We obtain a 1.3%~2.0% improvement over the strong baseline Cascade R-CNN on COCO val dataset. We achieve 56.5% AP on the COCO test-dev dataset with only bounding box annotations. Yali Li 0001, Shengjin Wang |
IEEE Trans. Image Process. | 2 |
| 2022 | Traffic Sign Recognition With Lightweight Two-Stage Model in Complex ScenesabstractTraffic sign recognition with high accuracy and real-time is an important part of the intelligent transportation system. In this article, based on large-scale traffic signs and the inherent conflict between location regression and classification of traffic signs, we propose a novel and flexible two-stage approach. It combines a lightweight superclass detector with a refinement classifier. The main contributions lie in three aspects: (1) We use locations and sizes of signs as prior knowledge to establish a probability distribution model. It can significantly decrease the search range of signs and improve the processing speed, as well as reducing false detection. (2) We propose a high-performance lightweight superclass detector. We introduce the Inception and Channel Attention, by generating multi-scale receptive fields and adaptively adjusting channel features. It alleviates the large scale variance challenge of objects and the interference of background information. Meanwhile, we present a merging Batch Normalization and multi-scale testing method to further improve detection performance. (3) We propose a refinement classifier based on similarity measure learning for the subclass classification. It increases the precision of discriminating similar subclasses and also improves the extensibility of our approach. Our two-stage approach is simple and effective, whose paradigm is different from others. Experiments on the Tsinghua-Tencent 100K dataset demonstrate the performance of our approach. Compared with the state-of-the-art methods, our method achieves competitive performance (92.16% mAP) with a lightweight detector ($6.49M $). The processing time is$0.150s $per frame, of which the speed is increased by 3 times compared with existing methods. Zhengshuai Wang, Yali Li 0001, Shengjin Wang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | A2-FPN: Attention Aggregation Based Feature Pyramid Network for Instance SegmentationabstractLearning pyramidal feature representations is crucial for recognizing object instances at different scales. Feature Pyramid Network (FPN) is the classic architecture to build a feature pyramid with high-level semantics throughout. However, intrinsic defects in feature extraction and fusion inhibit FPN from further aggregating more discriminative features. In this work, we propose Attention Aggregation based Feature Pyramid Network (A2-FPN), to improve multi-scale feature learning through attention-guided feature aggregation. In feature extraction, it extracts discriminative features by collecting-distributing multi-level global context features, and mitigates the semantic information loss due to drastically reduced channels. In feature fusion, it aggregates complementary information from adjacent features to generate location-wise reassembly kernels for content-aware sampling, and employs channel-wise reweighting to enhance the semantic consistency before element-wise addition. A2-FPN shows consistent gains on different instance segmentation frameworks. By replacing FPN with A2-FPN in Mask R-CNN, our model boosts the performance by 2.1% and 1.6% mask AP when using ResNet-50 and ResNet-101 as backbone, respectively. Moreover, A2-FPN achieves an improvement of 2.0% and 1.4% mask AP when integrated into the strong baselines such as Cascade Mask R-CNN and Hybrid Task Cascade. Yali Li 0001, Lu Fang 0001, Shengjin Wang |
CVPR | 4 |
| 2021 | Multi-Target Domain Adaptation With Collaborative Consistency LearningabstractRecently unsupervised domain adaptation for the semantic segmentation task has become more and more popular due to high-cost of pixel-level annotation on real-world images. However, most domain adaptation methods are only restricted to single-source-single-target pair, and can not be directly extended to multiple target domains. In this work, we propose a collaborative learning framework to achieve unsupervised multi-target domain adaptation. An unsupervised domain adaptation expert model is first trained for each source-target pair and is further encouraged to collaborate with each other through a bridge built between different target domains. These expert models are further improved by adding the regularization of making the consistent pixel-wise prediction for each sample with the same structured context. To obtain a single model that works across multiple target domains, we propose to simultaneously learn a student model which is trained to not only imitate the output of each expert on the corresponding target domain, but also to pull different expert close to each other with regularization on their weights. Extensive experiments demonstrate that the proposed method can effectively exploit rich structured information contained in both labeled source domain and multiple unlabeled target domains. Not only does it perform well across multiple target domains but also performs favorably against state-of-the-art unsupervised domain adaptation methods specially trained on a single source-target pair. Code is available at https://github.com/junpan19/MTDA. Takashi Isobe, Xu Jia 0012, Shuaijun Chen, Yongjie Shi, Jianzhuang Liu, Huchuan Lu, Shengjin Wang |
CVPR | 8 |
| 2021 | Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object DetectionabstractIn this paper, we delve into semi-supervised object detection where unlabeled images are leveraged to break through the upper bound of fully-supervised object detection. Previous semi-supervised methods based on pseudo labels are severely degenerated by noise and prone to overfit to noisy labels, thus are deficient in learning different unlabeled knowledge well. To address this issue, we propose a data-uncertainty guided multi-phase learning method for semisupervised object detection. We comprehensively consider divergent types of unlabeled images according to their difficulty levels, utilize them in different phases, and ensemble models from different phases together to generate ultimate results. Image uncertainty guided easy data selection and region uncertainty guided RoI Re-weighting are involved in multi-phase learning and enable the detector to concentrate on more certain knowledge. Through extensive experiments on PASCAL VOC and MS COCO, we demonstrate that our method behaves extraordinarily compared to baseline approaches and outperforms them by a large margin, more than 3% on VOC and 2% on COCO. Zhenyu Wang 0005, Yali Li 0001, Lu Fang 0001, Shengjin Wang |
CVPR | 5 |
| 2021 | Frame-Rate-Aware Aggregation for Efficient Video Super-ResolutionabstractVideo super-resolution, which aims at producing a high-resolution video from its corresponding low-resolution version, recently draws increasing attention. In contrast to the previous works that perform explicit motion estimation and compensation, we propose a novel deep neural network which performs implicit motion estimation with frame-rate-based temporal aggregation. Specifically, the input frames are first aggregated by a frame-rate-aware 3D convolution layer, where neighboring frames are integrated with the reference frame according to the corresponding frame rate. Then, the aggregated features are fed into several branches for further aggregation. Different branches correspond to a kind of motion rate, which provides complementary information to recover missing details in the reference frame. Extensive experiments demonstrate that our method is able to handle various motion types and achieves state-of-the-art performance on several benchmarks. In addition, our model is light-weight and requires an extremely less computational load than other state-of-the-art methods. Takashi Isobe, Shengjin Wang |
ICASSP | 3 |
| 2021 | Disentangled Representation for Age-Invariant Face Recognition: A Mutual Information Minimization PerspectiveabstractGeneral face recognition has seen remarkable progress in recent years. However, large age gap still remains a big challenge due to significant alterations in facial appearance and bone structure. Disentanglement plays a key role in partitioning face representations into identity-dependent and age-dependent components for age-invariant face recognition (AIFR). In this paper we propose a multi-task learning framework based on mutual information minimization (MT-MIM), which casts the disentangled representation learning as an objective of information constraints. The method trains a disentanglement network to minimize mutual information between the identity component and age component of the face image from the same person, and reduce the effect of age variations during the identification process. For quantitative measure of the degree of disentanglement, we verify that mutual information can represent as metric. The resulting identity-dependent representations are used for age-invariant face recognition. We evaluate MT-MIM on popular public-domain face aging datasets (FG-NET, MORPH Album 2, CACD and AgeDB) and obtained significant improvements over previous state-of-the-art methods. Specifically, our method exceeds the baseline models by over 0.4% on MORPH Album 2, and over 0.7% on CACD subsets, which are impressive improvements at the high accuracy levels of above 99% and an average of 94%. Xuege Hou, Yali Li 0001, Shengjin Wang |
ICCV | 3 |
| 2021 | Towards Discriminative Representation Learning for Unsupervised Person Re-identificationabstractIn this work, we address the problem of unsupervised domain adaptation for person re-ID where annotations are available for the source domain but not for target. Previous methods typically follow a two-stage optimization pipeline, where the network is first pre-trained on source and then fine-tuned on target with pseudo labels created by feature clustering. Such methods sustain two main limitations. (1) The label noise may hinder the learning of discriminative features for recognizing target classes. (2) The domain gap may hinder knowledge transferring from source to target. We propose three types of technical schemes to alleviate these issues. First, we propose a cluster-wise contrastive learning algorithm (CCL) by iterative optimization of feature learning and cluster refinery to learn noise-tolerant representations in the unsupervised manner. Second, we adopt a progressive domain adaptation (PDA) strategy to gradually mitigate the domain gap between source and target data. Third, we propose Fourier augmentation (FA) for further maximizing the class separability of re-ID models by imposing extra constraints in the Fourier space. We observe that these proposed schemes are capable of facilitating the learning of discriminative feature representations. Experiments demonstrate that our method consistently achieves notable improvements over the state-of-the-art unsupervised re-ID methods on multiple benchmarks, e.g., surpassing MMT largely by 8.1%, 9.9%, 11.4% and 11.1% mAP on the Market-to-Duke, Duke-to-Market, Market-to-MSMT and Duke-to-MSMT tasks, respectively. Takashi Isobe, Dong Li 0025, Shengjin Wang |
ICCV | 6 |
| 2021 | Partial Off-policy Learning: Balance Accuracy and Diversity for Human-Oriented Image CaptioningabstractHuman-oriented image captioning with both high diversity and accuracy is a challenging task in vision+language modeling. The reinforcement learning (RL) based frameworks promote the accuracy of image captioning, yet seriously hurt the diversity. In contrast, other methods based on variational auto-encoder (VAE) or generative adversarial network (GAN) can produce diverse yet less accurate captions. In this work, we devote our attention to promote the diversity of RL-based image captioning. To be specific, we devise a partial off-policy learning scheme to balance accuracy and diversity. First, we keep the model exposed to varied candidate captions by sampling from the initial state before RL launched. Second, a novel criterion named max-CIDEr is proposed to serve as the reward for promoting diversity. We combine the above-mentioned offpolicy strategy with the on-policy one to moderate the exploration effect, further balancing the diversity and accuracy for human-like image captioning. Experiments show that our method locates the closest to human performance in the diversity-accuracy space, and achieves the highest Pearson correlation as 0.337 with human performance. Jiahe Shi, Yali Li 0001, Shengjin Wang |
ICCV | 3 |
| 2021 | Mask Scene Text Recognizer
Haodong Shi, Liangrui Peng, Ruijie Yan, Shuman Han, Shengjin Wang |
ICDAR (4) | 6 |
| 2021 | Combating Noise: Semi-supervised Learning by Region Uncertainty QuantificationabstractSemi-supervised learning aims to leverage a large amount of unlabeled data for performance boosting. Existing works primarily focus on image classification. In this paper, we delve into semi-supervised learning for object detection, where labeled data are more labor-intensive to collect. Current methods are easily distracted by noisy regions generated by pseudo labels. To combat the noisy labeling, we propose noise-resistant semi-supervised learning by quantifying the region uncertainty. We first investigate the adverse effects brought by different forms of noise associated with pseudo labels. Then we propose to quantify the uncertainty of regions by identifying the noise-resistant properties of regions over different strengths. By importing the region uncertainty quantification and promoting multi-peak probability distribution output, we introduce uncertainty into training and further achieve noise-resistant learning. Experiments on both PASCAL VOC and MS COCO demonstrate the extraordinary performance of our method. Zhenyu Wang 0005, Yali Li 0001, Shengjin Wang |
NeurIPS | 4 |
| 2021 | Do Different Tracking Tasks Require Different Appearance Models?abstractTracking objects of interest in a video is one of the most popular and widely applicable problems in computer vision. However, with the years, a Cambrian explosion of use cases and benchmarks has fragmented the problem in a multitude of different experimental setups. As a consequence, the literature has fragmented too, and now novel approaches proposed by the community are usually specialised to fit only one specific setup. To understand to what extent this specialisation is necessary, in this work we present UniTrack, a solution to address five different tasks within the same framework. UniTrack consists of a single and task-agnostic appearance model, which can be learned in a supervised or self-supervised fashion, and multiple ``heads'' that address individual tasks and do not require training. We show how most tracking tasks can be solved within this framework, and that the same appearance model can be successfully used to obtain results that are competitive against specialised methods for most of the tasks considered. The framework also allows us to analyse appearance models obtained with the most recent self-supervised methods, thus extending their evaluation and comparison to a larger variety of important problems. Zhongdao Wang, Hengshuang Zhao, Yali Li 0001, Shengjin Wang, Philip Torr 0001, Luca Bertinetto |
NeurIPS | 4 |
| 2021 | EcRD: Edge-Cloud Computing Framework for Smart Road Damage Detection and WarningabstractRoad damages have caused numerous fatalities, thus the study of road damage detection, especially hazardous road damage detection and warning is critical for traffic safety. Existing road damage detection systems mainly process data at cloud, which suffers from a high latency caused by long-distance. Meanwhile, supervised machine learning algorithms are usually used in these systems requiring large precisely labeled data sets to achieve a good performance. In this article, we propose EcRD: an edge-cloud-based road damage detection and warning framework, that leverages the fast-responding advantage of edge and the large storage and computation resources advantages of cloud. There are three main contributions in this article: we first propose a simple yet efficient road segmentation algorithm to enable fast and accurate road area detection. Then, a light-weighted road damage detector is developed based on gray level co-occurrence matrix features at edge for rapid hazardous road damage detection and warning. Furthermore, a multitypes road damage detection model is introduced for long-term road management at cloud, embedded with a novel image generator based on cycle-consistent adversarial networks which automatically generates images with labels to further improve road damage detection accuracy. By comparing with the state-of-the-art, we demonstrate that the proposed EcRD can accurately detect both hazardous road damages at edge and multitypes road damages at cloud. Besides, it is around 579 times faster than cloud-based approaches without affecting users' experience and requiring very low storage and labeling cost. Yachao Yuan, Md. Saiful Islam 0011, Yali Yuan, Shengjin Wang, Thar Baker, Lutz M. Kolbe |
IEEE Internet Things J. | 4 |
| 2021 | Learning Part-based Convolutional Features for Person Re-IdentificationabstractPart-level features offer fine granularity for pedestrian image description. In this article, we generally aim to learn discriminative part-informed feature for person re-identification. Our contribution is two-fold. First, we introduce a general part-level feature learning method, named Part-based Convolutional Baseline (PCB). Given an image input, it outputs a convolutional descriptor consisting of several part-level features. PCB is general in that it is able to accommodate several part partitioning strategies, including pose estimation, human parsing and uniform part partitioning. In experiment, we show that the learned descriptor has a significantly higher discriminative ability than the global descriptor. Second, based on PCB, we propose refined part pooling (RPP), which allows the parts to be more precisely located. Our idea is that pixels within a well-located part should be similar to each other while being dissimilar with pixels from other parts. We call it within-part consistency. When a pixel-wise feature vector in a part is more similar to some other part, it is then an outlier, indicating inappropriate partitioning. RPP re-assigns these outliers to the parts they are closest to, resulting in refined parts with enhanced within-part consistency. RPP requires no part labels and is trained in a weakly supervised manner. Experiment confirms that RPP allows PCB to gain another round of performance boost. For instance, on the Market-1501 dataset, we achieve (77.4+4.2) percent mAP and (92.3+1.5) percent rank-1 accuracy, a competitive performance with the state of the art. Yifan Sun 0003, Liang Zheng 0001, Yali Li 0001, Yi Yang 0001, Qi Tian 0001, Shengjin Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Learning to Segment Video Object With Accurate BoundariesabstractVideo object segmentation has attracted considerable research interest these years. Top-performing video object segmentation methods mainly rely on fully convolutional neural networks which are specifically trained for predicting high-performance masks, resulting in a lack of preciseness in boundary details. This paper tackles the problem of predicting both mask-accurate and boundary-precise segmentation masks in videos. To solve this problem, we propose a simple and efficient network structure: the Mask-boundAry-Consistent Network (MAC-Net). TheMAC-Netis an end-to-end fully convolutional network, where both mask and boundaries are jointly optimized during training, enabling it to predict masks along with accurate boundaries. An inner-net boundary-computing module is incorporated in theMAC-Netfor producing spontaneously mask-consistent boundaries. We analyze the influence of parameter settings, network constructions of theMAC-Net, and compare with state-of-the-art algorithms on three widely-adopted datasets. Experimental results show that theMAC-Netachieves state-of-the-art performance, demonstrating the effectiveness of its mask-boundary-consistent network structure. We also propose that the boundary module inMAC-Nethas high compatibility, and can be easily adapted to other segmentation-related techniques. Jingchun Cheng, Yuhui Yuan, Yali Li 0001, Jingdong Wang 0001, Shengjin Wang |
IEEE Trans. Multim. | 5 |
| 2020 | Softmax Dissection: Towards Understanding Intra- and Inter-Class Objective for Embedding LearningabstractThe softmax loss and its variants are widely used as objectives for embedding learning applications like face recognition. However, the intra- and inter-class objectives in Softmax are entangled, therefore a well-optimized inter-class objective leads to relaxation on the intra-class objective, and vice versa. In this paper, we propose to dissect Softmax into independent intra- and inter-class objective (D-Softmax) with a clear understanding. It is straightforward to tune each part to the best state with D-Softmax as objective.Furthermore, we find the computation of the inter-class part is redundant and propose sampling-based variants of D-Softmax to reduce the computation cost. The face recognition experiments on regular-scale data show D-Softmax is favorably comparable to existing losses such as SphereFace and ArcFace. Experiments on massive-scale data show the fast variants significantly accelerates the training process (such as 64×) with only a minor sacrifice in performance, outperforming existing acceleration methods of Softmax in terms of both performance and efficiency. Lanqing He, Zhongdao Wang, Yali Li 0001, Shengjin Wang |
AAAI | 4 |
| 2020 | Revisiting Temporal Modeling for Video Super-resolution
Takashi Isobe, Shengjin Wang |
BMVC | 3 |
| 2020 | Video Super-Resolution With Temporal Group AttentionabstractVideo super-resolution, which aims at producing a high-resolution video from its corresponding low-resolution version, has recently drawn increasing attention. In this work, we propose a novel method that can effectively incorporate temporal information in a hierarchical way. The input sequence is divided into several groups, with each one corresponding to a kind of frame rate. These groups provide complementary information to recover missing details in the reference frame, which is further integrated with an attention module and a deep intra-group fusion module. In addition, a fast spatial alignment is proposed to handle videos with large motion. Extensive results demonstrate the capability of the proposed model in handling videos with various motion. It achieves favorable performance against state-of-the-art methods on several benchmark datasets. Takashi Isobe, Songjiang Li, Xu Jia 0012, Shanxin Yuan, Gregory Slabaugh, Chunjing Xu, Yali Li 0001, Shengjin Wang, Qi Tian 0001 |
CVPR | 8 |
| 2020 | Video Super-Resolution with Recurrent Structure-Detail Network
Takashi Isobe, Xu Jia 0012, Shuhang Gu, Songjiang Li, Shengjin Wang, Qi Tian 0001 |
ECCV (12) | 5 |
| 2020 | Towards Real-Time Multi-Object Tracking
Zhongdao Wang, Liang Zheng 0001, Yixuan Liu 0004, Yali Li 0001, Shengjin Wang |
ECCV (11) | 5 |
| 2020 | CycAs: Self-supervised Cycle Association for Learning Re-identifiable Descriptions
Zhongdao Wang, Liang Zheng 0001, Yixuan Liu 0004, Yifan Sun 0003, Yali Li 0001, Shengjin Wang |
ECCV (11) | 7 |
| 2020 | CS-R-FCN: Cross-Supervised Learning for Large-Scale Object DetectionabstractGeneric object detection is one of the most fundamental problems in computer vision, yet it is difficult to provide all the bounding-box-level annotations aiming at large-scale object detection for thousands of categories. In this paper, we present a novel cross-supervised learning pipeline for large-scale object detection, denoted as CS-R-FCN. First, we propose to utilize the data flow of image-level annotated images in the fully-supervised two-stage object detection framework, leading to cross-supervised learning combining bounding-box-level annotated data and image-level annotated data. Second, we introduce a semantic aggregation strategy utilizing the relationships among the cross-supervised categories to reduce the unreasonable mutual inhibition effects during the feature learning. Experimental results show that the proposed CS-R-FCN improves the mAP by a large margin compared to previous related works. Yali Li 0001, Shengjin Wang |
ICASSP | 3 |
| 2020 | Intra-Clip Aggregation For Video Person Re-IdentificationabstractVideo-based person re-identification has drawn massive attention in recent years due to its extensive applications in video surveillance. While deep learning based methods have led to significant progress, these methods are limited by ineffectively using complementary information, which is blamed on necessary data augmentation in training process. Data augmentation has been widely used to mitigate the overfitting trap and improve the ability of network representation. However, the previous methods adopt image-based data augmentation scheme to individually process the input frames, which corrupts the complementary information between consecutive frames and causes performance degradation. In this paper, we propose a novel video-based data augmentation scheme, termed as Synchronous Data Augmentation, to address the challenge above. In order to represent discriminative clip-level features, we also propose a cascade integration module which hierarchically aggregates the intra-clip features with a linear-nonlinear combining projection. Extensive experiments on three benchmark datasets demonstrate that our framework outperforms the most recent state-of-the-art methods. We also perform cross-dataset validation to prove the generality of our method. Takashi Isobe, Yali Li 0001, Shengjin Wang |
ICIP | 5 |
| 2020 | Fianet: Video Object Detection Via Joint Feature-Level And Instance-Level AggregationabstractVideo object detection task is challenging due to the nonrigid and rigid appearance deformations in videos. Most of the typical competitive methods are to enhance per-frame features through aggregating lots of previous and future frames. But feature-level aggregation isn't robust to rigid deformations such as occlusion and rare postures. In this paper, we propose an online video object detection method with joint feature-level aggregation and instance-level aggregation network (FIANet). Besides feature-level aggregation, we design a spatial-temporal instance calibration module (STIC) to aggregate the instance as a whole, which can reduce the interference of local distorted and missed pixels. Joint featurelevel and instance-level aggregation can work collaboratively to overcome different deformations. Only using less previous frames, our method can achieve 81.6% mAP with relatively high speed on ImageNet VID, which is state-of-the-art compared with causal and non-causal methods. Zhengshuai Wang, Yali Li 0001, Shengjin Wang |
ICME | 3 |
| 2020 | A Human-Computer Fusion Framework for Aircraft Recognition in Remote Sensing ImagesabstractExisting aircraft recognition methods usually regard recognition as an isolated classification problem, supposing that aircraft detection has finished. These methods use image slices each containing single aircraft as input, which is often not the case in practice. In order to recognize aircraft in remote sensing images that contain multiple objects and background, we propose a human-computer fusion framework that combines the advantages of human and computer. First, we propose candidate aircraft using human eye tracking, making use of the efficient and accurate search ability of human. Then, we propose a two-step recognition method, which simulates the recognition process of image analysts, to identify the types of candidate aircraft. In the first step, we recognize aircraft using fully connected features of a fine-tuned convolutional neural network (CNN). If it does not work, we go to the second step. In the second step, we use convolutional features extracted from the fine-tuned CNN for recognition. To improve the representational ability of convolutional features, we introduce the Fisher Vector encoding. Thorough experiments demonstrate that the proposed framework is effective, surpassing state-of-the-art methods on recognition performance. Xiaobin Li 0005, Bitao Jiang, Shengjin Wang, Yuze Fu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Progressive Representation Adaptation for Weakly Supervised Object LocalizationabstractWe address the problem of weakly supervised object localization where only image-level annotations are available for training object detectors. Numerous methods have been proposed to tackle this problem through mining object proposals. However, a substantial amount of noise in object proposals causes ambiguities for learning discriminative object models. Such approaches are sensitive to model initialization and often converge to undesirable local minimum solutions. In this paper, we propose to overcome these drawbacks by progressive representation adaptation with two main steps: 1) classification adaptation and 2) detection adaptation. In classification adaptation, we transfer a pre-trained network to a multi-label classification task for recognizing the presence of a certain object in an image. Through the classification adaptation step, the network learns discriminative representations that are specific to object categories of interest. In detection adaptation, we mine class-specific object proposals by exploiting two scoring strategies based on the adapted classification network. Class-specific proposal mining helps remove substantial noise from the background clutter and potential confusion from similar objects. We further refine these proposals using multiple instance learning and segmentation cues. Using these refined object bounding boxes, we fine-tune all the layer of the classification network and obtain a fully adapted detection network. We present detailed experimental validation on the PASCAL VOC and ILSVRC datasets. Experimental results demonstrate that our progressive representation adaptation algorithm performs favorably against the state-of-the-art methods. Dong Li 0025, Jia-Bin Huang 0001, Yali Li 0001, Shengjin Wang, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Adaptively Leverage Unlabeled Tracklets Based on Part Attention Model for Few-Example Re-IDabstractFew-example learning for video person re-ID is a challenging issue. Some studies use large unlabeled samples to mine more discriminative cues to overcome the visual scarcity. But how to develop a robust model to avoid overfitting and overcome noisy labels is still a remained problem. In this letter we focus on bridging between few labeled tracklets and numerous unlabeled tracklets. We leverage unlabeled tracklets by generating pseudo labels and adaptively joins them into the training set which consists of the labeled and unlabeled data. This work is distinguished by two key contributions. First, a novel model PAM (Part Attention Model) tailored for progressive learning is proposed. It has high accuracy and fast convergence. Second, we propose a novel sampling strategy ARS (Adaptively Relative Sampling). ARS adaptively filters out noisy labels and enlarges the training set. ARS-PAM achieves significant performance gains on four mainstream datasets. On the PRID2011 and iLIDS-VID dataset, ARS-PAM reaches 89.8%, 56.1% rank-1 accuracy, which exceeds the state-of-the-art by 5.5%, 17.5% respectively. Yali Li 0001, Shengjin Wang |
IEEE Signal Process. Lett. | 3 |
| 2020 | Node-Adaptive Multi-Graph Fusion Using Extreme Value TheoryabstractThis letter considers the problem of grouping data by their underlying categories with inputs from multiple sources, known as the multi-view clustering problem. One of the most fundamental challenges lies in how to benefit from the complementary information in the multi-view data, so that clustering on such data consistently achieves higher accuracy than clustering on each single-view component. In this letter, to tackle the multi-view clustering problem, we propose a novel approach to fuse multiple affinity graphs computed in each single view to a unified affinity graph, so that single-view affinity-based clustering methods can be accordingly applied on it. The edges in the unified affinity graph between a node and its neighbors are computed as weighted average over the corresponding edges from multiple single graphs, and the weights here are adaptive to each node, estimated using the Extreme Value Theory (EVT). Experiments on two challenging multi-view clustering tasks show that, combined with existing off-the-shelf single-view clustering algorithms, the proposed graph fusion method brings consistently performance gain compared with naive graph fusion baselines. Zhongdao Wang, Yali Li 0001, Shengjin Wang |
IEEE Signal Process. Lett. | 4 |
| 2020 | Unidirectional Representation-Based Efficient Dictionary LearningabstractDictionary learning (DL) has been widely studied for pattern classification. Most existing methods introduce multiple discriminative terms into objective functions for accuracy improvement, leading to complex learning frameworks and high computational burdens. This paper proposes a simple yet effective DL algorithm for classification, namely unidirectional representation dictionary learning (URDL). Unidirectional constraint is proposed to guide coefficient directions in the representation to be discriminative. Besides, direction-thresholding is proposed to exploit the direction property in the classification scheme. It suppresses the disturbance from undesired non-zero coefficients, and improves the representation discriminability. We adopt squared ℓ2-norm-based regularization for efficient coding, and systematically analyze the mechanism of the proposed method. Extensive experiments on five data sets are conducted, including object categorization, scene classification, face recognition, and fine-grained flower classification. The experimental results demonstrate that the proposed approach not only outperforms the state-of-the-art DL algorithms in terms of recognition accuracy significantly, but also exhibits a much higher computational efficiency. Xiudong Wang, Yali Li 0001, Shaodi You, Hongdong Li, Shengjin Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | HAR-Net: Joint Learning of Hybrid Attention for Single-Stage Object DetectionabstractObject detection has been a challenging task in computer vision. Although significant progress has been made in object detection with deep neural networks, the attention mechanism has yet to be fully developed. In this paper, we propose a hybrid attention mechanism for single-stage object detection. First, we present the modules of spatial attention, channel attention and aligned attention for single-stage object detection. In particular, dilated convolution layers with symmetrically fixed rates are stacked to learn spatial attention. A channel attention mechanism with the cross-level group normalization and squeeze-and-excitation operation is proposed. Aligned attention is constructed with organized deformable filters. Second, the three types of attention are unified to construct the hybrid attention mechanism. We then plug the hybrid attention into Retina-Net and propose the efficient single-stage HAR-Net for object detection. The attention modules and the proposed HAR-Net are evaluated on the COCO detection dataset. The experiments demonstrate that hybrid attention can significantly improve the detection accuracy and that the HAR-Net can achieve a state-of-the-art 45.8% mAP, thus outperforming existing single-stage object detectors. Yali Li 0001, Shengjin Wang |
IEEE Trans. Image Process. | 2 |
| 2020 | Fast Pedestrian Detection With Attention-Enhanced Multi-Scale RPN and Soft-Cascaded Decision TreesabstractPedestrian detection has attracted more attention in the fields of computer vision and artificial intelligence. A variety of real-world applications involving pedestrian detection have been promoted, such as Advanced Driving Assistant System (ADAS). Although both two-stage and single-stage deeply learned object detectors have shown outstanding performance for general object detection, they are still facing the problem of poor accuracy in single-class detection senario because they are designed to distinguish objects from different categories rather than pay attention to various appearances of pedestrians. Previous leading pedestrian detectors F-DNN and F-DNN v2 fuse several neural networks like SSD, VGG16 and GoogLeNet to generate ROIs and supress false alarms with cascaded structure, resulting in low miss rate but high complexity. In this paper we propose a novel framework called Attention-Enhanced Multi-Scale Region Proposal Network (AEMS-RPN) for ROI generation, which also acts as first-stage classification. Inspired by the success of traditional pedestrian detectors, we use soft-cascaded decision trees instead of cascaded deep neural networks to achieve high accuracy and fast detection speed simultaneously. The decision tree classifier is used and enables us to combine features from different layers with various resolutions for classification and incorporate effective bootstrapping for mining hard negatives. We test our method on several pedestrian detection datasets and the experimental results certify the effectiveness of the proposed AEMS-RPN. Compared with the state-of-the-art, we obtain the competitive accuracy with near real-time efficiency. Yali Li 0001, Shengjin Wang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Adversarial View-Consistent Learning for Monocular Depth Estimation
Yuwang Wang, Shengjin Wang |
BMVC | 3 |
| 2019 | Perceive Where to Focus: Learning Visibility-Aware Part-Level Features for Partial Person Re-IdentificationabstractThis paper considers a realistic problem in person re-identification (re-ID) task, i.e., partial re-ID. Under partial re-ID scenario, the images may contain a partial observation of a pedestrian. If we directly compare a partial pedestrian image with a holistic one, the extreme spatial misalignment significantly compromises the discriminative ability of the learned representation. We propose a Visibility-aware Part Model (VPM) for partial re-ID, which learns to perceive the visibility of regions through self-supervision. The visibility awareness allows VPM to extract region-level features and compare two images with focus on their shared regions (which are visible on both images). VPM gains two-fold benefit toward higher accuracy for partial re-ID. On the one hand, compared with learning a global feature, VPM learns region-level features and thus benefits from fine-grained information. On the other hand, with visibility awareness, VPM is capable to estimate the shared regions between two images and thus suppresses the spatial misalignment. Experimental results confirm that our method significantly improves the learned feature representation and the achieved accuracy is on par with the state of the art. Yifan Sun 0003, Yali Li 0001, Chi Zhang 0026, Shengjin Wang, Jian Sun 0001 |
CVPR | 6 |
| 2019 | Linkage Based Face Clustering via Graph Convolution NetworkabstractIn this paper, we present an accurate and scalable approach to the face clustering task. We aim at grouping a set of faces by their potential identities. We formulate this task as a link prediction problem: a link exists between two faces if they are of the same identity. The key idea is that we find the local context in the feature space around an instance (face) contains rich information about the linkage relationship between this instance and its neighbors. By constructing sub-graphs around each instance as input data, which depict the local context, we utilize the graph convolution network (GCN) to perform reasoning and infer the likelihood of linkage between pairs in the sub-graphs. Experiments show that our method is more robust to the complex distribution of faces than conventional methods, yielding favorably comparable results to state-of-the-art methods on standard face clustering benchmarks, and is scalable to large datasets. Furthermore, we show that the proposed method does not need the number of clusters as prior, is aware of noises and outliers, and can be extended to a multi-view version for more accurate clustering accuracy. Zhongdao Wang, Liang Zheng 0001, Yali Li 0001, Shengjin Wang |
CVPR | 4 |
| 2019 | Intention Oriented Image Captions With Guiding ObjectsabstractAlthough existing image caption models can produce promising results using recurrent neural networks (RNNs), it is difficult to guarantee that an object we care about is contained in generated descriptions, for example in the case that the object is inconspicuous in the image. Problems become even harder when these objects did not appear in training stage. In this paper, we propose a novel approach for generating image captions with guiding objects (CGO). The CGO constrains the model to involve a human-concerned object when the object is in the image. CGO ensures that the object is in the generated description while maintaining fluency. Instead of generating the sequence from left to right, we start the description with a selected object and generate other parts of the sequence based on this object. To achieve this, we design a novel framework combining two LSTMs in opposite directions. We demonstrate the characteristics of our method on MSCOCO where we generate descriptions for each detected object in the images. With CGO, we can extend the ability of description to the objects being neglected in image caption labels and provide a set of more comprehensive and diverse descriptions for an image. CGO shows advantages when applied to the task of describing novel objects. We show experimental results on both MSCOCO and ImageNet datasets. Evaluations show that our method outperforms the state-of-the-art models in the task with average F1 75.8, leading to better descriptions in terms of both content accuracy and fluency. Yali Li 0001, Shengjin Wang |
CVPR | 3 |
| 2019 | Deep Network with Pixel-Level Rectification and Robust Training for Handwriting RecognitionabstractOffline handwriting recognition is a well-known challenging task in the optical character recognition (OCR) field due to the difficulty caused by various unconstraint handwriting styles. In order to learn invariant feature representations for handwriting, we propose a novel method to incorporate pixel-level rectification into a CNN and RNN based model. We also propose an adjacent output mixup method for RNN layer's training to improve the generalization ability of the model, i.e., the previous output of an RNN layer is added to the current output with random weights. We additionally adopt a series of techniques including pre-training, data augmentation and language model, and further analyze their contributions to the improvement of the model performance. The proposed method performs well on three public benchmarks, including the IAM, Rimes and IFN/ENIT datasets. Shanyu Xiao, Liangrui Peng, Ruijie Yan, Shengjin Wang |
ICDAR | 4 |
| 2019 | Parallel-Structure-based Transfer Learning for Deep NIR-to-VIS Face Recognition
Yali Li 0001, Shengjin Wang |
ICIG (1) | 3 |
| 2019 | Cascade Attention: Multiple Feature Based Learning for Image CaptioningabstractMost recent researches in image captioning adopt attention mechanism based on encoder-decoder framework, where the attention module aligns input features for the decoder and boosts performance consequently. A common defect of traditional attention methods is that the inequality among different types of inputs is ignored, resulting in under-exploitation of certain informative features. In this paper, we propose a novel cascade attention module, which processes different types of input in a sequential manner. The cascade attention module enables inputs of higher priorities to affect the attention of other inputs so as to emphasize such inequality. We implement our model by introducing global feature of the image to the captioning process of R-CNN based frameworks, where such feature is rich of context information but takes few effects via traditional attention module. Experimental results demonstrate that our proposed method is able to exploit feature of different types, acquiring improvements on multiple automatic measurements. Jiahe Shi, Yali Li 0001, Shengjin Wang |
ICIP | 3 |
| 2019 | Dynamic temporal residual network for sequence modeling
Ruijie Yan, Liangrui Peng, Shanyu Xiao, Michael T. Johnson, Shengjin Wang |
Int. J. Document Anal. Recognit. | 5 |
| 2019 | Secondary Information Aware Facial Expression RecognitionabstractFacial expression recognition (FER) is a key factor in human behavior analysis. Most algorithms deal with FER as a pure classification problem, assuming that expressions are exclusive to each other. In this letter, the problem of FER is tackled from a more detailed view: learning to discriminate expressions with consideration of the secondary information. We propose the Secondary Information aware Facial Expression Network (SIFE-Net) to explore the latent components without auxiliary labeling, and we propose a novel dynamic weighting strategy to teach the SIFE-Net. In contrast to traditional classifiers trained with one-hot labels, the proposed SIFE-Net takes advantage of secondary expression information and has more rational feature distributions. We carry out extensive experiments and analysis on three widely-used FER datasets, i.e. the CK+ dataset, the JAFFE dataset, and the RAF dataset. Experimental results show that the SIFE-Net achieves state-of-the-art performance on all three datasets, which demonstrates the effectiveness of our method. Ye Tian 0019, Jingchun Cheng, Yali Li 0001, Shengjin Wang |
IEEE Signal Process. Lett. | 4 |
| 2019 | Open-World Person Re-Identification With Deep Hash Feature EmbeddingabstractMost existing person re-identification (re-id) methods are designed based on the artificial closed-set assumption that the probe and gallery identities are exactly overlapped with a small search pool. This leads to poor scalability in real-world applications where the task is often to re-id a small set of target people (i.e., watch-list) among a large search pool with unknown ID overlap, namely, an open-set deployment setting. In this paper, we firstly propose a new person re-id setting called Watch-List based Open-Set (WLOS) person re-id, which is characterised by the above open-set deployment and a watch-list available at the training stage. Then, we address such a under-studied WLOS problem by formulating a novel Task Dedicated Deep Hashing (TDDH) approach which learning a purpose-specific deep hash model particularly for the given target people in an efficient end-to-end manner. Extensive experiments on three large-scale re-id benchmarks are conducted to demonstrate the advantages and superiority of the TDDH over a wide range of the state-of-the-art hashing and re-id methods under the more realistic open-set setting. Yali Zhao, Yali Li 0001, Shengjin Wang |
IEEE Signal Process. Lett. | 3 |
| 2019 | Unsupervised Deep Hashing With Adaptive Feature Learning for Image RetrievalabstractThe hashing method is widely used for large-scale image retrieval due to its low time and space complexity. However, the existing deep hashing methods are mainly designed for labeled datasets. Without supervised information, retrieval performance on unlabeled datasets is not guaranteed. In this letter, we propose a novel deep hashing approach for unsupervised image retrieval applications. The contributions are two-fold. First, the pseudolabels are generated using their global features aggregated from the pretrained network and employed as self-supervised information to optimize the objective function of training. Second, adaptive feature learning is used in this deep hashing framework to perform simultaneous hash function learning and global features learning in an unsupervised manner. The experimental results validated the effectiveness of the proposed method, obtaining state-of-the-art performances on several public datasets such as CIFAR-10, Holidays, and Oxford5k. Yali Li 0001, Shengjin Wang |
IEEE Signal Process. Lett. | 3 |
| 2019 | Real-Time Head Pose Estimation and Face Modeling From a Depth ImageabstractWe address the issues of 3-D head pose estimation and face modeling from a depth image. Given a depth image, random forests are effective for estimating the location and orientation of a person's head. However, the accuracy of the estimation is not high enough. We propose using corrected regression votes. The corrected votes are obtained by considering the cooperation of all trees, leading to significant improvement of head pose estimation accuracy. Based on the head pose estimator, we present a face modeling system. In our system, the face model is generated by aligning a deformable face model to a depth image using an iterative closest point (ICP) algorithm. The novelty of our approach is that an optimal weight for each vertex is incorporated into the ICP algorithm with point to plane constraints. Experiments show that our system can automatically estimate the head pose and generate a realistic face model from a single depth image. We also provide a detailed evaluation that shows the benefits of our approach. Changwei Luo, Juyong Zhang, Jun Yu 0001, Chang Wen Chen, Shengjin Wang |
IEEE Trans. Multim. | 5 |
| 2018 | Learning from PhotoShop Operation Videos: The PSOV Dataset
Jingchun Cheng, Han-Kai Hsu, Hailin Jin, Shengjin Wang, Ming-Hsuan Yang 0001 |
ACCV (4) | 5 |
| 2018 | A Highly Accurate Feature Fusion Network For Vehicle Detection In Surveillance Scenarios
Yali Li 0001, Shengjin Wang |
BMVC | 3 |
| 2018 | Fast and Accurate Online Video Object Segmentation via Tracking PartsabstractOnline video object segmentation is a challenging task as it entails to process the image sequence timely and accurately. To segment a target object through the video, numerous CNN-based methods have been developed by heavily finetuning on the object mask in the first frame, which is time-consuming for online applications. In this paper, we propose a fast and accurate video object segmentation algorithm that can immediately start the segmentation process once receiving the images. We first utilize a part-based tracking method to deal with challenging factors such as large deformation, occlusion, and cluttered background. Based on the tracked bounding boxes of parts, we construct a region-of-interest segmentation network to generate part masks. Finally, a similarity-based scoring function is adopted to refine these object parts by comparing them to the visual information in the first frame. Our method performs favorably against state-of-the-art algorithms in accuracy on the DAVIS benchmark dataset, while achieving much faster runtime performance. Jingchun Cheng, Yi-Hsuan Tsai, Wei-Chih Hung, Shengjin Wang, Ming-Hsuan Yang 0001 |
CVPR | 4 |
| 2018 | Beyond Part Models: Person Retrieval with Refined Part Pooling (and A Strong Convolutional Baseline)
Yifan Sun 0003, Liang Zheng 0001, Yi Yang 0001, Qi Tian 0001, Shengjin Wang |
ECCV (4) | 5 |
| 2018 | Feature Learning for One-Shot Face RecognitionabstractOne-shot face recognition is a challenging open problem which requires recognizing novel identities from only one gallery face. One-shot classes are squeezed and neglected in the feature space for classification due to data imbalance. Moreover, training samples deficience is a major obstacle to intra-class clustering. In this paper, we propose a novel framework based on CNN of balancing regularizer and shifting center regeneration which regulates norms of weight vector into same scale and adjusts clustering center to deal with deficient training data. Comprehensive evaluations on MS-celeb-1M low-shot face dataset demonstrate that our methods improve one-shot face recognition notablely which achieve 88.78% coverage at precision=0.99 using restricted data without hybrid classifiers or multi-model. Moreover, experiments on LFW prove that CNN model trained with proposed methods can obtain more discriminative and compact feature representations. Since there are many identities that have only few training samples available online, our methods have great significance for improving data utilization and strengthening feature representation for face recognition. Lingxiao Wang 0009, Yali Li 0001, Shengjin Wang |
ICIP | 3 |
| 2018 | Attend and Align: Improving Deep Representations with Feature Alignment Layer for Person RetrievalabstractIn fine-grained recognition, object misalignment and background noise are two long-standing factors that influence the robustness of deep learning models. This paper mainly focuses on person re-identification (re-ID) and introduces a feature alignment layer (FAL) which alleviates the target misalignment and the background noise simultaneously. Through attention mechanism, FAL informs the underlying importance of each pixel on feature maps, i.e., whether the pixel is beneficial towards discriminating different persons. Then the discriminative regions relocate to the center and are stretched to fill the feature maps. Such an “attend and align” mechanism is specified into two steps: target position prediction and value assignment. In the first step, a pixel on feature maps learns to find a target position which is ID-discriminative. In the second step, the pixel is assigned with a new value using the context of the predicted position. Moreover, FAL can be easily plugged into a canonical Convolutional Neural Network (CNN) and learned in an end-to-end manner. In experiment, our method yields competitive results compared with the state-of-the-art approaches on three person re-ID datasets, Market-1501, DukeMTMC-reID and CUHK03. We also demonstrate that our method improves a competitive fine-grained recognition baseline on CUB-200-2011. Yifan Sun 0003, Yali Li 0001, Shengjin Wang |
ICPR | 4 |
| 2017 | SegFlow: Joint Learning for Video Object Segmentation and Optical FlowabstractThis paper proposes an end-to-end trainable network, SegFlow, for simultaneously predicting pixel-wise object segmentation and optical flow in videos. The proposed SegFlow has two branches where useful information of object segmentation and optical flow is propagated bidirectionally in a unified framework. The segmentation branch is based on a fully convolutional network, which has been proved effective in image segmentation task, and the optical flow branch takes advantage of the FlowNet model. The unified framework is trained iteratively offline to learn a generic notion, and fine-tuned online for specific objects. Extensive experiments on both the video object segmentation and optical flow datasets demonstrate that introducing optical flow improves the performance of segmentation and vice versa, against the state-of-the-art algorithms. Jingchun Cheng, Yi-Hsuan Tsai, Shengjin Wang, Ming-Hsuan Yang 0001 |
ICCV | 3 |
| 2017 | SVDNet for Pedestrian RetrievalabstractThis paper proposes the SVDNet for retrieval problems, with focus on the application of person re-identification (reID). We view each weight vector within a fully connected (FC) layer in a convolutional neuron network (CNN) as a projection basis. It is observed that the weight vectors are usually highly correlated. This problem leads to correlations among entries of the FC descriptor, and compromises the retrieval performance based on the Euclidean distance. To address the problem, this paper proposes to optimize the deep representation learning process with Singular Vector Decomposition (SVD). Specifically, with the restraint and relaxation iteration (RRI) training scheme, we are able to iteratively integrate the orthogonality constraint in CNN training, yielding the so-called SVDNet. We conduct experiments on the Market-1501, CUHK03, and DukeMTMC-reID datasets, and show that RRI effectively reduces the correlation among the projection vectors, produces more discriminative FC descriptors, and significantly improves the re-ID accuracy. On the Market-1501 dataset, for instance, rank-1 accuracy is improved from 55.3% to 80.5% for CaffeNet, and from 73.8% to 82.3% for ResNet-50. Yifan Sun 0003, Liang Zheng 0001, Weijian Deng, Shengjin Wang |
ICCV | 4 |
| 2017 | Orientation Invariant Feature Embedding and Spatial Temporal Regularization for Vehicle Re-identificationabstractIn this paper, we tackle the vehicle Re-identification (ReID) problem which is of great importance in urban surveillance and can be used for multiple applications. In our vehicle ReID framework, an orientation invariant feature embedding module and a spatial-temporal regularization module are proposed. With orientation invariant feature embedding, local region features of different orientations can be extracted based on 20 key point locations and can be well aligned and combined. With spatial-temporal regularization, the log-normal distribution is adopted to model the spatial-temporal constraints and the retrieval results can be refined. Experiments are conducted on public vehicle ReID datasets and our proposed method achieves state-of-the-art performance. Investigations of the proposed framework is conducted, including the landmark regressor and comparisons with attention mechanism. Both the orientation invariant feature embedding and the spatio-temporal regularization achieve considerable improvements. Zhongdao Wang, Luming Tang, Xihui Liu, Zhuliang Yao, Shuai Yi, Shengjin Wang, Hongsheng Li 0001, Xiaogang Wang 0001 |
ICCV | 8 |
| 2017 | Residual Recurrent Neural Network with Sparse Training for Offline Arabic Handwriting RecognitionabstractDeep Recurrent Neural Networks (RNN) have been suffering from the overfitting problem due to the model redundancy of the network structures. We propose a novel temporal and spatial residual learning method for RNN, followed with sparse training by weight pruning to gain sparsity in network parameters. For a Long Short-Term Memory (LSTM) network, we explore the combination schemes and parameter settings for temporal and spatial residual learning with sparse training. Experiments are carried out on the IFN/ENIT database. For the character error rate on the testing set e while training with sets a, b, c, d, the previously reported best result is 13.42%, and the proposed configuration of temporal residual learning followed with sparse training achieves the state-of-the-art result 12.06%. Ruijie Yan, Liangrui Peng, GuangXiang Bin, Shengjin Wang |
ICDAR | 4 |
| 2017 | A highly accurate facial region network for unconstrained face detectionabstractIn this paper, a new face detection method with very high accuracy is proposed. We introduce a novel facial region network to detect faces in unconstrained conditions. Firstly, a face proposal net is raised to generate possible face regions in the input image. Then, a novel weighted grid feature is applied to calculate features of face regions. Owing to that, faces with large pose variation and severe occlusion can be detected correctly. Furthermore, we use millions of general object data to pre-train the network to enhance the robustness of the extracted feature. Our method is evaluated on several public face detection datasets and achieves state-of-the-art performance on all of them. Specially, our method demonstrates a very high recall rate of 96.4% when false positives are 300 on the challenging FDDB benchmark, ranking first not only in the academic list but also in the commercial list which is much more competitive than the previous one. Han Shu, Dangdang Chen, Yali Li 0001, Shengjin Wang |
ICIP | 4 |
| 2017 | Object Detection Using Convolutional Neural Networks in a Coarse-to-Fine MannerabstractObject detection in remote sensing images has long been studied, but it remains challenging due to the diversity of objects and the complexity of backgrounds. In this letter, we propose an object detection method using convolutional neural networks (CNNs) in a coarse-to-fine manner. In the coarse step, coarse candidate regions that may contain objects are proposed. In the fine step, fine candidate regions are cropped from coarse candidate regions, and are classified as objects or backgrounds. We design a concise and efficient framework that can propose fewer candidate regions and extract more discriminative features. The framework consists of two eight-layer CNNs that are well designed and powerful. To use CNNs to detect inshore ships, image samples are required, each of which should contain only one ship. However, the traditional image cropping method cannot generate such samples. To solve this problem, we present an orientation-free image cropping method that can generate trapezium rather than rectangle samples, making inshore ship detection by CNN feasible. Experimental results on Google Earth images demonstrate that the proposed method outperforms existing state-of-the-art methods. Xiaobin Li 0005, Shengjin Wang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Robust ImageGraph: Rank-Level Feature Fusion for Image SearchabstractRecently, feature fusion has demonstrated its effectiveness in image search. However, bad features and inappropriate parameters usually bring about false positive images, i.e., outliers, leading to inferior performance. Therefore, a major challenge of fusion scheme is how to be robust to outliers. Towards this goal, this paper proposes a rank-level framework for robust feature fusion. First, we define Rank Distance to measure the relevance of images at rank level. Based on it, Bayes similarity is introduced to evaluate the retrieval quality of individual features, through which true matches tend to obtain higher weight than outliers. Then, we construct the directed ImageGraph to encode the relationship of images. Each image is connected to its K nearest neighbors with an edge, and the edge is weighted by Bayes similarity. Multiple rank lists resulted from different methods are merged via ImageGraph. Furthermore, on the fused ImageGraph, local ranking is performed to re-order the initial rank lists. It aims at local optimization, and thus is more robust to global outliers. Extensive experiments on four benchmark data sets validate the effectiveness of our method. Besides, the proposed method outperforms two popular fusion schemes, and the results are competitive to the state-of-the-art. Ziqiong Liu, Shengjin Wang, Liang Zheng 0001, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Weakly Supervised Object Localization with Progressive Domain AdaptationabstractWe address the problem of weakly supervised object localization where only image-level annotations are available for training. Many existing approaches tackle this problem through object proposal mining. However, a substantial amount of noise in object proposals causes ambiguities for learning discriminative object models. Such approaches are sensitive to model initialization and often converge to an undesirable local minimum. In this paper, we address this problem by progressive domain adaptation with two main steps: classification adaptation and detection adaptation. In classification adaptation, we transfer a pre-trained network to our multi-label classification task for recognizing the presence of a certain object in an image. In detection adaptation, we first use a mask-out strategy to collect class-specific object proposals and apply multiple instance learning to mine confident candidates. We then use these selected object proposals to fine-tune all the layers, resulting in a fully adapted detection network. We extensively evaluate the localization performance on the PASCAL VOC and ILSVRC datasets and demonstrate significant performance improvement over the state-of-the-art methods. Dong Li 0025, Jia-Bin Huang 0001, Yali Li 0001, Shengjin Wang, Ming-Hsuan Yang 0001 |
CVPR | 4 |
| 2016 | Unsupervised Visual Representation Learning by Graph-Based Consistent Constraints
Dong Li 0025, Wei-Chih Hung, Jia-Bin Huang 0001, Shengjin Wang, Narendra Ahuja, Ming-Hsuan Yang 0001 |
ECCV (4) | 4 |
| 2016 | MARS: A Video Benchmark for Large-Scale Person Re-Identification
Liang Zheng 0001, Zhi Bie, Yifan Sun 0003, Jingdong Wang 0001, Chi Su, Shengjin Wang, Qi Tian 0001 |
ECCV (6) | 6 |
| 2016 | Bagging regularized common spatial pattern with hybrid motor imagery and myoelectric signalabstractCommon Spatial Pattern(CSP) is a widely used algorithm in BCI application. However, it is sensitive to noise and artifact. In this paper, we propose a bagging regularized common spatial pattern (Bagging RCSP) approach for BCI with hybrid motor imagery and myoelectric signal. We divide the training samples into packets and choose training packets by Bagging to extract RCSP features. Furthermore, LDA is used to project the feature vector to lower space. In the end, a classification algorithm based on NNC is adopted. The Off-line experiment on BCI competition III attests Bagging RCSP versatile. The accuracy increases by 3%-5% in average than RCSP-A results. Furthermore, we designed and realized an online BCI system based on Bagging RCSP and evaluated through experiment involving four experimenters performing the BCI system of catching the apples. The results show the effectiveness of the proposed approach and the real time BCI system. Hongchuan Liu, Yali Li 0001, Hongma Liu, Shengjin Wang |
ICASSP | 4 |
| 2016 | A Multi-stage Method for Chinese Text Detection in News VideosabstractWith the rapid increase of on-line video resources, there is an urgent demand for text detection and recognition technologies to build content-based video indexing and retrieval systems. Chinese news video texts contain highly condensed and rich information, but the low resolution of videos on the Internet and the complexity of Chinese character structures bring challenges for text detection. In this paper, we present a multi-stage scheme for Chinese news video text detection. We propose an improved Stroke Width Transform (SWT) method by incorporating text color consistency constraint for candidate text blocks generation. Then we use “divide and conquer” strategy to distinguish candidate text blocks into three sub-spaces according to their geometric shapes and size. For each sub-space, a neural network is designed to filter the candidates into text or non-text blocks. Finally, the text blocks are merged into text lines based on the stroke width, color and other heuristic information. Experimental results on self-collected Chinese news video dataset and ICDAR 2013 dataset show that the proposed method is effective to detect both news video captions and scene texts. Liangrui Peng, Shengjin Wang |
KES | 3 |
| 2016 | Accurate Image Search with Multi-Scale Contextual Evidences
Liang Zheng 0001, Shengjin Wang, Jingdong Wang 0001, Qi Tian 0001 |
Int. J. Comput. Vis. | 2 |
| 2016 | Fine-residual VLAD for image retrieval
Ziqiong Liu, Shengjin Wang, Qi Tian 0001 |
Neurocomputing | 2 |
| 2016 | Person Re-Identification by Discriminative Selection in Video RankingabstractCurrent person re-identification (ReID) methods typically rely on single-frame imagery features, whilst ignoring space-time information from image sequences often available in the practical surveillance scenarios. Single-frame (single-shot) based visual appearance matching is inherently limited for person ReID in public spaces due to the challenging visual ambiguity and uncertainty arising from non-overlapping camera views where viewing condition changes can cause significant people appearance variations. In this work, we present a novel model to automatically select the most discriminative video fragments from noisy/incomplete image sequences of people from which reliable space-time and appearance features can be computed, whilst simultaneously learning a video ranking function for person ReID. Using the PRID 2011, iLIDS-VID, and HDA+ image sequence datasets, we extensively conducted comparative evaluations to demonstrate the advantages of the proposed model over contemporary gait recognition, holistic image sequence matching and state-of-the-art single-/multi-shot ReID methods. Taiqing Wang, Shaogang Gong, Xiatian Zhu, Shengjin Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | A Boosting Approach to Exploit Instance Correlations for Multi-Instance ClassificationabstractWe propose a Boosting approach for multi-instance (MI) classification. Lp-norm is integrated to localize the witness instances and formulate the bag scores from classifier outputs. The contributions are twofold. First, a flexible and concise model for Boosting is proposed by the Lp-norm localization and exponential loss optimization. The scores for bag-level classification are directly fused from the instance feature space without probabilistic assumptions. Second, gradient and Newton descent optimizations are applied to derive the weak learners for Boosting. In particular, the instance correlations are exploited by fitting the weights and Newton updates for the weak learner construction. The final Boosted classifiers are the sums of iteratively chosen weak learners. Experiments demonstrate that the proposed Lp-norm-localized Boosting approach significantly improves the MI classification performance. Compared with the state of the art, the approach achieves the highest MI classification accuracy on 7/10 benchmark data sets. Yali Li 0001, Shengjin Wang, Qi Tian 0001, Xiaoqing Ding |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Query-adaptive late fusion for image search and person re-identificationabstractFeature fusion has been proven effective [35, 36] in image search. Typically, it is assumed that the to-be-fused heterogeneous features work well by themselves for the query. However, in a more realistic situation, one does not know in advance whether a feature is effective or not for a given query. As a result, it is of great importance to identify feature effectiveness in a query-adaptive manner. Liang Zheng 0001, Shengjin Wang, Ziqiong Liu, Qi Tian 0001 |
CVPR | 2 |
| 2015 | Fast Orthogonal Projection Based on Kronecker ProductabstractWe propose a family of structured matrices to speed up orthogonal projections for high-dimensional data commonly seen in computer vision applications. In this, a structured matrix is formed by the Kronecker product of a series of smaller orthogonal matrices. This achieves O(dlogd) computational complexity and O(logd) space complexity for d-dimensional data, a drastic improvement over the standard unstructured projections whose computational and space complexities are both O(d^2). The proposed structured matrices are applicable to a number of application domains, and are faster and more compact than other structured matrices used in the past. We also introduce an efficient learning procedure for optimizing such matrices in a data dependent fashion. We demonstrate the significant advantages of the proposed approach in solving the approximate nearest neighbor (ANN) image search problem with both binary embedding and quantization. We find that the orthogonality plays a very important role in solving ANN problem, since the random orthogonal Kronecker projection has already provided promising performance. Comprehensive experiments show that the proposed approach can achieve similar or better accuracy as the existing state-of-the-art but with significantly less time and memory. Xu Zhang 0022, Felix X. Yu, Sanjiv Kumar, Shengjin Wang, Shih-Fu Chang |
ICCV | 5 |
| 2015 | Scalable Person Re-identification: A BenchmarkabstractThis paper contributes a new high quality dataset for person re-identification, named "Market-1501". Generally, current datasets: 1) are limited in scale, 2) consist of hand-drawn bboxes, which are unavailable under realistic settings, 3) have only one ground truth and one query image for each identity (close environment). To tackle these problems, the proposed Market-1501 dataset is featured in three aspects. First, it contains over 32,000 annotated bboxes, plus a distractor set of over 500K images, making it the largest person re-id dataset to date. Second, images in Market-1501 dataset are produced using the Deformable Part Model (DPM) as pedestrian detector. Third, our dataset is collected in an open system, where each identity has multiple images under each camera. As a minor contribution, inspired by recent advances in large-scale image search, this paper proposes an unsupervised Bag-of-Words descriptor. We view person re-identification as a special task of image search. In experiment, we show that the proposed descriptor yields competitive accuracy on VIPeR, CUHK03, and Market-1501 datasets, and is scalable on the large-scale 500k dataset. Liang Zheng 0001, Liyue Shen, Shengjin Wang, Jingdong Wang 0001, Qi Tian 0001 |
ICCV | 4 |
| 2015 | Selective parts for fine-grained recognitionabstractClassical visual bag-of-words approaches tackle fine-grained recognition using global features which discard spatial location of features. In this paper, we propose a novel part-based approach to distinguish fine-grained categories. This work is distinguished by two contributions. First, a fully automatic technique for selecting mid-level parts from large amounts of candidate regions without any part supervised information is presented. We call the selected parts by discriminative mining algorithm as selective parts. Second, a general effective evaluation criterion of quantifying part discriminability is built, which leads to joint selection process. For classification, feature ensembles are constructed based on global object and selective parts. Experimental results demonstrate the particular effectiveness of selective parts for fine-grained recognition on bird species on the Caltech UCSD Birds (CUB) dataset. Dong Li 0025, Yali Li 0001, Shengjin Wang |
ICIP | 3 |
| 2015 | A survey of recent advances in visual feature detection
Yali Li 0001, Shengjin Wang, Qi Tian 0001, Xiaoqing Ding |
Neurocomputing | 2 |
| 2015 | Salient region detection via simple local and global contrast representation
Shengjin Wang |
Neurocomputing | 2 |
| 2015 | Cognitive pedestrian detector: Adapting detector to specific scene by transferring attributes
Xu Zhang 0022, Shengjin Wang |
Neurocomputing | 4 |
| 2015 | Tensor index for large scale image retrieval
Liang Zheng 0001, Shengjin Wang, Peizhen Guo, Hanyue Liang, Qi Tian 0001 |
Multim. Syst. | 2 |
| 2015 | Feature representation for statistical-learning-based object detection: A review
Yali Li 0001, Shengjin Wang, Qi Tian 0001, Xiaoqing Ding |
Pattern Recognit. | 2 |
| 2015 | Beyond χ2 Difference: Learning Optimal Metric for Boundary DetectionabstractThis letter focuses on solving the challenging problem of detecting natural image boundaries. A boundary usually refers to the border between two regions with different semantic meanings. Therefore, a measurement of dissimilarity between image regions plays a pivotal role in boundary detection of natural images. To improve the performance of boundary detection, a Learning-based Boundary Metric (LBM) is proposed to replace χ2difference adopted by the classical algorithm mPb. Compared with χ2difference, LBM is composed of a single layer neural network and an RBF kernel, and is fine-tuned by supervised learning rather than human-crafted. It is more effective in describing the dissimilarity between natural image regions while tolerating large variance of image data. After substituting χ2difference with LBM, the F-measure metric of mPb on the BSDS500 benchmark is increased from 0.69 to 0.71. Moreover, when image features are computed on a single scale, the proposed LBM algorithm still achieves competitive results compared with mPb, which makes use of multi-scale image features. Shengjin Wang |
IEEE Signal Process. Lett. | 2 |
| 2015 | Fast Image Retrieval: Query Pruning and Early TerminationabstractEfficiency is of great importance for image retrieval systems. For this pragmatic issue, this paper proposes a fast image retrieval framework to speed up the online retrieval process. To this end, an impact score for local features is proposed in the first place, which considers multiple properties of a local feature, including TF-IDF, scale, saliency, and ambiguity. Then, to decrease memory consumption, the impact score is quantized to an integer, which leads to a novel inverted index organization, called Q-Index. Importantly, based on the impact score, two closely complementary strategies are introduced: query pruning and early termination. On one hand, query pruning discards less important features in the query. On the other hand, early termination visits indexed features only with high impact scores, resulting in the partial traversing of the inverted index. Our approach is tested on two benchmark datasets populated with an additional 1 million images to account as negative examples. Compared with full traversal of the inverted index, we show that our system is capable of visiting less than 10% of the “should-visit” postings, thus achieving a significant speed-up in query time while providing competitive retrieval accuracy. Liang Zheng 0001, Shengjin Wang, Ziqiong Liu, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2014 | Packing and Padding: Coupled Multi-index for Accurate Image RetrievalabstractIn Bag-of-Words (BoW) based image retrieval, the SIFT visual word has a low discriminative power, so false positive matches occur prevalently. Apart from the information loss during quantization, another cause is that the SIFT feature only describes the local gradient distribution. To address this problem, this paper proposes a coupled Multi-Index (c-MI) framework to perform feature fusion at indexing level. Basically, complementary features are coupled into a multi-dimensional inverted index. Each dimension of c-MI corresponds to one kind of feature, and the retrieval process votes for images similar in both SIFT and other feature spaces. Specifically, we exploit the fusion of local color feature into c-MI. While the precision of visual match is greatly enhanced, we adopt Multiple Assignment to improve recall. The joint cooperation of SIFT and color features significantly reduces the impact of false positive matches. Extensive experiments on several benchmark datasets demonstrate that c-MI improves the retrieval accuracy significantly, while consuming only half of the query time compared to the baseline. Importantly, we show that c-MI is well complementary to many prior techniques. Assembling these methods, we have obtained an mAP of 85.8% and N-S score of 3.85 on Holidays and Ukbench datasets, respectively, which compare favorably with the state-of-the-arts. Liang Zheng 0001, Shengjin Wang, Ziqiong Liu, Qi Tian 0001 |
CVPR | 2 |
| 2014 | Bayes Merging of Multiple Vocabularies for Scalable Image RetrievalabstractIn the Bag-of-Words (BoW) model, the vocabulary is of key importance. Typically, multiple vocabularies are generated to correct quantization artifacts and improve recall. However, this routine is corrupted by vocabulary correlation, i.e., overlapping among different vocabularies. Vocabulary correlation leads to an over-counting of the indexed features in the overlapped area, or the intersection set, thus compromising the retrieval accuracy. In order to address the correlation problem while preserve the benefit of high recall, this paper proposes a Bayes merging approach to down-weight the indexed features in the intersection set. Through explicitly modeling the correlation problem in a probabilistic view, a joint similarity on both image- and feature-level is estimated for the indexed features in the intersection set. We evaluate our method on three benchmark datasets. Albeit simple, Bayes merging can be well applied in various merging tasks, and consistently improves the baselines on multi-vocabulary merging. Moreover, Bayes merging is efficient in terms of both time and memory cost, and yields competitive performance with the state-of-the-art methods. Liang Zheng 0001, Shengjin Wang, Wengang Zhou 0001, Qi Tian 0001 |
CVPR | 2 |
| 2014 | Person Re-identification by Video Ranking
Taiqing Wang, Shaogang Gong, Xiatian Zhu, Shengjin Wang |
ECCV (4) | 4 |
| 2014 | Visual reranking with improved image graphabstractThis paper introduces an improved reranking method for the Bag-of-Words (BoW) based image search. Built on [1], a directed image graph robust to outlier distraction is proposed. In our approach, the relevance among images is encoded in the image graph, based on which the initial rank list is refined. Moreover, we show that the rank-level feature fusion can be adopted in this reranking method as well. Taking advantage of the complementary nature of various features, the reranking performance is further enhanced. Particularly, we exploit the reranking method combining the BoW and color information. Experiments on two benchmark datasets demonstrate that our method yields significant improvements and the reranking results are competitive to the state-of-the-art methods. Ziqiong Liu, Shengjin Wang, Liang Zheng 0001, Qi Tian 0001 |
ICASSP | 2 |
| 2014 | Submanifold DecompositionabstractLow-dimensional structures embedded in high-dimensional data space can be extracted by spectral analysis and manifold learning. Standard approaches to manifold learning are usually based on the assumption that there is a dominant low-dimensional manifold, while other variations are considered with minor priority. We instead consider the scenario that a pair of distinct manifolds intertwined in the same high-dimensional space, which can be decomposed for analysis. The core of this new method is a novel submanifold decomposition (SMD) algorithm. This paper has three contributions: 1) a submanifold framework is proposed to model the high-dimensional dataset, which is dominated by more than one factor; 2) a nonlinear manifold decomposition method, SMD, is presented to extract two intertwined manifolds from a dataset in a discriminative manner; and 3) in order to solve the out-of-sample problem of nonlinear SMD, a linear extension of SMD is developed, which is effective to extract two linear submanifolds. We demonstrate that comparing with the existing manifold learning methods that only extract one dominant manifold, the proposed SMD and its linear extension are capable of extracting a pair of submanifolds discriminatively and effectively. Moreover, the two extracted manifolds can complement each other to enhance the representation performance. Extensive experiments on both artificial data and real data demonstrate that the proposed method outperforms the state-of-the-art manifold learning algorithms in visual recognition tasks. Ya Su, Sheng Li 0001, Shengjin Wang, Yun Fu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Detecting Human Action as the Spatio-Temporal Tube of Maximum Mutual InformationabstractHuman action detection in complex scenes is a challenging problem due to its high-dimensional search space and dynamic backgrounds. To achieve efficient and accurate action detection, we represent a video sequence as a collection of feature trajectories and model human action as the spatio-temporal tube (ST-tube) of maximum mutual information. First, a random forest is built to evaluate the mutual information of feature trajectories toward the action class, and then a one-order Markov model is introduced to recursively infer the action regions at consecutive frames. By exploring the time-continuity property of feature trajectories, the action region is efficiently inferred at large temporal intervals. Finally, we obtain an ST-tube by concatenating the consecutive action regions bounding the human bodies. Compared with the popular spatio-temporal cuboid action model, the proposed ST-tube model is not only more efficient, but also more accurate in action localization. Experimental results on the KTH, CMU and UCF sports datasets validate the superiority of our approach over the state-of-the-art methods in both localization accuracy and time efficiency. Taiqing Wang, Shengjin Wang, Xiaoqing Ding |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Learning Cascaded Shared-Boost Classifiers for Part-Based Object DetectionabstractThis paper focuses on the problem of detecting a number of different class objects in images. We present a novel part-based model for object detection with cascaded classifiers. The coarse root and fine part classifiers are combined into the model. Different from the existing methods which learn root and part classifiers independently, we propose a shared-Boost algorithm to jointly train multiple classifiers. This paper is distinguished by two key contributions. The first is to introduce a new definition of shared features for similar pattern representation among multiple classifiers. Based on this, a shared-Boost algorithm which jointly learns multiple classifiers by reusing the shared feature information is proposed. The second contribution is a method for constructing a discriminatively trained part-based model, which fuses the outputs of cascaded shared-Boost classifiers as high-level features. The proposed shared-Boost-based part model is applied for both rigid and deformable object detection experiments. Compared with the state-of-the-art method, the proposed model can achieve higher or comparable performance. In particular, it can lift up the detection rates in low-resolution images. Also the proposed procedure provides a systematic framework for information reusing among multiple classifiers for part-based object detection. Yali Li 0001, Shengjin Wang, Qi Tian 0001, Xiaoqing Ding |
IEEE Trans. Image Process. | 2 |
| 2014 | Coupled Binary Embedding for Large-Scale Image RetrievalabstractVisual matching is a crucial step in image retrieval based on the bag-of-words (BoW) model. In the baseline method, two keypoints are considered as a matching pair if their SIFT descriptors are quantized to the same visual word. However, the SIFT visual word has two limitations. First, it loses most of its discriminative power during quantization. Second, SIFT only describes the local texture feature. Both drawbacks impair the discriminative power of the BoW model and lead to false positive matches. To tackle this problem, this paper proposes to embed multiple binary features at indexing level. To model correlation between features, a multi-IDF scheme is introduced, through which different binary features are coupled into the inverted file. We show that matching verification methods based on binary features, such as Hamming embedding, can be effectively incorporated in our framework. As an extension, we explore the fusion of binary color feature into image retrieval. The joint integration of the SIFT visual word and binary features greatly enhances the precision of visual matching, reducing the impact of false positive matches. Our method is evaluated through extensive experiments on four benchmark datasets (Ukbench, Holidays, DupImage, and MIR Flickr 1M). We show that our method significantly improves the baseline approach. In addition, large-scale experiments indicate that the proposed method requires acceptable memory usage and query time compared with other approaches. Further, when global color feature is integrated, our method yields competitive performance with the state-of-the-arts. Liang Zheng 0001, Shengjin Wang, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | \(\mathcal {L}_p\) -Norm IDF for Scalable Image RetrievalabstractThe inverse document frequency (IDF) is prevalently utilized in the bag-of-words-based image retrieval application. The basic idea is to assign less weight to terms with high frequency, and vice versa. However, in the conventional IDF routine, the estimation of visual word frequency is coarse and heuristic. Therefore, its effectiveness is largely compromised and far from optimal. To address this problem, this paper introduces a novel IDF family by the use of Lp-norm pooling technique. Carefully designed, the proposed IDF considers the term frequency, document frequency, the complexity of images, as well as the codebook information. We further propose a parameter tuning strategy, which helps to produce optimal balancing between TF and pIDF weights, yielding the so-called Lp-norm IDF (pIDF). We show that the conventional IDF is a special case of our generalized version, and two novel IDFs, i.e., the average IDF and the max IDF, can be defined from the concept of pIDF. Further, by counting for the term-frequency in each image, the proposed pIDF helps to alleviate the visual word burstiness phenomenon. Our method is evaluated through extensive experiments on four benchmark data sets (Oxford 5K, Paris 6K, Holidays, and Ukbench). We show that the pIDF works well on large scale databases and when the codebook is trained on irrelevant data. We report an mean average precision improvement of as large as +13.0% over the baseline TF-IDF approach on a 1M data set. In addition, the pIDF has a wide application scope varying from buildings to general objects and scenes. When combined with postprocessing steps, we achieve competitive results compared with the state-of-the-art methods. In addition, since the pIDF is computed offline, no extra computation or memory cost is introduced to the system at all. Liang Zheng 0001, Shengjin Wang, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Lp-Norm IDF for Large Scale Image SearchabstractThe Inverse Document Frequency (IDF) is prevalently utilized in the Bag-of-Words based image search. The basic idea is to assign less weight to terms with high frequency, and vice versa. However, the estimation of visual word frequency is coarse and heuristic. Therefore, the effectiveness of the conventional IDF routine is marginal, and far from optimal. To tackle this problem, this paper introduces a novel IDF expression by the use of Lp-norm pooling technique. Carefully designed, the proposed IDF takes into account the term frequency, document frequency, the complexity of images, as well as the codebook information. Optimizing the IDF function towards optimal balancing between TF and pIDF weights yields the so-called Lp-norm IDF (pIDF). We show that the conventional IDF is a special case of our generalized version, and two novel IDFs, i.e. the average IDF and the max IDF, can also be derived from our formula. Further, by counting for the term-frequency in each image, the proposed Lp-norm IDF helps to alleviate the visual word burstiness phenomenon. Our method is evaluated through extensive experiments on three benchmark datasets (Oxford 5K, Paris 6K and Flickr 1M). We report a performance improvement of as large as 27.1% over the baseline approach. Moreover, since the Lp-norm IDF is computed offline, no extra computation or memory cost is introduced to the system at all. Liang Zheng 0001, Shengjin Wang, Ziqiong Liu, Qi Tian 0001 |
CVPR | 2 |
| 2013 | Visual Phraselet: Refining Spatial Constraints for Large Scale Image SearchabstractAbstract—The Bag-of-Words (BoW) model is prone to the deficiency of spatial constraints among visual words. The state of the art methods encode spatial information via visual phrases. However, these methods discard the spatial context among visual phrases instead. To address the problem, this letter introduces a novel visual concept, the Visual Phraselet, as a kind of similarity measurement between images. The visual phraselet refers to the spatial consistent group of visual phrases. In a simple yet effective manner, visual phraselet filters out false visual phrase matches, and is much more discriminative than both visual word and visual phrase. To boost the discovery of visual phraselets, we apply the soft quantization scheme. Our method is evaluated through extensive experiments on three benchmark datasets (Oxford 5 K, Paris 6 K and Flickr 1 M). We report significant improvements as large as 54.6 % over the baseline approach, thus validating the concept of visual phraselet. Index Terms—Image search, spatial constraint, visual phrase, vi-sual phraselet. I. Liang Zheng 0001, Shengjin Wang |
IEEE Signal Process. Lett. | 2 |
| 2012 | Face Swapping under Large Pose Variations: A 3D Model Based ApproachabstractTraditional face swapping technologies require the faces of source images and target images have similar pose and appearance (usually frontal). This limits its applications. This paper presents a method for face swapping based on personalized 3D head models. This framework builds a personalized 3D head model from a frontal face and can be rendered at any pose to match the characters in the image we want to swap. The 3D head model is constructed by a user uploaded frontal view face image. This construction process goes through face alignment and feature point matching. The final personalized 3D head is built by deforming a standard 3D head model using radial basis function to match the specific person. To make the synthesized face seamlessly blended into the image, color transfer and multi-resolution spline technique are applied. We use the proposed technique to create personalized storybook where the characters are replaced with a user's face and promising results are obtained. The system can be used in face de-identification as well. Shengjin Wang, Qian Lin 0001 |
ICME | 2 |
| 2012 | Evaluation of canonical correlation analysis: A Correlation Generation Model
Ya Su, Shengjin Wang, Yun Fu 0001 |
ICPR | 2 |
| 2012 | Submanifold decomposition
Ya Su, Shengjin Wang, Yun Fu 0001 |
ICPR | 2 |
| 2012 | Face replacement with large-pose differencesabstractIn this paper, we present a novel face replacement system exchanging faces with large-pose differences. Traditional 2D image based face replacement can only replace faces with similar pose and appearance. This significantly limits the application of face replacement. In this paper, we propose to build a 3D head model from a single frontal face photo. The automatically constructed 3D head can be rendered under arbitrary poses and illuminations. This makes it possible to do swapping for faces with large pose variations. In the demo, the user captures a frontal face image using a capture device such as a webcam or a smartphone, and then the algorithm can automatically build the 3D model using feature detection, face alignment and reconstruction. This 3D model is used to swap to any other target face photo the user selects. While our system is automatic, we also provide interactive tools for the user to adjust the feature detection to enhance the results. Qian Lin 0001, Shengjin Wang |
ACM Multimedia | 4 |
| 2011 | Contextual image searchabstractIn this paper, we propose a novel image search scheme, contextual image search. Different from conventional image search schemes that present a separate interface (e.g., text input box) to allow users to submit a query, the new search scheme enables users to search images by only masking a few words when they are reading through Web pages or other documents. Rather than merely making use of the explicit query input that is often not sufficient to express user's search intent, our approach explores the context information to better understand the search intent with two key steps: query augmenting and search results reranking using context, and expects to obtain better search results. Beyond contextual Web search, the context in our case is much richer and includes images besides texts. In addition to this type of search scheme, called contextual image search with text input, we also present another type of scheme, called contextual image search with image input, to allow users to select an image as the search query from Web pages or other documents they are reading. The key idea is to use the search-to-annotation technique and the contextual textual query mining scheme to determine the corresponding textual query, to finally get semantically similar search results. Experiments show that the proposed schemes make image search more convenient and the search results are more relevant to user intention. Wenhao Lu, Jingdong Wang 0001, Xian-Sheng Hua 0001, Shengjin Wang, Shipeng Li 0001 |
ACM Multimedia | 4 |
| 2010 | Person-independent head pose estimation based on random forest regressionabstractIn this paper, a novel approach for person-independent head pose estimation in gray-level images is presented. There are two steps of the proposed method. In order to preserve similar patterns of faces under various poses, a novel multi-view face detector using tree-structured cascaded-Adaboost classifiers is applied. Furthermore, based on the cropped face images, randomized regression trees are learned and applied to estimate head pose precisely. Experiments show that our method achieves better pose estimation results in both horizontal and vertical orientations in comparison with the reported result with skin color information. Yali Li 0001, Shengjin Wang, Xiaoqing Ding |
ICIP | 2 |
| 2010 | Part Detection, Description and Selection Based on Hidden Conditional Random FieldsabstractIn this paper, the problem of part detection, description and selection is discussed. This problem is crucial in the learning algorithms of part-based models, but can't be solved well when some candidate parts are extracted from background. This paper studies this problem and introduces a new algorithm, HCRF-PS (Hidden Conditional Random Fields for Part Selection), for part detection, description, especially selection. Our algorithm is distinguished for its power to optimize multiple kinds of information at the same time, including texture, color, location and part label. Finally, we did some experiments with HCRF-PS algorithm which give good results on both virtual and real data. Wenhao Lu, Shengjin Wang, Xiaoqing Ding |
ICPR | 2 |
| 2010 | Eye/eyes tracking based on a unified deformable template and particle filtering
Yali Li 0001, Shengjin Wang, Xiaoqing Ding |
Pattern Recognit. Lett. | 2 |
| 2010 | Action and Gait Recognition From Recovered 3-D Human JointsabstractA common viewpoint-free framework that fuses pose recovery and classification for action and gait recognition is presented in this paper. First, a markerless pose recovery method is adopted to automatically capture the 3-D human joint and pose parameter sequences from volume data. Second, multiple configuration features (combination of joints) and movement features (position, orientation, and height of the body) are extracted from the recovered 3-D human joint and pose parameter sequences. A hidden Markov model (HMM) and an exemplar-based HMM are then used to model the movement features and configuration features, respectively. Finally, actions are classified by a hierarchical classifier that fuses the movement features and the configuration features, and persons are recognized from their gait sequences with the configuration features. The effectiveness of the proposed approach is demonstrated with experiments on the Institut National de Recherche en Informatique et Automatique Xmas Motion Acquisition Sequences data set. Junxia Gu, Xiaoqing Ding, Shengjin Wang, Youshou Wu |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2009 | Vehicle Detection and Tracking in Relatively Crowded ConditionsabstractAiming at vehicle detection and tracking problems in video monitoring and controlling system, this paper mainly studies vehicle detection and tracking problems in conditions of high traffic density in daytime. This paper is distinguished by two key contributions. First, we develop an improvement — SEAP(Simple but Efficient After Process) which checks the detection results in an accurate way and is an after process of Adaboost [1] detector which used to detect car in every frame. Second, we propose a tracking algorithm named 4-states tracking algorithm based on Kalman[5] linear filter. Tracking results turn unsteady as traffic density grows higher because of much more false positives and false negatives appear. However, 4-states tracking algorithm can solve this problem in an easy way by introducing FSM (Finite State Machine) into tracking algorithm. Finally, we implement a real-time vehicle detection and tracking system with the upper methods. Experiments give good results in relative crowded Conditions. Wenhao Lu, Shengjin Wang, Xiaoqing Ding |
SMC | 2 |
| 2008 | Adaptive particle filter with body part segmentation for full body trackingabstractThis paper presents a novel approach for marker-less 3D full body pose tracking using adaptive particle filter. Firstly, the search space decomposition strategy and body part segmentation method are used to reduce the calculation complexity due to the large degrees of freedom. Then an adaptive particle filter is adopted to track each body part. This new technique is a significant improvement over the standard particle filter with the advantage of adaptive particle number for each body part. Experimental results on tracking several challenging action sequences have shown that the proposed 3D full body tracker is able to effectively handle rapid no-linear movements, large changes of viewpoint, and different actors. The average errors of joint position are from 0.56 to 1.13 voxel in these action sequences. Junxia Gu, Xiaoqing Ding, Shengjin Wang, Youshou Wu |
FG | 3 |
| 2008 | Full body tracking-based human action recognitionabstractIn this paper, we present a novel method for human action recognition with the combined global movement feature and local configuration feature. The human action is represented as a sequence of joints in the 4D spatio-temporal space, and modeled by two HMMs, a conventional HMM for global movement feature and an exemplar-based HMM for configuration feature. Firstly, an adaptive particle filter is adopted to track the marker-less actor's 3D joints. Then, the combined features are extracted from the full body tracking results. Finally, the actions are classified by fusing two HMMs. The effectiveness of the proposed algorithm is demonstrated with experiments on 7 actions by 12 actors. The results prove robustness of the proposed method with respect to viewpoints and actors. Junxia Gu, Xiaoqing Ding, Shengjin Wang, Youshou Wu |
ICPR | 3 |
| 2008 | Visual Tracker Using Sequential Bayesian Learning: Discriminative, Generative, and HybridabstractThis paper presents a novel solution to track a visual object under changes in illumination, viewpoint, pose, scale, and occlusion. Under the framework of sequential Bayesian learning, we first develop a discriminative model-based tracker with a fast relevance vector machine algorithm, and then, a generative model-based tracker with a novel sequential Gaussian mixture model algorithm. Finally, we present a three-level hierarchy to investigate different schemes to combine the discriminative and generative models for tracking. The presented hierarchical model combination contains the learner combination (at level one), classifier combination (at level two), and decision combination (at level three). The experimental results with quantitative comparisons performed on many realistic video sequences show that the proposed adaptive combination of discriminative and generative models achieves the best overall performance. Qualitative comparison with some state-of-the-art methods demonstrates the effectiveness and efficiency of our method in handling various challenges during tracking. Yun Lei, Xiaoqing Ding, Shengjin Wang |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2007 | Multi-view Moving Human Detection and Correspondence Based on Object Occupancy Random Field
Xiaoqing Ding, Shengjin Wang, Youshou Wu |
ISNN (3) | 3 |
| 2006 | Object Detection Via Fusion of Global Classifier and Part-Based Classifier
Shengjin Wang, Xiaoqing Ding |
ISNN (2) | 2 |
| 2004 | Recognition of 3-D objects in multiple statuses based on Markov random field modelsabstractA general framework is presented to realize 3D object recognition, invariant to object scaling, deformation, rotation, occlusion, and viewpoint change. This framework utilizes densely sampled grids, with different resolutions, to represent the local information of the input image. A Markov random field (MRF) model is then created to model the geometric distribution of the object key nodes. Flexible matching, which is aimed at finding the accurate correspondence map between the key points of two images, is performed by combining the local similarities and the geometric relations together using the highest confidence first (HCF) method. Afterwards, a global similarity is calculated for object recognition. Experimental results on the Coil-100 object database are presented. The excellent recognition rates achieved in all the experiments indicate that our approach is well-suited for appearance-based recognition. Xiaoqing Ding, Shengjin Wang |
ICASSP (5) | 3 |