Xi Li 0001

dblp:46/2311-1 · DBLP profile ↗
← Back
237ranked-venue papers
25as first author
108since 2021 · last 2026
0000-0003-3023-1662ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 154 · 16 first-author · 66 since 2021Artificial intelligence and machine learning · 136 · 16 first-author · 64 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-authorComputer networks · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Text-driven 3D human motion generation for pose estimation using dual-transformer architecture
Rizwan Abbas, Hua Gao, Xi Li 0001
Comput. Aided Des.3
2026 MovieChat+: Question-Aware Sparse Memory for Long Video Question Answering
abstract
Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific vision tasks. Yet, existing methods either employ complex spatial-temporal modules or rely heavily on additional perception models to extract temporal features for video understanding, performing well only on short videos. For long videos, the computational complexity and memory costs associated with long-term temporal connections are significantly increased, posing additional challenges. Leveraging the hierarchical memory structure of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination, we propose MovieChat within a training-free memory consolidation mechanism to overcome these challenges, which transfers dense frames from short-term memory into sparse tokens in long-term memory by temporally merging adjacent frames. We lift pre-trained large multi-modal models for understanding long videos without additional trainable modules, employing a zero-shot approach. Additionally, in our new version, MovieChat+, we design an enhanced training-free vision-question matching-based memory consolidation mechanism to better anchor predictions to relevant visual content. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1 K benchmark with 1 K long video, 2 K temporal grounding labels, and 14 K manual annotations.
Enxin Song, Wenhao Chai, Tian Ye 0001, Jenq-Neng Hwang, Xi Li 0001, Gaoang Wang
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 CamI2V-Epipolar: Epipolar-Constrained Block Sparse Attention for Camera-Controlled Image-to-Video Diffusion Model
abstract
Recently, camera pose has emerged as a physics-informed condition for video diffusion models. Existing methods that directly adopt 3D full attention for cross-frame feature interaction achieve moderate camera controllability with extensive training, yet camera controllability and geometric consistency remain a challenge in complex scenarios such as large camera rotation or movement. This underscores the value of integrating physical priors into attention mechanisms. Instead of pixel-level attention masking, we propose a key innovation to infuse epipolar geometric priors into the block selection of video block sparse attention. This design resolves memory and speed bottlenecks at high resolutions like 768P, enables compatibility with FlashAttention-3, and establishes a principled framework for embedding physical priors into sparse attention mechanisms. Experiments on static RealEstate10 K and dynamic RealCam-Vid datasets demonstrate that our method outperforms the state-of-the-art RealCam-I2V, effectively enhancing camera controllability while preserving generation quality and generalization.
Guangcong Zheng, Xi Li 0001
IEEE Signal Process. Lett.2
2026 CWPS: Efficient Channel-Wise Parameter Sharing for Knowledge Transfer
abstract
Knowledge transfer aims to apply existing knowledge to different tasks or new data, and it has extensive applications in multi-domain and Multi-Task Learning. The key to this task is quickly identifying a fine-grained object for knowledge sharing and efficiently transferring knowledge. Current methods, such as fine-tuning, layer-wise parameter sharing, and task-specific adapters, only offer coarse-grained sharing solutions and struggle to effectively search for shared parameters, thus hindering the performance and efficiency of knowledge transfer. To address these issues, we propose Channel-Wise Parameter Sharing (CWPS), a novel fine-grained parameter-sharing method for knowledge transfer, which is efficient for parameter sharing, comprehensive, and plug-and-play. For the coarse-grained problem, we first achieve fine-grained parameter sharing by refining the granularity of shared parameters from the level of layers to the level of neurons. The knowledge learned from previous tasks can be utilized through the explicit composition of the model neurons. Besides, we promote an effective search strategy to minimize computational costs, simplifying the selection of shared weights. In addition, our CWPS has strong composability and generalization ability, which theoretically can be applied to any network consisting of linear and convolution layers. We introduce several datasets in both Incremental Learning and Multi-Task Learning scenarios. Our method has achieved state-of-the-art precision-to-parameter ratio performance with various backbones, demonstrating its efficiency and versatility.
Mingxuan Cui, Xuewei Li 0003, Cunzheng Wang, Gaoang Wang, Chenyi Zhuang, Jinjie Gu, Xiubo Liang, Xi Li 0001
IEEE Trans. Image Process.9
2026 Exploring Vision-Based Active 3D Object Detection by Informativeness Characterization
abstract
Vision-based 3D object detection (3DOD) gains lots of attention due to its low cost for deployment compared to Lidar-based tasks, while it suffers from labor-expensive data annotations. At the same time, active learning (AL) has shown great potential in reducing annotation costs in related tasks, which can maximize model performance within very limited labeled data. In this paper, we explore active learning for vision-based 3DOD for the first time. Inspired by the entropy analysis, we involve three concerns to characterize the sample informativeness: sample diversity in input space, feature informativeness in BEV space, and result distribution in prediction space. Based on these concerns, we propose a novel AL framework named HMAD, which utilizes Height Modeling and Adaptive Diversity-based sampling for comprehensive informativeness characterization. In HMAD, we first propose a novel height-guided adversarial module in BEV space, which measures the informativeness of height modeling for 2D-to-3D mapping in an adversarial manner. Furthermore, Budget-aware SpatioTemporal diversity Sampling (BSTS) and Class Balance Sampling (CBS) are proposed to adaptively measure the sample informativeness in input and prediction space, respectively. Finally, the three components are integrated into a two-stage sampling strategy, with which the most informative samples can be selected and annotated for the next iteration. Experiments evidence that HMAD achieves comparable performances by only using 50% annotated training data, and can generalize well on different conditions.
Yiming Wu 0006, Yehao Lu, Xuewei Li 0003, Xiubo Liang, Xi Li 0001
IEEE Trans. Image Process.7
2025 Exploring Unbiased Deepfake Detection via Token-Level Shuffling and Mixing
abstract
The generalization problem is broadly recognized as a critical challenge in detecting deepfakes. Most previous work believes that the generalization gap is caused by the differences among various forgery methods. However, our investigation reveals that the generalization issue can still occur when forgery-irrelevant factors shift. In this work, we identify two biases that detectors may also be prone to overfitting: position bias and content bias, as depicted in Fig. 1. For the position bias, we observe that detectors are prone to “lazily” depending on the specific positions within an image (e.g., central regions even no forgery). As for content bias, we argue that detectors may potentially and mistakenly utilize forgery-unrelated information for detection (e.g., background, and hair). To intervene in these biases, we propose two branches for shuffling and mixing with tokens in the latent space of transformers. For the shuffling branch, we rearrange the tokens and corresponding position embedding for each image while maintaining the local correlation. For the mixing branch, we randomly select and mix the tokens in the latent space between two images with the same label within the mini-batch to recombine the content information. During the learning process, we align the outputs of detectors from different branches in both feature space and logit space. Contrastive losses for features and divergence losses for logits are applied to obtain unbiased feature representation and classifiers. We demonstrate and verify the effectiveness of our method through extensive experiments on widely used evaluation datasets.
Xinghe Fu, Zhiyuan Yan 0002, Taiping Yao, Shen Chen 0004, Xi Li 0001
AAAI5
2025 CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
abstract
Customized video generation aims to generate high-quality videos guided by text prompts and subject's reference images. However, since it is only trained on static images, the fine-tuning process of subject learning disrupts abilities of video diffusion models (VDMs) to combine concepts and generate motions. To restore these abilities, some methods use additional video similar to the prompt to fine-tune or guide the model. This requires frequent changes of guiding videos and even re-tuning of the model when generating different motions, which is very inconvenient for users. In this paper, we propose CustomCrafter, a novel framework that preserves the model's motion generation and conceptual combination abilities without additional video and fine-tuning to recovery. For preserving conceptual combination ability, we design a plug-and-play module to update few parameters in VDMs, enhancing the model's ability to capture the appearance details and the ability of concept combinations for new subjects. For motion generation, we observed that VDMs tend to restore the motion of video in the early stage of denoising, while focusing on the recovery of subject details in the later stage. Therefore, we propose Dynamic Weighted Video Sampling Strategy. Using the pluggability of our subject learning modules, we reduce the impact of this module on motion generation in the early stage of denoising, preserving the ability to generate motion of VDMs. In the later stage of denoising, we restore this module to repair the appearance details of the specified subject, thereby ensuring the fidelity of the subject's appearance. Experimental results show that our method has a significant improvement compared to previous methods.
Yong Zhang 0034, Xintao Wang 0002, Xianpan Zhou, Guangcong Zheng, Zhongang Qi, Ying Shan, Xi Li 0001
AAAI8
2025 PiD: Generalized AI-Generated Images Detection with Pixelwise Decomposition Residuals
abstract
Fake images, created by recently advanced generative models, have become increasingly indistinguishable from real ones, making their detection crucial, urgent, and challenging. This paper introduces PiD (Pixelwise Decomposition Residuals), a novel detection method that focuses on residual signals within images. Generative models are designed to optimize high-level semantic content (principal components), often overlooking low-level signals (residual components). PiD leverages this observation by disentangling residual components from images, encouraging the model to uncover more underlying and general forgery clues independent of semantic content. Compared to prior approaches that rely on reconstruction techniques or high-frequency information, PiD is computationally efficient and does not rely on any generative models for reconstruction. Specifically, PiD operates at the pixel level, mapping the pixel vector to another color space (e.g., YUV) and then quantizing the vector. The pixel vector is mapped back to the RGB space and the quantization loss is taken as the residual for AIGC detection. Our experiment results are striking and highly surprising: PiD achieves 98% accuracy on the widely used GenImage benchmark, highlighting the effectiveness and generalization performance.
Xinghe Fu, Zhiyuan Yan 0002, Taiping Yao, Yandan Zhao, Shouhong Ding, Xi Li 0001
ICML7
2025 SDSP: Scalable and Diverse Synthetic Pairwise Text Generation from Web Corpus Using Large Language Model
Xiaoxu Wu, Xi Li 0001, Aleksei Timofeev, Yinfei Yang, Si Liu 0001, Jiulong Shan
ICONIP (1)2
2025 Collaborating Vision, Depth, and Thermal Signals for Multi-Modal Tracking: Dataset and Algorithm
abstract
Existing multi-modal object tracking approaches primarily focus on dual-modal paradigms, such as RGB-Depth or RGB-Thermal, yet remain challenged in complex scenarios due to limited input modalities. To address this gap, this work introduces a novel multi-modal tracking task that leverages three complementary modalities, including visible RGB, Depth (D), and Thermal Infrared (TIR), aiming to enhance robustness in complex scenarios. To support this task, we construct a new multi-modal tracking dataset, coined RGBDT500, which consists of 500 videos with synchronised frames across the three modalities. Each frame provides spatially aligned RGB, depth, and thermal infrared images with precise object bounding box annotations.Furthermore, we propose a novel multi-modal tracker, dubbed RDTTrack.RDTTrack integrates tri-modal information for robust tracking by leveraging a pretrained RGB-only tracking model and prompt learning techniques.In specific, RDTTrack fuses thermal infrared and depth modalities under a proposed orthogonal projection constraint, then integrates them with RGB signals as prompts for the pre-trained foundation tracking model, effectively harmonising tri-modal complementary cues.The experimental results demonstrate the effectiveness and advantages of the proposed method, showing significant improvements over existing dual-modal approaches in terms of tracking accuracy and robustness in complex scenarios. The dataset and source code are publicly available at https://xuefeng-zhu5.github.io/RGBDT500.
Xuefeng Zhu 0003, Tianyang Xu 0001, Yifan Pan, Jinjie Gu, Xi Li 0001, Jiwen Lu, Xiaojun Wu 0001, Josef Kittler
NeurIPS5
2025 FusionBooster: A Unified Image Fusion Boosting Paradigm
Chunyang Cheng, Tianyang Xu 0001, Xiaojun Wu 0001, Hui Li 0037, Xi Li 0001, Josef Kittler
Int. J. Comput. Vis.5
2025 OCCO: LVM-Guided Infrared and Visible Image Fusion Framework Based on Object-Aware and Contextual Contrastive Learning
Hui Li 0037, Congcong Bian, Zeyang Zhang 0002, Xiaoning Song, Xi Li 0001, Xiaojun Wu 0001
Int. J. Comput. Vis.5
2025 Context-based emotion recognition: A survey
Rizwan Abbas, Bingnan Ni, Ruhui Ma, Yehao Lu, Xi Li 0001
Neurocomputing6
2025 Enhancing Tiny Object Detection Using Guided Object Inference Slicing (GOIS): An efficient dynamic adaptive framework for fine-tuned and non-fine-tuned deep learning models
Muhammad Muzammul, Xuewei Li 0003, Xi Li 0001
Neurocomputing3
2025 Emotion recognition in live broadcasting: a multimodal deep learning framework
Rizwan Abbas, Björn W. Schuller, Xi Li 0001
Multim. Syst.5
2025 Synth-CLIP: Synthetic data make CLIP generalize better in data-limited scenarios
Mushui Liu, Ziqian Lu, Jun Dan, Yunlong Yu 0001, Yingming Li, Xi Li 0001, Jungong Han
Neural Networks7
2025 Variational Adapter: Improving CLIP in Data-Imbalanced Scenarios
abstract
In this paper, we propose the Prompt-based Variational Adapter (PVA), a novel approach designed to fine-tune the pre-trained Vision-Language Models (VLMs) in data-imbalanced scenarios. Unlike existing methods that focus primarily on pairwise alignment of visual-text relationships during fine-tuning, PVA relaxes pairwise explicit constrains and emphasizes the harmonization of visual and text modality distributions, enhancing generalization and cross-modal understanding. To realize this harmonization, we develop two variational adapters, which are appended separately to the visual and text encoders. These adapters transform the feature embeddings into latent spaces that implicitly align with the corresponding modality distributions. We then adopt a divide-and-conquer strategy, dividing classes into data-abundant and data-limited sets to reduce prediction bias. Within each set, we independently fine-tune the models by incorporating both the model’s original general knowledge and specialized knowledge gained from training samples. Extensive experiments across two data-imbalanced scenarios validate the superiority of our approach, establishing a new state-of-the-art on popular benchmarks.
Ziqian Lu, Mushui Liu, Yunlong Yu 0001, Xi Li 0001, Jungong Han
IEEE Trans. Circuits Syst. Video Technol.5
2025 Scale- and Shape-Aware Network With Prediction Decoupling for Building Fine-Grained Change Detection
abstract
Building change detection (BCD) is a hot topic in geoscience and remote sensing (RS) with widespread applications. However, most existing BCD methods only focus on areas where changes have occurred, but ignore the change statuses. To address this problem, a building fine-grained change detection (BFCD) task is further explored in this work, which aims to judge the time-related “disappeared”, “appeared”, and “rebuilt” change types of buildings. Meanwhile, a scale- and shape-aware network (S2Net) with prediction decoupling is designed. Firstly, a prediction decoupling framework with dual decoders is built to ensure the prediction consistency with the temporal order of bi-temporal images. Secondly, considering the rebuilt type is the changes between building instances, which are often reflected in the scale and shape differences of the buildings. Thereby, a scale-aware module (ScAM) and a shape-aware module (ShAM) are designed. These two modules help extract the discriminative features of buildings with different scales and shapes for subsequent change detection (CD). In addition, two BCD datasets widely used, LEVIR-CD+ and WHU-CD, are relabeled in this work to support the study of BFCD. Experimental results show that S2Net achieves competitive performance, and its effectiveness is confirmed. The code and datasets will be publicly available at https://github.com/ptdoge/S2Net.
Chengcai Leng, Xi Li 0001, Irene Cheng 0001, Anup Basu, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2025 CLASH: Complementary Learning With Neural Architecture Search for Gait Recognition
abstract
Gait recognition, which aims at identifying individuals by their walking patterns, has achieved great success based on silhouette. The binary silhouette sequence encodes the walking pattern within the sparse boundary representation. Therefore, most pixels in the silhouette are under-sensitive to the walking pattern since the sparse boundary lacks dense spatial-temporal information, which is suitable to be represented with dense texture. To enhance the sensitivity to the walking pattern while maintaining the robustness of recognition, we present a Complementary Learning with neural Architecture SearcH (CLASH) framework, consisting of walking pattern sensitive gait descriptor named dense spatial-temporal field (DSTF) and neural architecture search based complementary learning (NCL). Specifically, DSTF transforms the representation from the sparse binary boundary into the dense distance-based texture, which is sensitive to the walking pattern at the pixel level. Further, NCL presents a task-specific search space for complementary learning, which mutually complements the sensitivity of DSTF and the robustness of the silhouette to represent the walking pattern effectively. Extensive experiments demonstrate the effectiveness of the proposed methods under both in-the-lab and in-the-wild scenarios. On CASIA-B, we achieve rank-1 accuracy of 98.8%, 96.5%, and 89.3% under three conditions. On OU-MVLP, we achieve rank-1 accuracy of 91.9%. Under the latest in-the-wild datasets, we outperform the latest silhouette-based methods by 16.3% and 19.7% on Gait3D and GREW, respectively.
Huanzhang Dou, Xi Li 0001
IEEE Trans. Image Process.5
2025 Associate Everything Detected: Facilitating Tracking-by-Detection to the Unknown
abstract
Multi-object tracking (MOT) emerges as a pivotal and highly promising branch in the field of computer vision. Classical closed-vocabulary MOT (CV-MOT) methods aim to track objects of predefined categories. Recently, some open-vocabulary MOT (OV-MOT) methods have successfully addressed the problem of tracking unknown categories. However, we found that the CV-MOT and OV-MOT methods each struggle to excel in the tasks of the other. In this paper, we present a unified framework, Associate Everything Detected (AED), that simultaneously tackles CV-MOT and OV-MOT by integrating with any off-the-shelf detector and supports unknown categories. Different from existing tracking-by-detection MOT methods, AED gets rid of prior knowledge (e.g., motion cues) and relies solely on highly robust feature learning to handle complex trajectories in OV-MOT tasks while keeping excellent performance in CV-MOT tasks. Specifically, we model the association task as a similarity decoding problem and propose a sim-decoder with an association-centric learning mechanism. The sim-decoder calculates similarities in three aspects: spatial, temporal, and cross-clip. Subsequently, association-centric learning leverages these threefold similarities to ensure that the extracted features are appropriate for continuous tracking and robust enough to generalize to unknown categories. Compared with existing powerful OV-MOT and CV-MOT methods, AED achieves superior performance on TAO, SportsMOT, and DanceTrack without any prior knowledge. Our code is available at https://github.com/balabooooo/AED.
Zimeng Fang, Shuyuan Zhu, Xi Li 0001
IEEE Trans. Image Process.5
2025 Faces Blind Your Eyes: Unveiling the Content-Irrelevant Synthetic Artifacts for Deepfake Detection
abstract
Data synthesis methods have shown promising results in general deepfake detection tasks. This is attributed to the inherent blending process in deepfake creation, which leaves behind distinct synthetic artifacts. However, the existence of content-irrelevant artifacts has not been explicitly explored in the deepfake synthesis. Unveiling content-irrelevant synthetic artifacts helps uncover general deepfake features and enhances the generalization capability of detection models. To capture the content-irrelevant synthetic artifacts, we propose a learning framework incorporating a synthesis process for diverse contents and specially designed learning strategies that encourage using content-irrelevant forgery information across deepfake images. From the data perspective, we disentangle the blending operation from face data and propose a universal synthetic module that generates images from various classes with common synthetic artifacts. From the learning perspective, a domain-adaptive learning head is introduced to filter out forgery-irrelevant features and optimize the decision on deepfake face detection. To efficiently learn the content-irrelevant artifacts for detection with a large sampling space, we propose a batch-wise sample selection strategy that actively mines the hard samples based on their effect on the adaptive decision boundary. Extensive cross-dataset experiments show that our method achieves state-of-the-art performance in general deepfake detection.
Xinghe Fu, Benzun Fu, Shen Chen 0004, Taiping Yao, Shouhong Ding, Xiubo Liang, Xi Li 0001
IEEE Trans. Image Process.8
2025 Relationship-Incremental Scene Graph Generation by a Divide-and-Conquer Pipeline With Feature Adapter
abstract
As a challenging computer vision task, Scene Graph Generation (SGG) finds the latent semantic relationships among objects from a given image, which may be limited by the datasets and real-world scenarios. In this paper, we consider a novel incremental learning task called Relationship-Incremental Scene Graph Generation (RISGG) that learns the semantic relationships among objects in an incremental way. Compared with classic Class-Incremental Learning (CIL) problem, RISGG suffers from its special issues: 1) Old class shift - the relationship-labeled object pair may have different labels during different learning sessions; 2) Background shift - the relationship-unlabeled object pair may not be a real unlabeled one. In this work, we address the above issues from the following aspects. First, we present a Divide-and-Conquer (DaC) pipeline to deal with the old class shift via decoupling the recognition of relationship classes and recognizing relationships individually. In this way, label confusion and interaction among different relationships are eliminated during training. Second, we propose a Feature Adapter (FA) to bridge the feature space gap between the current session and the previous one and use our extra supervision to mine old relationship information in the current session. Our proposed network combined DaC and FA, abbreviated DaCFA-Net, for RISGG. Experimental results on the benchmark dataset demonstrate the significant performance gain of DaCFA-Net in RISGG. It gains about 20% improvement against the SGG baselines on the popular VG dataset.
Xuewei Li 0003, Guangcong Zheng, Yunlong Yu 0001, Naye Ji, Xi Li 0001
IEEE Trans. Image Process.5
2025 Decoupling Discriminative Attributes for Few-Shot Fine-Grained Recognition
abstract
Few-shot fine-tuning of pre-trained vision-language models (VLMs) for downstream tasks has gained widespread attention for reducing data annotation efforts while maintaining high performance. However, we observe that VLMs excel in excluding most incorrect classes in fine-grained recognition tasks, but struggles with a small set of confusing categories, which are typically highly similar subspecies. Existing few-shot fine-tuning methods attempt to directly recognize the correct category among all predefined classes, limiting their ability to capture discriminative features for those confusing categories. This raises an intriguing question: Can we specifically extract useful information from confusing classes to enhance fine-grained recognition performance? Based on this insight, we propose a hierarchical few-shot fine-tuning framework to address the severe confusion problem while ensuring the interpretability, namely Attribute-Decoupled Discriminator (AttrDD). Instead of thinking once among all classes, AttrDD employs a two-stage recognition, "think through" then "think smart". Specifically, in the first phase, a representative VLM, CLIP, is fine-tuned to select the Top-K confusing classes. In the second phase, we leverage the knowledge of large language models (LLMs) to generate fixed format descriptions of attribute differences between these confusing classes via in-context learning. Attribute-decoupled classifications are then conducted to capture fine-grained discriminative features. To achieve parameter-efficient fine-tuning, we introduce a lightweight attention adapter for each phase to align image features with task-specific textual features and LLM-generated textual features. Extensive experiments on 9 fine-grained recognition benchmarks demonstrate that AttrDD consistently outperforms existing baselines by wide margins.
Yehao Lu, Chaoxiang Cai, Wei Su 0009, Guangcong Zheng, Xuewei Li 0003, Xi Li 0001
IEEE Trans. Image Process.7
2025 HeightFormer: Explicit Height Modeling Without Extra Data for Camera-Only 3D Object Detection in Bird's Eye View
abstract
Vision-based Bird's Eye View (BEV) representation is an emerging perception formulation for autonomous driving. The core challenge is to construct BEV space with multi-camera features, which is a one-to-many ill-posed problem. Diving into all previous BEV representation generation methods, we found that most of them fall into two types: modeling depths in image views or modeling heights in the BEV space, mostly in an implicit way. In this work, we propose to explicitly model heights in the BEV space, which needs no extra data like LiDAR and can fit arbitrary camera rigs and types compared to modeling depths. Theoretically, we give proof of the equivalence between height-based methods and depth-based methods. Considering the equivalence and some advantages of modeling heights, we propose HeightFormer, which models heights and uncertainties in a self-recursive way. Without any extra data, the proposed HeightFormer could estimate heights in BEV accurately. Benchmark results show that the performance of HeightFormer achieves SOTA compared with those camera-only methods.
Yiming Wu 0006, Zequn Qin, Xinhai Zhao, Xi Li 0001
IEEE Trans. Image Process.5
2025 GCSTG: Generating Class-Confusion-Aware Samples With a Tree-Structure Graph for Few-Shot Object Detection
abstract
Few-Shot Object Detection (FSOD) aims to detect the objects of novel classes using only a few manually annotated samples. With the few novel class samples, learning the inter-class relationships among foreground and constructing the corresponding class hierarchy in FSOD is a challenging task. The poor construction of the class hierarchy will result in the inter-class confusion problem, which has been identified as a primary cause of inferior performance in novel classes by recent FSOD methods. In this work, we further find that the intra-super-class confusion, where samples are misclassified as classes within their associated super-classes, is the main challenge in solving the confusion problem. To solve this issue, this work generates class-confusion-aware samples with a pre-defined tree-structure graph, for helping models to construct a precise class hierarchy. In precise, for generating class-confusion-aware samples, we add the noise into available samples and update the noise to maximize confidence scores on associated confusion categories of samples. Then, a confusion-aware curriculum learning strategy is proposed to make generated samples gradually participate in the training, which benefits the model convergence while learning the generated samples. Experimental results show that our method can be used as a plug-in in recent FSOD methods and consistently improve the model performance.
Longrong Yang, Hanbin Zhao, Hongliang Li 0001, Liang Qiao 0001, Xi Li 0001
IEEE Trans. Image Process.6
2025 Balancing Feature Alignment and Uniformity for Few-Shot Classification
abstract
In Few-Shot Learning (FSL), the objective is to correctly recognize new samples from novel classes with only a few available samples per class. Existing methods in FSL primarily focus on learning transferable knowledge from base classes by maximizing the information between feature representations and their corresponding labels. However, this approach may suffer from the "supervision collapse" issue, which arises due to a bias towards the base classes. In this paper, we propose a solution to address this issue by preserving the intrinsic structure of the data and enabling the learning of a generalized model for the novel classes. Following the InfoMax principle, our approach maximizes two types of mutual information (MI): between the samples and their feature representations, and between the feature representations and their class labels. This allows us to strike a balance between discrimination (capturing class-specific information) and generalization (capturing common characteristics across different classes) in the feature representations. To achieve this, we adopt a unified framework that perturbs the feature embedding space using two low-bias estimators. The first estimator maximizes the MI between a pair of intra-class samples, while the second estimator maximizes the MI between a sample and its augmented views. This framework effectively combines knowledge distillation between class-wise pairs and enlarges the diversity in feature representations. By conducting extensive experiments on popular FSL benchmarks, our proposed approach achieves comparable performances with state-of-the-art competitors. For example, we achieved an accuracy of 69.53% on the miniImageNet dataset and 77.06% on the CIFAR-FS dataset for the 5-way 1-shot task.
Yunlong Yu 0001, Dingyi Zhang, Zhong Ji, Xi Li 0001, Jungong Han, Zhongfei Zhang
IEEE Trans. Image Process.4
2025 GAMA-Pose: Graph-Aware Multi-Representation Aggregation for 3D Human Pose Estimation
abstract
Monocular 3D human pose estimation presents a considerable challenge owing to the intrinsic depth ambiguity associated with single-camera observations. Existing methods primarily rely on mean per joint position error (MPJPE) loss to train models for the conversion from 2D to 3D coordinates. However, empirical analysis reveals that models trained solely with point-based supervision may produce biomechanically implausible poses or exhibit significant depth ambiguity, even when achieving low MPJPE. This limitation arises from the fact that point-based loss only considers individual joint locations without accounting for inter-joint relationships. Fortunately, edges of human pose encode critical prior knowledge, including skeleton connectivity and biomechanical distributions. Explicitly modeling edge representations enables the model to overcome the constraints associated with point-only approaches, reducing the uncertainty in the optimization process of the 2D-3D inverse mapping and directly constraining depth ambiguity. Therefore, we propose the Graph-Aware Multi-Representation Aggregation (GAMA-Pose) framework that jointly predicts points and edges, with their fusion serving as the final output. To ensure the accuracy of edge predictions and mitigate depth ambiguity, Anti-Depth-Ambiguity Loss (ADA-Loss) is introduced to supervise the properties of edges and give direct supervision on depth ambiguity. Correspondingly, edge-based metrics are proposed to quantify the error of predicted edges. Experiments conducted on Human3.6M and MPI-INF-3DHP datasets demonstrate that GAMA-Pose effectively addresses the limitations of models relying solely on point constraints, mitigates depth ambiguity, enhances the accuracy of both point and edge predictions, and achieves state-of-the-art (SOTA) performance on both datasets.
Songran Zhou, Xuewei Li 0003, Xiubo Liang, Naye Ji, Xi Li 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2024 SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion Model
abstract
Controllable spherical panoramic image generation holds substantial applicative potential across a variety of domains. However, it remains a challenging task due to the inherent spherical distortion and geometry characteristics, resulting in low-quality content generation. In this paper, we introduce a novel framework of SphereDiffusion to address these unique challenges, for better generating high-quality and precisely controllable spherical panoramic images. For the spherical distortion characteristic, we embed the semantics of the distorted object with text encoding, then explicitly construct the relationship with text-object correspondence to better use the pre-trained knowledge of the planar images. Meanwhile, we employ a deformable technique to mitigate the semantic deviation in latent space caused by spherical distortion. For the spherical geometry characteristic, in virtue of spherical rotation invariance, we improve the data diversity and optimization objectives in the training process, enabling the model to better learn the spherical geometry characteristic. Furthermore, we enhance the denoising process of the diffusion model, enabling it to effectively use the learned geometric characteristic to ensure the boundary continuity of the generated images. With these specific techniques, experiments on Structured3D dataset show that SphereDiffusion significantly improves the quality of controllable spherical image generation and relatively reduces around 35% FID on average.
Xuewei Li 0003, Zhongang Qi, Xintao Wang 0002, Ying Shan, Xi Li 0001
AAAI7
2024 ScanFormer: Referring Expression Comprehension by Iteratively Scanning
abstract
Referring Expression Comprehension (REC) aims to localize the target objects specified by free-form natural language descriptions in images. While state-of-the-art methods achieve impressive performance, they perform a dense perception of images, which incorporates redundant visual regions unrelated to linguistic queries, leading to additional computational overhead. This inspires us to explore a question: can we eliminate linguistic-irrelevant redundant visual regions to improve the efficiency of the model? Existing relevant methods primarily focus on fundamental visual tasks, with limited exploration in vision-language fields. To address this, we propose a coarse-to-fine iterative perception framework, called ScanFormer. It can iteratively exploit the image scale pyramid to extract linguistic-relevant visual patches from top to bottom. In each iteration, irrelevant patches are discarded by our designed informativeness prediction. Furthermore, we propose a patch selection strategy for discarded patches to accelerate inference. Experiments on widely used datasets, namely Ref COCO, Ref COCO+, Ref COCO g, and ReferItGame, verify the effectiveness of our method, which can strike a balance between accuracy and efficiency.
Wei Su 0009, Peihan Miao 0002, Huanzhang Dou, Xi Li 0001
CVPR4
2024 BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-Based Roadside 3D Object Detection
abstract
Vision-based roadside 3D object detection has attracted rising attention in autonomous driving domain, since it en-compasses inherent advantages in reducing blind spots and expanding perception range. While previous work mainly focuses on accurately estimating depth or height for 2D-to-3D mapping, ignoring the position approximation error in the voxel pooling process. Inspired by this insight, we propose a novel voxel pooling strategy to reduce such error, dubbed BEVSpread. Specifically, instead of bringing the image features contained in a frustum point to a single BEV grid, BEVSpread considers each frustum point as a source and spreads the image features to the surrounding BEV grids with adaptive weights. To achieve superior prop- agation performance, a specific weight function is designed to dynamically control the decay speed of the weights according to distance and depth. Aided by customized CUDA parallel acceleration, BEVSpread achieves comparable inference time as the original voxel pooling. Extensive experiments on two large-scale roadside benchmarks demonstrate that, as a plug-in, BEVSpread can significantly improve the performance of existing frustum-based BEV methods by a large margin of (1.12, 5.26, 3.01) AP in vehicle, pedestrian and cyclist. The source code will be made publicly available at BEVSpread.
Yehao Lu, Guangcong Zheng, Shuigen Zhan, Xiaoqing Ye, Zichang Tan, Jingdong Wang 0001, Gaoang Wang, Xi Li 0001
CVPR9
2024 Improving Zero-Shot Generalization for CLIP with Variational Adapter
Ziqian Lu, Mushui Liu, Yunlong Yu 0001, Xi Li 0001
ECCV (20)5
2024 RCS-Prompt: Learning Prompt to Rearrange Class Space for Prompt-Based Continual Learning
Longrong Yang, Hanbin Zhao, Yunlong Yu 0001, Xiaodong Zeng, Xi Li 0001
ECCV (47)5
2024 MRCAN: Multi-scale Region Correlation-driven Adaptive Normalization for Image Harmonization
abstract
Image composition serves as a vital data augmentation technique commonly utilized for intelligent model training. In order to facilitate the optimization of image composition and improve the authenticity and efficacy of the composite images, this paper delves into the composite image harmonization technique to adjust the appearance of the foreground to be harmonious with the background. Current methods usually overlook the interrelation between the foreground and background content, typically transferring the style directly from the whole background to the foreground. In addition, conventional normalization methods were prone to a degradation in image quality of the foreground region after harmonization. To address these issues, we propose an innovative Adaptive Normalization method, which modulates the mean and standard deviation of foreground region specifically and pointedly, to harmonize the appearance while simultaneously preserving the texture structures of original foreground by considering the local information from the deviation maps. Moreover, we incorporate a Multi-scale Region Correlation-driven strategy to explore the correlation between the foreground and the background content, enabling the foreground object to fit the background semantics more seamlessly. Both ablation and comparison experiments on the iHarmony4 dataset demonstrate the effectiveness of our proposed method as well as the superiority of our model over other state-of-the-art image harmonization methods.
Luwen Duan, Hongliang Lou, Xi Li 0001
SMC5
2024 A Survey of Multimodal Controllable Diffusion Models
Guangcong Zheng, Tian-Rui Yang, Jingdong Wang 0001, Xi Li 0001
J. Comput. Sci. Technol.6
2024 Multimodal pre-train then transfer learning approach for speaker recognition
Summaira Jabeen, Amin Muhammad Shoib, Xi Li 0001
Multim. Tools Appl.3
2024 Ultra Fast Deep Lane Detection With Hybrid Anchor Driven Ordinal Classification
abstract
Modern methods mainly regard lane detection as a problem of pixel-wise segmentation, which is struggling to address the problems of efficiency and challenging scenarios like severe occlusions and extreme lighting conditions. Inspired by human perception, the recognition of lanes under severe occlusions and extreme lighting conditions is mainly based on contextual and global information. Motivated by this observation, we propose a novel, simple, yet effective formulation aiming at ultra fast speed and the problem of challenging scenarios. Specifically, we treat the process of lane detection as an anchor-driven ordinal classification problem using global features. First, we represent lanes with sparse coordinates on a series of hybrid (row and column) anchors. With the help of the anchor-driven representation, we then reformulate the lane detection task as an ordinal classification problem to get the coordinates of lanes. Our method could significantly reduce the computational cost with the anchor-driven representation. Using the large receptive field property of the ordinal classification formulation, we could also handle challenging scenarios. Extensive experiments on four lane detection datasets show that our method could achieve state-of-the-art performance in terms of both speed and accuracy. A lightweight version could even achieve 300+ frames per second(FPS). Our code is at https://github.com/cfzd/Ultra-Fast-Lane-Detection-v2.
Zequn Qin, Xi Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 MgSvF: Multi-Grained Slow versus Fast Framework for Few-Shot Class-Incremental Learning
abstract
As a challenging problem, few-shot class-incremental learning (FSCIL) continually learns a sequence of tasks, confronting the dilemma between slow forgetting of old knowledge and fast adaptation to new knowledge. In this paper, we concentrate on this "slow versus fast" (SvF) dilemma to determine which knowledge components to be updated in a slow fashion or a fast fashion, and thereby balance old-knowledge preservation and new-knowledge adaptation. We propose a multi-grained SvF learning strategy to cope with the SvF dilemma from two different grains: intra-space (within the same feature space) and inter-space (between two different feature spaces). The proposed strategy designs a novel frequency-aware regularization to boost the intra-space SvF capability, and meanwhile develops a new feature space composition operation to enhance the inter-space SvF learning performance. With the multi-grained SvF learning strategy, our method outperforms the state-of-the-art approaches by a large margin.
Hanbin Zhao, Yongjian Fu 0002, Mintong Kang, Qi Tian 0001, Fei Wu 0001, Xi Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Reading order detection in visually-rich documents with multi-modal layout-aware relation prediction
Liang Qiao 0001, Zhanzhan Cheng, Yunlu Xu, Xi Li 0001
Pattern Recognit.6
2024 GaitMPL: Gait Recognition With Memory-Augmented Progressive Learning
abstract
Gait recognition aims at identifying the pedestrians at a long distance by their biometric gait patterns. It is inherently challenging due to the various covariates and the properties of silhouettes (textureless and colorless), which result in two kinds of pair-wise hard samples: the same pedestrian could have distinct silhouettes (intra-class diversity) and different pedestrians could have similar silhouettes (inter-class similarity). In this work, we propose to solve the hard sample issue with a Memory-augmented Progressive Learning network (GaitMPL), including Dynamic Reweighting Progressive Learning module (DRPL) and Global Structure-Aligned Memory bank (GSAM). Specifically, DRPL reduces the learning difficulty of hard samples by easy-to-hard progressive learning. GSAM further augments DRPL with a structure-aligned memory mechanism, which maintains and models the feature distribution of each ID. Experiments on two commonly used datasets, CASIA-B and OU-MVLP, demonstrate the effectiveness of GaitMPL. On CASIA-B, we achieve the state-of-the-art performance, i.e., 88.0% on the most challenging condition (Clothing) and 93.3% on the average condition, which outperforms the other methods by at least 3.8% and 1.4%, respectively. Code will be available at https://github.com/WhiteDOU/GaitMPL https://github.com/WhiteDOU/GaitMPL.
Huanzhang Dou, Zequn Qin, Xi Li 0001
IEEE Trans. Image Process.6
2024 MLMG-SGG: Multilabel Scene Graph Generation With Multigrained Features
abstract
As an important and challenging problem in computer vision, scene graph generation (SGG) aims to find out the underlying semantic relationships among objects from a given image for scene understanding. Usually, prevalent SGG approaches adopt a learning pipeline with the assumption that there exists only a single relationship for a particular object pair. Considering the common phenomenon that a pair of objects can be attached by multiple relationships, we propose a multi-label scene graph generation pipeline with multi-grained features (MLMG-SGG), which formulates the relationship detection as a multi-label classification problem during training while generating multigraphs at inference time. In order to better model the fine-grained relationships, the proposed pipeline encodes the feature representation of SGG on different spatial scales by a specially designed Multi-Grained Module (MGM), resulting in the multi-grained (i.e., object-level and region-level) features of objects. Experimental results over the benchmark dataset demonstrate the significant performance gain of the proposed pipeline used as a plug-in for the state-of-the-art methods.
Xuewei Li 0003, Peihan Miao 0002, Xi Li 0001
IEEE Trans. Image Process.4
2024 IDNet: Information Decomposition Network for Fast Panoptic Segmentation
abstract
Traditional CNN-based pipelines for panoptic segmentation decompose the task into two subtasks, i.e., instance segmentation and semantic segmentation. In this way, they extract information with multiple branches, perform two subtasks separately and finally fuse the results. However, excessive feature extraction and complicated processes make them time-consuming. We propose IDNet to decompose panoptic segmentation at information level. IDNet only extracts two kinds of information and directly completes panoptic segmentation task, saving the efforts to extract extra information and to fuse subtasks. By decomposing panoptic segmentation into category information and location information and recomposing them with a serial pipeline, the process for panoptic segmentation is simplified greatly and unified with regard to stuff and things. We also adopt two correction losses specially designed for our serial pipeline, guaranteeing the overall predicting performance. As a result, IDNet strikes a better balance between effectiveness and efficiency, achieving the fastest inference speed of 24.2 FPS at a resolution of 800×1333 on a Tesla V100 GPU and a PQ of 43.8, which is comparable in one-stage CNN-based methods. The code will be released at https://github.com/AronLin/IDNet.
Guangchen Lin, Xi Li 0001
IEEE Trans. Image Process.4
2024 Self-Paced Multi-Grained Cross-Modal Interaction Modeling for Referring Expression Comprehension
abstract
As an important and challenging problem in vision-language tasks, referring expression comprehension (REC) generally requires a large amount of multi-grained information of visual and linguistic modalities to realize accurate reasoning. In addition, due to the diversity of visual scenes and the variation of linguistic expressions, some hard examples have much more abundant multi-grained information than others. How to aggregate multi-grained information from different modalities and extract abundant knowledge from hard examples is crucial in the REC task. To address aforementioned challenges, in this paper, we propose a Self-paced Multi-grained Cross-modal Interaction Modeling framework, which improves the language-to-vision localization ability through innovations in network structure and learning mechanism. Concretely, we design a transformer-based multi-grained cross-modal attention, which effectively utilizes the inherent multi-grained information in visual and linguistic encoders. Furthermore, considering the large variance of samples, we propose a self-paced sample informativeness learning to adaptively enhance the network learning for samples containing abundant multi-grained information. The proposed framework significantly outperforms state-of-the-art methods on widely used datasets, such as RefCOCO, RefCOCO+, RefCOCOg, and ReferItGame datasets, demonstrating the effectiveness of our method.
Peihan Miao 0002, Wei Su 0009, Gaoang Wang, Xuewei Li 0003, Xi Li 0001
IEEE Trans. Image Process.5
2024 Epoch-Evolving Gaussian Process Guided Learning for Classification
abstract
The conventional mini-batch gradient descent algorithms are usually trapped in the local batch-level distribution information, resulting in the ``zig-zag'' effect in the learning process. To characterize the correlation information between the batch-level distribution and the global data distribution, we propose a novel learning scheme called epoch-evolving Gaussian process guided learning (GPGL) to encode the global data distribution information in a non-parametric way. Upon a set of class-aware anchor samples, our GP model is built to estimate the class distribution for each sample in mini-batch through label propagation from the anchor samples to the batch samples. The class distribution, also named the context label, is provided as a complement for the ground-truth one-hot label. Such a class distribution structure has a smooth property and usually carries a rich body of contextual information that is capable of speeding up the convergence process. With the guidance of the context label and ground-truth label, the GPGL scheme provides a more efficient optimization through updating the model parameters with a triangle consistency loss. Furthermore, our GPGL scheme can be generalized and naturally applied to the current deep models, outperforming the state-of-the-art optimization methods on six benchmark datasets.
Jiabao Cui, Xuewei Li 0003, Hanbin Zhao, Hui Wang 0107, Bin Li 0038, Xi Li 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Temporal-Frequency Attention Focusing for Time Series Extrinsic Regression via Auxiliary Task
abstract
Time series extrinsic regression (TSER) aims at predicting numeric values based on the knowledge of the entire time series. The key to solving the TSER problem is to extract and use the most representative and contributed information from raw time series. To build a regression model that focuses on those information suitable for the extrinsic regression characteristic, there are two major issues to be addressed. That is, how to quantify the contributions of those information extracted from raw time series and then how to focus the attention of the regression model on those critical information to improve the model's regression performance. In this article, a multitask learning framework called temporal-frequency auxiliary task (TFAT) is designed to solve the mentioned problems. To explore the integral information from the time and frequency domains, we decompose the raw time series into multiscale subseries in various frequencies via a deep wavelet decomposition network. To address the first problem, the transformer encoder with the multihead self-attention mechanism is integrated in our TFAT framework to quantify the contribution of temporal-frequency information. To address the second problem, an auxiliary task in a manner of self-supervised learning is proposed to reconstruct the critical temporal-frequency features so as to focusing the regression model's attention on those essential information for facilitating TSER performance. We estimated three kinds of attention distribution on those temporal-frequency features to perform auxiliary task. To evaluate the performances of our method under various application scenarios, the experiments are carried out on the 12 datasets of the TSER problem. Also, ablation studies are used to examine the effectiveness of our method.
Lei Ren 0001, Tingyu Mo, Xuejun Cheng, Xi Li 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Unsupervised Domain Adaptation With Class-Aware Memory Alignment
abstract
Unsupervised domain adaptation (UDA) is to make predictions on unlabeled target domain by learning the knowledge from a label-rich source domain. In practice, existing UDA approaches mainly focus on minimizing the discrepancy between different domains by mini-batch training, where only a few instances are accessible at each iteration. Due to the randomness of sampling, such a batch-level alignment pattern is unstable and may lead to misalignment. To alleviate this risk, we propose class-aware memory alignment (CMA) that models the distributions of the two domains by two auxiliary class-aware memories and performs domain adaptation on these predefined memories. CMA is designed with two distinct characteristics: class-aware memories that create two symmetrical class-aware distributions for different domains and two reliability-based filtering strategies that enhance the reliability of the constructed memory. We further design a unified memory-based loss to jointly improve the transferability and discriminability of features in the memories. State-of-the-art (SOTA) comparisons and careful ablation studies show the effectiveness of our proposed CMA.
Hui Wang 0107, Liangli Zheng, Hanbin Zhao, Shijian Li, Xi Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 Referring Expression Comprehension Using Language Adaptive Inference
abstract
Different from universal object detection, referring expression comprehension (REC) aims to locate specific objects referred to by natural language expressions. The expression provides high-level concepts of relevant visual and contextual patterns, which vary significantly with different expressions and account for only a few of those encoded in the REC model. This leads us to a question: do we really need the entire network with a fixed structure for various referring expressions? Ideally, given an expression, only expression-relevant components of the REC model are required. These components should be small in number as each expression only contains very few visual and contextual clues. This paper explores the adaptation between expressions and REC models for dynamic inference. Concretely, we propose a neat yet efficient framework named Language Adaptive Dynamic Subnets (LADS), which can extract language-adaptive subnets from the REC model conditioned on the referring expressions. By using the compact subnet, the inference can be more economical and efficient. Extensive experiments on RefCOCO, RefCOCO+, RefCOCOg, and Referit show that the proposed method achieves faster inference speed and higher accuracy against state-of-the-art approaches.
Wei Su 0009, Peihan Miao 0002, Huanzhang Dou, Yongjian Fu 0002, Xi Li 0001
AAAI5
2023 GaitGCI: Generative Counterfactual Intervention for Gait Recognition
abstract
Gait is one of the most promising biometrics that aims to identify pedestrians from their walking patterns. However, prevailing methods are susceptible to confounders, resulting in the networks hardly focusing on the regions that re-flect effective walking patterns. To address this fundamen-tal problem in gait recognition, we propose a Generative Counterfactual Intervention framework, dubbed GaitGCI, consisting of Counterfactual Intervention Learning (CIL) and Diversity-Constrained Dynamic Convolution (DCDC). CIL eliminates the impacts of confounders by maximizing the likelihood difference between factual/counterfactual attention while DCDC adaptively generates sample-wise factual/counterfactual attention to efficiently perceive the sample-wise properties. With matrix decomposition and diversity constraint, DCDC guarantees the model to be efficient and effective. Extensive experiments indicate that proposed GaitGCI.· 1) could effectively focus on the discrimi-native and interpretable regions that reflect gait pattern; 2) is model-agnostic and could be plugged into existing models to improve performance with nearly no extra cost; 3) efficiently achieves state-of-the-art performance on arbitrary scenarios (in-the-lab and in-the-wild).
Huanzhang Dou, Wei Su 0009, Yunlong Yu 0001, Yining Lin, Xi Li 0001
CVPR6
2023 RWSC-Fusion: Region-Wise Style-Controlled Fusion Network for the Prohibited X-ray Security Image Synthesis
abstract
Automatic prohibited item detection in security inspection X-ray images is necessary for transportation. The abundance and diversity of the X-ray security images with prohibited item, termed as prohibited X-ray security images, are essential for training the detection model. In order to solve the data insufficiency, we propose a Region-Wise Style-Controlled Fusion (RWSC-Fusion) network, which superimposes the prohibited items onto the normal X-ray security images, to synthesize the prohibited X-ray security images. The proposed RWSC-Fusion innovates both network structure and loss functions to generate more realistic X-ray security images. Specifically, a RWSC-Fusion module is designed to enable the region-wise fusion by controlling the appearance of the overlapping region with novel modulation parameters. In addition, an Edge-Attention (EA) module is proposed to effectively improve the sharpness of the synthetic images. As for the unsupervised loss function, we propose the Luminance loss in Logarithmic form (LL) and Correlation loss of Saturation Difference (CSD), to optimize the fused X-ray security images in terms of luminance and saturation. We evaluate the authenticity and the training effect of the synthetic X-ray security images on private and public SIXray dataset. The results confirm that our synthetic images are reliable enough to augment the prohibited X-ray security images.
Luwen Duan, Lijian Mao, Jianping Xiong, Xi Li 0001
CVPR6
2023 Language Adaptive Weight Generation for Multi-Task Visual Grounding
abstract
Although the impressive performance in visual grounding, the prevailing approaches usually exploit the visual backbone in a passive way, i.e., the visual backbone extracts features with fixed weights without expression-related hints. The passive perception may lead to mismatches (e.g., redundant and missing), limiting further performance improvement. Ideally, the visual backbone should actively extract visual features since the expressions already provide the blueprint of desired visual features. The active perception can take expressions as priors to extract relevant visual features, which can effectively alleviate the mismatches. Inspired by this, we propose an active perception Visual Grounding framework based on Language Adaptive Weights, called VG-LAW. The visual backbone serves as an expression-specific feature extractor through dynamic weights generated for various expressions. Benefiting from the specific and relevant visual features extracted from the language-aware visual backbone, VG-LAW does not require additional modules for cross-modal interaction. Along with a neat multi-task head, VG-LAW can be competent in referring expression comprehension and segmentation jointly. Extensive experiments on four representative datasets, i.e., RefCOCO, RefCOCO+, RefCOCOg, and ReferItGame, validate the effectiveness of the proposed framework and demonstrate state-of-the-art performance.
Wei Su 0009, Peihan Miao 0002, Huanzhang Dou, Gaoang Wang, Liang Qiao 0001, Zheyang Li, Xi Li 0001
CVPR7
2023 LayoutDiffusion: Controllable Diffusion Model for Layout-to-Image Generation
abstract
Recently, diffusion models have achieved great success in image synthesis. However, when it comes to the layout-to-image generation where an image often has a complex scene of multiple objects, how to make strong control over both the global layout map and each detailed object remains a challenging task. In this paper, we propose a diffusion model named LayoutDiffusion that can obtain higher generation quality and greater controllability than the previous works. To overcome the difficult multimodal fusion of image and layout, we propose to construct a structural image patch with region information and transform the patched image into a special layout to fuse with the normal layout in a unified form. Moreover, Layout Fusion Module (LFM) and Object-aware Cross Attention (OaCA) are proposed to model the relationship among multiple objects and designed to be object-aware and position-sensitive, allowing for precisely controlling the spatial related information. Extensive experiments show that our LayoutDiffusion out-performs the previous SOTA methods on FID, CAS by relatively 46.35%,26.70% on COCO-stuff and 44.29%,41.82% on VG. Code is available at https://github.com/ZGCTroy/LayoutDiffusion.
Guangcong Zheng, Xianpan Zhou, Xuewei Li 0003, Zhongang Qi, Ying Shan, Xi Li 0001
CVPR6
2023 UniFusion: Unified Multi-view Fusion Transformer for Spatial-Temporal Representation in Bird's-Eye-View
abstract
Bird’s eye view (BEV) representation is a new perception formulation for autonomous driving, which is based on spatial fusion. Further, temporal fusion is also introduced in BEV representation and gains great success. In this work, we propose a new method that unifies both spatial and temporal fusion and merges them into a unified mathematical formulation. The unified fusion could not only provide a new perspective on BEV fusion but also brings new capabilities. With the proposed unified spatial-temporal fusion, our method could support long-range fusion, which is hard to achieve in conventional BEV methods. Moreover, the BEV fusion in our work is temporal-adaptive and the weights of temporal fusion are learnable. In contrast, conventional methods mainly use fixed and equal weights for temporal fusion. Besides, the proposed unified fusion could avoid information lost in conventional BEV fusion methods and make full use of features. Extensive experiments and ablation studies on the NuScenes dataset show the effectiveness of the proposed method and our method gains the state-of-the-art performance in the map and vehicle segmentation task.
Zequn Qin, Xiaozhi Chen, Xi Li 0001
ICCV5
2023 Bridging Cross-task Protocol Inconsistency for Distillation in Dense Object Detection
abstract
Knowledge distillation (KD) has shown potential for learning compact models in dense object detection. However, the commonly used softmax-based distillation ignores the absolute classification scores for individual categories. Thus, the optimum of the distillation loss does not necessarily lead to the optimal student classification scores for dense object detectors. This cross-task protocol inconsistency is critical, especially for dense object detectors, since the foreground categories are extremely imbalanced. To address the issue of protocol differences between distillation and classification, we propose a novel distillation method with cross-task consistent protocols, tailored for the dense object detection. For classification distillation, we address the cross-task protocol inconsistency problem by formulating the classification logit maps in both teacher and student models as multiple binary-classification maps and applying a binary-classification distillation loss to each map. For localization distillation, we design an IoU-based Localization Distillation Loss that is free from specific network structures and can be compared with existing localization distillation losses. Our proposed method is simple but effective, and experimental results demonstrate its superiority over existing methods. Code is available at https://github.com/TinyTigerPan/BCKD.
Longrong Yang, Xianpan Zhou, Xuewei Li 0003, Liang Qiao 0001, Zheyang Li, Ziwei Yang 0004, Gaoang Wang, Xi Li 0001
ICCV8
2023 SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation
abstract
As an important and challenging problem in computer vision, PAnoramic Semantic Segmentation (PASS) gives complete scene perception based on an ultra-wide angle of view. Usually, prevalent PASS methods with 2D panoramic image input focus on solving image distortions but lack consideration of the 3D properties of original 360 degree data. Therefore, their performance will drop a lot when inputting panoramic images with the 3D disturbance. To be more robust to 3D disturbance, we propose our Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation (SGAT4PASS), considering 3D spherical geometry knowledge. Specifically, a spherical geometry-aware framework is proposed for PASS. It includes three modules, i.e., spherical geometry-aware image projection, spherical deformable patch embedding, and a panorama-aware loss, which takes input images with 3D disturbance into account, adds a spherical geometry-aware constraint on the existing deformable patch embedding, and indicates the pixel density of original 360 degree data, respectively. Experimental results on Stanford2D3D Panoramic datasets show that SGAT4PASS significantly improves performance and robustness, with approximately a 2% increase in mIoU, and when small 3D disturbances occur in the data, the stability of our performance is improved by an order of magnitude. Our code and supplementary material are available at https://github.com/TencentARC/SGAT4PASS.
Xuewei Li 0003, Zhongang Qi, Gaoang Wang, Ying Shan, Xi Li 0001
IJCAI6
2023 DenseDINO: Boosting Dense Self-Supervised Learning with Token-Based Point-Level Consistency
abstract
In this paper, we propose a simple yet effective transformer framework for self-supervised learning called DenseDINO to learn dense visual representations. To exploit the spatial information that the dense prediction tasks require but neglected by the existing self-supervised transformers, we introduce point-level supervision across views in a novel token-based way. Specifically, DenseDINO introduces some extra input tokens called reference tokens to match the point-level features with the position prior. With the reference token, the model could maintain spatial consistency and deal with multi-object complex scene images, thus generalizing better on dense prediction tasks. Compared with the vanilla DINO, our approach obtains competitive performance when evaluated on classification in ImageNet and achieves a large margin (+7.2% mIoU) improvement in semantic segmentation on PascalVOC under the linear probing protocol for segmentation.
Yike Yuan, Xinghe Fu, Yunlong Yu 0001, Xi Li 0001
IJCAI4
2023 Finding Cycles in Graph: A Unified Approach for Various NER Tasks
abstract
Named Entity Recognition (NER) is the task of recognizing the entities' locations and types in text, which can be generally categorized into flat NER, overlapped NER, and discontinuous NER. Most previous methods are usually designed specifically for one of the tasks, such as sequence labeling approaches for flat NER and span-based models for overlapped NER. Recently, some new work has begun to propose the unified NER framework that can addresses all three scenarios simultaneously. However, there still has room for improvement in some complex scenarios (long/discontinuous entity). In this paper, we propose a concise framework that supports all types of NER tasks, where entities can be represented by unique cycles that are formed by the directed edges among tokens in the graph. The model integrates a Graph Feature Enhancement module to extract correlations at both the node-level and edge-level. At the node-level, the features are enhanced in the binary and ternary token relations. In edge-level, the model will go further to enhance the relations among token pairs using deformable convolutions. Furthermore, to benefit the completeness of cycle formation, we also propose a novel Cycle Loss that optimizes the independent edge classification in the group of cycles from a global perspective. Experimental results show that our model can achieve competitive and even new state-of-the-art performance on eight popular NER benchmarks, including flat NER, overlapped NER, and discontinuous NER.
Liang Qiao 0001, Xi Li 0001
IJCNN4
2023 Personalized Behavior-Aware Transformer for Multi-Behavior Sequential Recommendation
abstract
Sequential Recommendation (SR) captures users' dynamic preferences by modeling how users transit among items. However, SR models that utilize only single type of behavior interaction data encounter performance degradation when the sequences are short. To tackle this problem, we focus on Multi-Behavior Sequential Recommendation (MBSR) in this paper, which aims to leverage time-evolving heterogeneous behavioral dependencies for better exploring users' potential intents on the target behavior. Solving MBSR is challenging. On the one hand, users exhibit diverse multi-behavior patterns due to personal characteristics. On the other hand, there exists comprehensive co-influence between behavior correlations and item collaborations, the intensity of which is deeply affected by temporal factors. To tackle these challenges, we propose a Personalized Behavior-Aware Transformer framework (PBAT) for MBSR problem, which models personalized patterns and multifaceted sequential collaborations in a novel way to boost recommendation performance. PBAT includes two main modules, i.e, personalized behavior pattern generator and behavior-aware collaboration extractor. First, PBAT develops a personalized behavior pattern generator in the representation layer, which extracts dynamic and discriminative behavior patterns for sequential learning. Second, PBAT reforms the self-attention layer with a behavior-aware collaboration extractor, which introduces a fused behavior-aware attention mechanism for incorporating both behavioral and temporal impacts into collaborative transitions. We conduct experiments on three benchmark datasets and the results demonstrate the effectiveness and interpretability of our framework.
Jiajie Su, Chaochao Chen 0001, Zibin Lin, Xi Li 0001, Weiming Liu 0005
ACM Multimedia4
2023 Forgery face detection via adaptive learning from multiple experts
Xinghe Fu, Shengming Li, Yike Yuan, Bin Li 0038, Xi Li 0001
Neurocomputing5
2023 Adaptive cooperative exploration for reinforcement learning from imperfect demonstrations
Fuxian Huang, Naye Ji, Huajian Ni, Shijian Li, Xi Li 0001
Pattern Recognit. Lett.5
2023 Uncertainty-Aware Scene Graph Generation
Xuewei Li 0003, Guangcong Zheng, Yunlong Yu 0001, Xi Li 0001
Pattern Recognit. Lett.5
2023 Fuzzy Semantics for Arbitrary-Shaped Scene Text Detection
abstract
To robustly detect arbitrary-shaped scene texts, bottom-up methods are widely explored for their flexibility. Due to the highly homogeneous texture and cluttered distribution of scene texts, it is nontrivial for segmentation-based methods to discover the separatrixes between adjacent instances. To effectively separate nearby texts, many methods adopt the seed expansion strategy that segments shrunken text regions as seed areas, and then iteratively expands the seed areas into intact text regions. In seek of a more straightforward way that does not rely on seed area segmentation and avoid possible error accumulation brought by iterative processing, we propose a redundancy removal strategy. In this work, we directly explore two types of fuzzy semantics-text and separatrix-that do not possess specific boundaries, and separate cluttered instances by excluding the separatrix pixels from text regions. To deal with the fuzzy semantic boundaries, we also conduct reliability analysis in both optimization and inference stage to suppress false positive pixels at ambiguous locations. Experiments on benchmark datasets demonstrate the effectiveness of our method.
Xiaogang Xu 0001, Xi Li 0001
IEEE Trans. Image Process.4
2023 Elastic Knowledge Distillation by Learning From Recollection
abstract
Model performance can be further improved with the extra guidance apart from the one-hot ground truth. To achieve it, recently proposed recollection-based methods utilize the valuable information contained in the past training history and derive a "recollection" from it to provide data-driven prior to guide the training. In this article, we focus on two fundamental aspects of this method, i.e., recollection construction and recollection utilization. Specifically, to meet the various demands of models with different capacities and at different training periods, we propose to construct a set of recollections with diverse distributions from the same training history. After that, all the recollections collaborate together to provide guidance, which is adaptive to different model capacities, as well as different training periods, according to our similarity-based elastic knowledge distillation (KD) algorithm. Without any external prior to guide the training, our method achieves a significant performance gain and outperforms the methods of the same category, even as well as KD with well-trained teacher. Extensive experiments and further analysis are conducted to demonstrate the effectiveness of our method.
Yongjian Fu 0002, Hanbin Zhao, Wenfu Wang, Weihao Fang, Yueting Zhuang, Xi Li 0001
IEEE Trans. Neural Networks Learn. Syst.8
2023 A Review on Methods and Applications in Multimodal Deep Learning
abstract
Deep Learning has implemented a wide range of applications and has become increasingly popular in recent years. The goal of multimodal deep learning (MMDL) is to create models that can process and link information using various modalities. Despite the extensive development made for unimodal learning, it still cannot cover all the aspects of human learning. Multimodal learning helps to understand and analyze better when various senses are engaged in the processing of information. This article focuses on multiple types of modalities, i.e., image, video, text, audio, body gestures, facial expressions, physiological signals, flow, RGB, pose, depth, mesh, and point cloud. Detailed analysis of the baseline approaches and an in-depth study of recent advancements during the past five years (2017 to 2021) in multimodal deep learning applications has been provided. A fine-grained taxonomy of various multimodal deep learning methods is proposed, elaborating on different applications in more depth. Last, main issues are highlighted separately for each domain, along with their possible future research directions.
Summaira Jabeen, Xi Li 0001, Amin Muhammad Shoib, Omar El Farouk Bourahla
ACM Trans. Multim. Comput. Commun. Appl.2
2023 D3T-GAN: Data-Dependent Domain Transfer GANs for Image Generation with Limited Data
abstract
As an important and challenging problem, image generation with limited data aims at generating realistic images through training a GAN model given few samples. A typical solution is to transfer a well-trained GAN model from a data-rich source domain to the data-deficient target domain. In this paper, we propose a novel self-supervised transfer scheme termed D 3 T-GAN, addressing the cross-domain GANs transfer in limited image generation. Specifically, we design two individual strategies to transfer knowledge between generators and discriminators, respectively. To transfer knowledge between generators, we conduct a data-dependent transformation, which projects target samples into the latent space of source generator and reconstructs them back. Then, we perform knowledge transfer from transformed samples to generated samples. To transfer knowledge between discriminators, we design a multi-level discriminant knowledge distillation from the source discriminator to the target discriminator on both the real and fake samples. Extensive experiments show that our method improves the quality of generated images and achieves the state-of-the-art FID scores on commonly used datasets.
Xintian Wu, Yiming Wu 0005, Xi Li 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2023 A Large-Scale Synthetic Gait Dataset Towards in-the-Wild Simulation and Comparison Study
abstract
Gait recognition has a rapid development in recent years. However, current gait recognition focuses primarily on ideal laboratory scenes, leaving the gait in the wild unexplored. One of the main reasons is the difficulty of collecting in-the-wild gait datasets, which must ensure diversity of both intrinsic and extrinsic human gait factors. To remedy this problem, we propose to construct a large-scale gait dataset with the help of controllable computer simulation. In detail, to diversify the intrinsic factors of gait, we generate numerous characters with diverse attributes and associate them with various types of walking styles. To diversify the extrinsic factors of gait, we build a complicated scene with a dense camera layout. Then we design an automatic generation toolkit under Unity3D for simulating the walking scenarios and capturing the gait data. As a result, we obtain a dataset simulating towards the in-the-wild scenario, called VersatileGait, which has more than one million silhouette sequences of 10,000 subjects with diverse scenarios. VersatileGait possesses several nice properties, including huge dataset size, diverse pedestrian attributes, complicated camera layout, high-quality annotations, small domain gap with the real one, good scalability for new demands, and no privacy issues. By conducting a series of experiments, we first explore the effects of different factors on gait recognition. We further illustrate the effectiveness of using our dataset to pre-train models, which obtain considerable performance gain on CASIA-B, OU-MVLP, and CASIA-E. Besides, we show the great potential of the fine-grained labels other than the ID label in improving the efficiency and effectiveness of models. Our dataset and its corresponding generation toolkit are available at https://github.com/peterzpy/VersatileGait.
Huanzhang Dou, Wenhu Zhang, Zequn Qin, Dongping Hu, Xi Li 0001
ACM Trans. Multim. Comput. Commun. Appl.8
2022 MonoGround: Detecting Monocular 3D Objects from the Ground
abstract
Monocular 3D object detection has attracted great attention for its advantages in simplicity and cost. Due to the ill-posed 2D to 3D mapping essence from the monocular imaging process, monocular 3D object detection suffers from inaccurate depth estimation and thus has poor 3D detection results. To alleviate this problem, we propose to introduce the ground plane as a prior in the monocular 3d object detection. The ground plane prior serves as an additional geometric condition to the ill-posed mapping and an extra source in depth estimation. In this way, we can get a more accurate depth estimation from the ground. Meanwhile, to take full advantage of the ground plane prior, we propose a depth-align training strategy and a precise two-stage depth inference method tailored for the ground plane prior. It is worth noting that the introduced ground plane prior requires no extra data sources like LiDAR, stereo images, and depth information. Extensive experiments on the KITTI benchmark show that our method could achieve state-of-the-art results compared with other methods while maintaining a very fast speed. Our code, models, and training logs are available at https://github.com/cfzd/MonoGround.
Zequn Qin, Xi Li 0001
CVPR2
2022 Dynamic Low-Resolution Distillation for Cost-Efficient End-to-End Text Spotting
Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Xi Li 0001
ECCV (28)6
2022 MetaGait: Learning to Learn an Omni Sample Adaptive Representation for Gait Recognition
Huanzhang Dou, Wei Su 0009, Yunlong Yu 0001, Xi Li 0001
ECCV (5)5
2022 SP-Net: Slowly Progressing Dynamic Inference Networks
Wenhu Zhang, Shihao Su, Hui Wang 0107, Zhenwei Miao, Xin Zhan, Xi Li 0001
ECCV (11)7
2022 Adaptive Cross-domain Learning for Generalizable Person Re-identification
Huanzhang Dou, Yunlong Yu 0001, Xi Li 0001
ECCV (14)4
2022 Saliency Hierarchy Modeling via Generative Kernels for Salient Object Detection
Wenhu Zhang, Liangli Zheng, Xintian Wu, Xi Li 0001
ECCV (28)5
2022 RBC: Rectifying the Biased Context in Continual Semantic Segmentation
Hanbin Zhao, Fengyu Yang 0003, Xinghe Fu, Xi Li 0001
ECCV (34)4
2022 Entropy-Driven Sampling and Training Scheme for Conditional Diffusion Generation
Guangcong Zheng, Shengming Li, Hui Wang 0107, Taiping Yao, Shouhong Ding, Xi Li 0001
ECCV (22)7
2022 Explicitly Modeling Importance and Coherence for Timeline Summarization
abstract
Timeline summarization (TLS) identifies major events and generates short summaries on how the event evolves in a period of time. Existing timeline summarization methods generate summaries by considering the coverage and diversity of the content and temporized information but ignore the importance and coherence of sentences used in summary. However, ignoring such information often causes missing important facts in the generated TLS and confuses users. We propose a better approach for TLS by explicitly optimizing importance and coherence on top of coverage and diversity. We apply our approach to both direct and pipeline TLS frameworks. Experimental results show that our approach achieves better performance when compared with two state-of-the-art TLS methods.
Qianren Mao, Jianxin Li 0002, JiaZheng Wang, Xi Li 0001, Zheng Wang 0001
ICASSP4
2022 End-to-End Compound Table Understanding with Multi-Modal Modeling
abstract
Table is a widely used data form in webpages, spreadsheets, or PDFs to organize and present structural data. Although studies on table structure recognition have been successfully used to convert image-based tables into digital structural formats, solving many real problems still relies on further understanding of the table, such as cell relationship extraction. The current datasets related to table understanding are all based on the digit format. To boost research development, we release a new benchmark named ComFinTab with rich annotations that support both table recognition and understanding tasks. Unlike previous datasets containing the basic tables, ComFinTab contains a large ratio of compound tables, which is much more challenging and requires methods using multiple information sources. Based on the dataset, we also propose a uniform, concise task form with the evaluation metric to better evaluate the model's performance on the table understanding task in compound tables. Finally, a framework named CTUNet is proposed to integrate the compromised visual, semantic, and position features with a graph attention network, which can solve the table recognition task and the challenging table understanding task as a whole. Experimental results compared with some previous advanced table understanding methods demonstrate the effectiveness of our proposed model. Code and dataset are available at \urlhttps://github.com/hikopensource/DAVAR-Lab-OCR.
Zaisheng Li, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Xi Li 0001
ACM Multimedia8
2022 Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image Generation
abstract
As a challenging task, text-to-image generation aims to generate photo-realistic and semantically consistent images according to the given text descriptions. Existing methods mainly extract the text information from only one sentence to represent an image and the text representation effects the quality of the generated image well. However, directly utilizing the limited information in one sentence misses some key attribute descriptions, which are the crucial factors to describe an image accurately. To alleviate the above problem, we propose an effective text representation method with the complements of attribute information. Firstly, we construct an attribute memory to jointly control the text-to-image generation with sentence input. Secondly, we explore two update mechanisms, sample-aware and sample-joint mechanisms, to dynamically optimize a generalized attribute memory. Furthermore, we design an attribute-sentence-joint conditional generator learning scheme to align the feature embeddings among multiple representations, which promotes the cross-modal network training. Experimental results illustrate that the proposed method obtains substantial performance improvements on both the CUB (FID from 14.81 to 8.57) and COCO (FID from 21.42 to 12.39) datasets.
Xintian Wu, Hanbin Zhao, Liangli Zheng, Shouhong Ding, Xi Li 0001
ACM Multimedia5
2022 Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search
Peng Sun 0011, Wenhu Zhang, Congli Song, Xi Li 0001
Int. J. Comput. Vis.6
2022 PcmNet: Position-sensitive context modeling network for temporal action localization
Hanbin Zhao, Guangchen Lin, Songcen Xu, Xi Li 0001
Neurocomputing6
2022 Structure-conditioned adversarial learning for unsupervised domain adaptation
Hui Wang 0107, Hanbin Zhao, Fei Wu 0001, Xi Li 0001
Neurocomputing6
2022 Attend and select: A segment selective transformer for microblog hashtag generation
Qianren Mao, Xi Li 0001, Jianxin Li 0002
Knowl. Based Syst.2
2022 TapLab: A Fast Framework for Semantic Video Segmentation Tapping Into Compressed-Domain Knowledge
abstract
Real-time semantic video segmentation is a challenging task due to the strict requirements of inference speed. Recent approaches mainly devote great efforts to reducing the model size for high efficiency. In this paper, we rethink this problem from a different viewpoint: using knowledge contained in compressed videos. We propose a simple and effective framework, dubbed TapLab, to tap into resources from the compressed domain. Specifically, we design a fast feature warping module using motion vectors for acceleration. To reduce the noise introduced by motion vectors, we design a residual-guided correction module and a residual-guided frame selection module using residuals. TapLab significantly reduces redundant computations of the state-of-the-art fast semantic image segmentation models, running 3 to 10 times faster with controllable accuracy degradation. The experimental results show that TapLab achieves 70.6 percent mIoU on the Cityscapes dataset at 99.8 FPS with a single GPU card for the 1024×2048 videos. A high-speed version even reaches the speed of 160+ FPS. Code will be available soon at https://github.com/Sixkplus/TapLab.
Junyi Feng, Xi Li 0001, Fei Wu 0001, Qi Tian 0001, Ming-Hsuan Yang 0001, Haibin Ling
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 CoDiNet: Path Distribution Modeling With Consistency and Diversity for Dynamic Routing
abstract
Dynamic routing networks, aimed at finding the best routing paths in the networks, have achieved significant improvements to neural networks in terms of accuracy and efficiency. In this paper, we see dynamic routing networks in a fresh light, formulating a routing method as a mapping from a sample space to a routing space. From the perspective of space mapping, prevalent methods of dynamic routing did not take into account how inference paths would be distributed in the routing space. Thus, we propose a novel method, termed CoDiNet, to model the relationship between a sample space and a routing space by regularizing the distribution of routing paths with the properties of consistency and diversity. Specifically, samples with similar semantics should be mapped into the same area in routing space, while those with dissimilar semantics should be mapped into different areas. Moreover, we design a customizable dynamic routing module, which can strike a balance between accuracy and efficiency. When deployed upon ResNet models, our method achieves higher performance and effectively reduces average computational cost on four widely used datasets.
Zequn Qin, Xi Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Unified curiosity-Driven learning with smoothed intrinsic reward estimation
Fuxian Huang, Weichao Li 0003, Jiabao Cui, Yongjian Fu 0002, Xi Li 0001
Pattern Recognit.5
2022 Memory-efficient distribution-guided experience sampling for policy consolidation
Fuxian Huang, Weichao Li 0003, Yining Lin, Naye Ji, Shijian Li, Xi Li 0001
Pattern Recognit. Lett.6
2022 Reparameterized attention for convolutional neural networks
Yiming Wu 0006, Yunlong Yu 0001, Xi Li 0001
Pattern Recognit. Lett.4
2022 Forgery-Domain-Supervised Deepfake Detection With Non-Negative Constraint
abstract
Fake faces produced by deepfake techniques have attracted public concerns in recent years. Deepfake detection is a binary classification task that distinguishes fake faces from real ones. As the training data for deepfake detection is usually generated from real faces via various face forgery methods, it is difficult for a single binary decision boundary to distinguish fake faces. Besides that, the learned features often involve irrelevant information for identifying fake faces that are generated from diverse forgery methods. To deal with such challenges, unlike existing approaches that regard fake detection as a binary classification, we re-model the task as a multiclass forgery-domain classification task, where each forgery method is treated as a distinct class. This simplifies the complex decision boundary brought by the diversity of forgery patterns and provides more forgery-relevant information for the learning process. In addition, we introduce a non-negative constrained learning framework composed of non-negative features and a non-negative constrained classifier (NCC) to block irrelevant features with zero weights and enhance forgery-relevant features with positive weights, leading to a sparse structure of the classifier. Furthermore, to capture subtle and discriminative forgery-relevant features, we propose an integration module over augmented faces based on cross-attention. We demonstrate that our approach achieves competitive performance and generalization ability on widely-used benchmarks through extensive experiments.
Yike Yuan, Xinghe Fu, Gaoang Wang, Xi Li 0001
IEEE Signal Process. Lett.5
2022 Progressive Multistage Learning for Discriminative Tracking
abstract
Visual tracking is typically solved as a discriminative learning problem that usually requires high-quality samples for online model adaptation. It is a critical and challenging problem to evaluate the training samples collected from previous predictions and employ sample selection by their quality to train the model. To tackle the above problem, we propose a joint discriminative learning scheme with the progressive multistage optimization policy of sample selection for robust visual tracking. The proposed scheme presents a novel time-weighted and detection-guided self-paced learning strategy for easy-to-hard sample selection, which is capable of tolerating relatively large intraclass variations while maintaining interclass separability. Such a self-paced learning strategy is jointly optimized in conjunction with the discriminative tracking process, resulting in robust tracking results. Experiments on the benchmark datasets demonstrate the effectiveness of the proposed learning framework.
Weichao Li 0003, Xi Li 0001, Omar El Farouk Bourahla, Fuxian Huang, Fei Wu 0001, Wei Liu 0005, Hongmin Liu 0001
IEEE Trans. Cybern.2
2022 Bias-Eliminated Semantic Refinement for Any-Shot Learning
abstract
When training samples are scarce, the semantic embedding technique, i. e., describing class labels with attributes, provides a condition to generate visual features for unseen objects by transferring the knowledge from seen objects. However, semantic descriptions are usually obtained in an external paradigm, such as manual annotation, resulting in weak consistency between descriptions and visual features. In this paper, we refine the coarse-grained semantic description for any-shot learning tasks, i. e., zero-shot learning (ZSL), generalized zero-shot learning (GZSL), and few-shot learning (FSL). A new model, namely, the semantic refinement Wasserstein generative adversarial network (SRWGAN) model, is designed with the proposed multihead representation and hierarchical alignment techniques. Unlike conventional methods, semantic refinement is performed with the aim of identifying a bias-eliminated condition for disjoint-class feature generation and is applicable in both inductive and transductive settings. We extensively evaluate model performance on six benchmark datasets and observe state-of-the-art results for any-shot learning; e. g., we obtain 70.2% harmonic accuracy for the Caltech UCSD Birds (CUB) dataset and 82.2% harmonic accuracy for the Oxford Flowers (FLO) dataset in the standard GZSL setting. Various visualizations are also provided to show the bias-eliminated generation of SRWGAN. Our code is available. 1.
Liangjun Feng, Chunhui Zhao 0001, Xi Li 0001
IEEE Trans. Image Process.3
2022 Global Context Assisted Structure-Aware Vehicle Retrieval
abstract
In vehicle retrieval, the vehicle patch should first be localized to remove the irrelevant background information. Moreover, the negative samples are much more prevalent than the positive samples, and the information from the negative samples is not fully exploited in the triple loss. What we need is a way to incorporate global knowledge and structure information to address these two issues. Therefore, we introduce a local-global context network for landmark alignment to update the predicted results by using the semantic information and the local compatibility and propose a structure-aware quadruple loss to use multiple and diverse negative samples in retrieval. Experiments on the VehicleID and the ENJOYOR vehicle retrieval datasets demonstrate that our approach obtains accuracy comparable to state-of-the-art approaches in vehicle retrieval.
Guohua Cheng, Shihao Yu, Xi Li 0001, Jianyuan Li, Bailin Yang
IEEE Trans. Intell. Transp. Syst.5
2022 Memory-Efficient Class-Incremental Learning for Image Classification
abstract
With the memory-resource-limited constraints, class-incremental learning (CIL) usually suffers from the "catastrophic forgetting" problem when updating the joint classification model on the arrival of newly added classes. To cope with the forgetting problem, many CIL methods transfer the knowledge of old classes by preserving some exemplar samples into the size-constrained memory buffer. To utilize the memory buffer more efficiently, we propose to keep more auxiliary low-fidelity exemplar samples, rather than the original real-high-fidelity exemplar samples. Such a memory-efficient exemplar preserving scheme makes the old-class knowledge transfer more effective. However, the low-fidelity exemplar samples are often distributed in a different domain away from that of the original exemplar samples, that is, a domain shift. To alleviate this problem, we propose a duplet learning scheme that seeks to construct domain-compatible feature extractors and classifiers, which greatly narrows down the above domain gap. As a result, these low-fidelity auxiliary exemplar samples have the ability to moderately replace the original exemplar samples with a lower memory cost. In addition, we present a robust classifier adaptation scheme, which further refines the biased classifier (learned with the samples containing distillation label knowledge about old classes) with the help of the samples of pure true class labels. Experimental results demonstrate the effectiveness of this work against the state-of-the-art approaches. We will release the code, baselines, and training statistics for all models to facilitate future research.
Hanbin Zhao, Hui Wang 0107, Yongjian Fu 0002, Fei Wu 0001, Xi Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2022 What and Where: Learn to Plug Adapters via NAS for Multidomain Learning
abstract
As an important and challenging problem, multidomain learning (MDL) typically seeks a set of effective lightweight domain-specific adapter modules plugged into a common domain-agnostic network. Usually, existing ways of adapter plugging and structure design are handcrafted and fixed for all domains before model learning, resulting in learning inflexibility and computational intensiveness. With this motivation, we propose to learn a data-driven adapter plugging strategy with neural architecture search (NAS), which automatically determines where to plug for those adapter modules. Furthermore, we propose an NAS-adapter module for adapter structure design in an NAS-driven learning scheme, which automatically discovers effective adapter module structures for different domains. Experimental results demonstrate the effectiveness of our MDL model against existing approaches under the conditions of comparable performance.
Hanbin Zhao, Yongjian Fu 0002, Hui Wang 0107, Omar El Farouk Bourahla, Xi Li 0001
IEEE Trans. Neural Networks Learn. Syst.7
2021 Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal Fusion
abstract
RGB-D salient object detection (SOD) is usually formulated as a problem of classification or regression over two modalities, i.e., RGB and depth. Hence, effective RGB-D feature modeling and multi-modal feature fusion both play a vital role in RGB-D SOD. In this paper, we propose a depth-sensitive RGB feature modeling scheme using the depth-wise geometric prior of salient objects. In principle, the feature modeling scheme is carried out in a depth-sensitive attention module, which leads to the RGB feature enhancement as well as the background distraction reduction by capturing the depth geometry prior. More-over, to perform effective multi-modal feature fusion, we further present an automatic architecture search approach for RGB-D SOD, which does well in finding out a feasible architecture from our specially designed multi-modal multi-scale search space. Extensive experiments on seven standard benchmarks demonstrate the effectiveness of the proposed approach against the state-of-the-art.
Peng Sun 0011, Wenhu Zhang, Xi Li 0001
CVPR5
2021 FcaNet: Frequency Channel Attention Networks
abstract
Attention mechanism, especially channel attention, has gained great success in the computer vision field. Many works focus on how to design efficient channel attention mechanisms while ignoring a fundamental problem, i.e., channel attention mechanism uses scalar to represent channel, which is difficult due to massive information loss. In this work, we start from a different view and regard the channel representation problem as a compression process using frequency analysis. Based on the frequency analysis, we mathematically prove that the conventional global average pooling is a special case of the feature decomposition in the frequency domain. With the proof, we naturally generalize the compression of the channel attention mechanism in the frequency domain and propose our method with multi-spectral channel attention, termed as FcaNet. FcaNet is simple but effective. We can change a few lines of code in the calculation to implement our method within existing channel attention methods. Moreover, the proposed method achieves state-of-the-art results compared with other channel attention methods on image classification, object detection, and instance segmentation tasks. Our method could consistently outperform the baseline SENet, with the same number of parameters and the same computational cost. Our code and models are publicly available at https://github.com/cfzd/FcaNet.
Zequn Qin, Fei Wu 0001, Xi Li 0001
ICCV4
2021 RDI-Net: Relational Dynamic Inference Networks
abstract
Dynamic inference networks, aimed at promoting computational efficiency, go along an adaptive executing path for a given sample. Prevalent methods typically assign a router for each convolutional block and sequentially make block-by-block executing decisions, without considering the relations during the dynamic inference. In this paper, we model the relations for dynamic inference from two aspects: the routers and the samples. We design a novel type of router called the relational router to model the relations among routers for a given sample. In principle, the current relational router aggregates the contextual features of preceding routers by graph convolution and propagates its router features to subsequent ones, making the executing decision for the current block in a long-range manner. Furthermore, we model the relation between samples by introducing a Sample Relation Module (SRM), encouraging correlated samples to go along correlated executing paths. As a whole, we call our method the Relational Dynamic Inference Network (RDI-Net). Extensive experiments on CIFAR-10/100 and ImageNet show that RDI-Net achieves state-of-the-art performance and computational cost reduction.
Shihao Su, Zequn Qin, Xi Li 0001
ICCV5
2021 MGH: Metadata Guided Hypergraph Modeling for Unsupervised Person Re-identification
abstract
As a challenging task, unsupervised person ReID aims to match the same identity with query images which does not require any labeled information. In general, most existing approaches focus on the visual cues only, leaving potentially valuable auxiliary metadata information (e.g., spatio-temporal context) unexplored. In the real world, such metadata is normally available alongside captured images, and thus plays an important role in separating several hard ReID matches. With this motivation in mind, we propose MGH, a novel unsupervised person ReID approach that uses meta information to construct a hypergraph for feature learning and label refinement. In principle, the hypergraph is composed of camera-topology-aware hyperedges, which can model the heterogeneous data correlations across cameras. Taking advantage of label propagation on the hypergraph, the proposed approach is able to effectively refine the ReID results, such as correcting the wrong labels or smoothing the noisy labels. Given the refined results, we further present a memory-based listwise loss to directly optimize the average precision in an approximate manner. Extensive experiments on three benchmarks demonstrate the effectiveness of the proposed approach against the state-of-the-art.
Yiming Wu 0005, Xintian Wu, Xi Li 0001
ACM Multimedia3
2021 When Video Classification Meets Incremental Classes
abstract
With the rapid development of social media, tremendous videos with new classes are generated daily, which raise an urgent demand for video classification methods that can continuously update new classes while maintaining the knowledge of old videos with limited storage and computing resources. In this paper, we summarize this task as Class-Incremental Video Classification (CIVC) and propose a novel framework to address it. As a subarea of incremental learning tasks, the challenge of catastrophic forgetting is unavoidable in CIVC. To better alleviate it, we utilize some characteristics of videos. First, we decompose the spatio-temporal knowledge before distillation rather than treating it as a whole in the knowledge transfer process; trajectory is also used to refine the decomposition. Second, we propose a dual granularity exemplar selection method to select and store representative video instances of old classes and key-frames inside videos under a tight storage budget. We benchmark our method and previous SOTA class-incremental learning methods on Something-Something V2 and Kinetics datasets, and our method outperforms previous methods significantly.
Hanbin Zhao, Shihao Su, Yongjian Fu 0002, Zibo Lin, Xi Li 0001
ACM Multimedia6
2021 Event prediction based on evolutionary event ontology knowledge
Qianren Mao, Xi Li 0001, Hao Peng 0001, Jianxin Li 0002, Dongxiao He
Future Gener. Comput. Syst.2
2021 Real-Time Semantic Segmentation via Auto Depth, Downsampling Joint Decision and Feature Aggregation
Peng Sun 0011, Jiaxiang Wu 0001, Peiwen Lin, Junzhou Huang, Xi Li 0001
Int. J. Comput. Vis.6
2021 Anytime Recognition with Routing Convolutional Networks
abstract
Achieving an automatic trade-off between accuracy and efficiency for a single deep neural network is highly desired in time-sensitive computer vision applications. To achieve anytime prediction, existing methods only embed fixed exits to neural networks and make the predictions with the fixed exits for all the samples (refer to the "latest-all" strategy). However, it is observed that the latest exit within a time budget does not always provide a more accurate prediction than the earlier exits for testing samples of various difficulties, making the "latest-all" strategy a sub-optimal solution. Motivated by this, we propose to improve the anytime prediction accuracy by allowing each sample to adaptively select its own optimal exit within a specific time budget. Specifically, we propose a new Routing Convolutional Network (RCN). For any given time budget, it adaptively selects the optimal layer as exit for a specific testing sample. To learn an optimal policy for sample routing, a Q-network is embedded into the RCN at each exit, considering both potential information gain and time-cost. To further boost the anytime prediction accuracy, the exits and the Q-networks are optimized alternately to mutually boost each other under the cost-sensitive environment. Apart from applying to whole image classification, RCN can also be adapted to dense prediction tasks, e.g., scene parsing, to achieve the pixel-level anytime prediction. Extensive experimental results on CIFAR-10, CIFAR-100, and ImageNet classification benchmarks, and Cityscapes scene parsing benchmark demonstrate the efficacy of the proposed RCN for anytime recognition.
Zequn Jie, Peng Sun 0011, Xi Li 0001, Jiashi Feng, Wei Liu 0005
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Progressive Class-Based Expansion Learning for Image Classification
abstract
In this paper, we propose a novel image process scheme called class-based expansion learning for image classification, which aims at improving the supervision-stimulation frequency for the samples of the confusing classes. Class-based expansion learning takes a bottom-up growing strategy in a class-based expansion optimization fashion, which pays more attention to the quality of learning the fine-grained classification boundaries for the preferentially selected classes. Besides, we develop a class confusion criterion to select the confusing class preferentially for training. In this way, the classification boundaries of the confusing classes are frequently stimulated, resulting in a fine-grained form. Experimental results demonstrate the effectiveness of the proposed scheme on several benchmarks.
Hui Wang 0107, Hanbin Zhao, Xi Li 0001
IEEE Signal Process. Lett.3
2021 Efficient Person Search via Expert-Guided Knowledge Distillation
abstract
The person search problem aims to find the target person in the scene images, which presents high demands for both effectiveness and efficiency. In this paper, we present a unified person search framework which jointly handles the two demands for real-world applications. We explore the technique of knowledge distillation (KD), which allows the student network to share capabilities of the deep expert networks with much fewer parameters and less computing time. To achieve this, we describe an efficient person search network and a set of deep and well-engineered expert networks, to build a tiny and compact model that can approximate the representations of the expert networks in a multitask learning manner. We present extensive experiments on three customized student networks with different scales of networks and show strong performance compared to the state-of-the-art methods on both mean average precision and top-1 accuracies. We further demonstrate the efficiency of the proposed network at 120 frames/s in the feedforward time with only a little sacrifice on the accuracy.
Xi Li 0001, Zhongfei Zhang
IEEE Trans. Cybern.2
2021 Multitask Identity-Aware Image Steganography via Minimax Optimization
abstract
High-capacity image steganography, aimed at concealing a secret image in a cover image, is a technique to preserve sensitive data, e.g., faces and fingerprints. Previous methods focus on the security during transmission and subsequently run a risk of privacy leakage after the restoration of secret images at the receiving end. To address this issue, we propose a framework, called Multitask Identity-Aware Image Steganography (MIAIS), to achieve direct recognition on container images without restoring secret images. The key issue of the direct recognition is to preserve identity information of secret images into container images and make container images look similar to cover images at the same time. Thus, we introduce a simple content loss to preserve the identity information, and design a minimax optimization to deal with the contradictory aspects. We demonstrate that the robustness results can be transferred across different cover datasets. In order to be flexible for the secret image restoration in some cases, we incorporate an optional restoration network into our method, providing a multitask framework. The experiments under the multitask scenario show the effectiveness of our framework compared with other visual information hiding methods and state-of-the-art high-capacity image steganography methods. The code is available at https://github.com/jiabaocui/MIAIS.
Jiabao Cui, Liangli Zheng, Cuizhu Bao, Jupeng Xia, Xi Li 0001
IEEE Trans. Image Process.7
2021 ResKD: Residual-Guided Knowledge Distillation
abstract
Knowledge distillation, aimed at transferring the knowledge from a heavy teacher network to a lightweight student network, has emerged as a promising technique for compressing neural networks. However, due to the capacity gap between the heavy teacher and the lightweight student, there still exists a significant performance gap between them. In this article, we see knowledge distillation in a fresh light, using the knowledge gap, or the residual, between a teacher and a student as guidance to train a much more lightweight student, called a res-student. We combine the student and the res-student into a new student, where the res-student rectifies the errors of the former student. Such a residual-guided process can be repeated until the user strikes the balance between accuracy and cost. At inference time, we propose a sample-adaptive strategy to decide which res-students are not necessary for each sample, which can save computational cost. Experimental results show that we achieve competitive performance with 18.04%, 23.14%, 53.59%, and 56.86% of the teachers' computational costs on the CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet datasets. Finally, we do thorough theoretical and empirical analysis for our method.
Xuewei Li 0003, Omar El Farouk Bourahla, Fei Wu 0001, Xi Li 0001
IEEE Trans. Image Process.5
2021 Multitask Non-Autoregressive Model for Human Motion Prediction
abstract
Human motion prediction, which aims at predicting future human skeletons given the past ones, is a typical sequence-to-sequence problem. Therefore, extensive efforts have been devoted to exploring different RNN-based encoder-decoder architectures. However, by generating target poses conditioned on the previously generated ones, these models are prone to bringing issues such as error accumulation problem. In this paper, we argue that such issue is mainly caused by adopting autoregressive manner. Hence, a novel Non-AuToregressive model (NAT) is proposed with a complete non-autoregressive decoding scheme, as well as a context encoder and a positional encoding module. More specifically, the context encoder embeds the given poses from temporal and spatial perspectives. The frame decoder is responsible for predicting each future pose independently. The positional encoding module injects positional signal into the model to indicate the temporal order. Besides, a multitask training paradigm is presented for both low-level human skeleton prediction and high-level human action recognition, resulting in the considerable improvement for the prediction task. Our approach is evaluated on Human3.6M and CMU-Mocap benchmarks and outperforms state-of-the-art autoregressive methods.
Bin Li 0038, Zhongfei Zhang, Hailin Feng, Xi Li 0001
IEEE Trans. Image Process.5
2021 Efficient Style-Corpus Constrained Learning for Photorealistic Style Transfer
abstract
Photorealistic style transfer is a challenging task, which demands the stylized image remains real. Existing methods are still suffering from unrealistic artifacts and heavy computational cost. In this paper, we propose a novel Style-Corpus Constrained Learning (SCCL) scheme to address these issues. The style-corpus with the style-specific and style-agnostic characteristics simultaneously is proposed to constrain the stylized image with the style consistency among different samples, which improves photorealism of stylization output. By using adversarial distillation learning strategy, a simple fast-to-execute network is trained to substitute previous complex feature transforms models, which reduces the computational cost significantly. Experiments demonstrate that our method produces rich-detailed photorealistic images, with 13 ~ 50 times faster than the state-of-the-art method (WCT2).
Yingxu Qiao, Jiabao Cui, Fuxian Huang, Hongmin Liu 0001, Cuizhu Bao, Xi Li 0001
IEEE Trans. Image Process.6
2021 Condition-Aware Comparison Scheme for Gait Recognition
abstract
As an important and challenging problem, gait recognition has gained considerable attention. It suffers from confounding conditions, that is, it is sensitive to camera views, dressing types and so on. Interestingly, it is observed that, under different conditions, local body parts contribute differently to recognition performance. In this paper, we propose a condition-aware comparison scheme to measure gait pairs' similarity via a novel module named Instructor. Also, we present a geometry-guided data augmentation approach (Dresser) to enrich dressing conditions. Furthermore, to enhance the gait representation, we propose to model temporal local information from coarse to fine. Our model is evaluated on two popular benchmarks, CASIA-B and OULP. Results show that our method outperforms current state-of-the-art methods, especially in the cross-condition scenario.
Haoqian Wu, Yongjian Fu 0002, Bin Li 0038, Xi Li 0001
IEEE Trans. Image Process.5
2021 F³A-GAN: Facial Flow for Face Animation With Generative Adversarial Networks
abstract
Formulated as a conditional generation problem, face animation aims at synthesizing continuous face images from a single source image driven by a set of conditional face motion. Previous works mainly model the face motion as conditions with 1D or 2D representation (e.g., action units, emotion codes, landmark), which often leads to low-quality results in some complicated scenarios such as continuous generation and large-pose transformation. To tackle this problem, the conditions are supposed to meet two requirements, i.e., motion information preserving and geometric continuity. To this end, we propose a novel representation based on a 3D geometric flow, termed facial flow, to represent the natural motion of the human face at any pose. Compared with other previous conditions, the proposed facial flow well controls the continuous changes to the face. After that, in order to utilize the facial flow for face editing, we build a synthesis framework generating continuous images with conditional facial flows. To fully take advantage of the motion information of facial flows, a hierarchical conditional framework is designed to combine the extracted multi-scale appearance features from images and motion features from flows in a hierarchical manner. The framework then decodes multiple fused features back to images progressively. Experimental results demonstrate the effectiveness of our method compared to other state-of-the-art methods.
Xintian Wu, Qihang Zhang, Yiming Wu 0005, Lingyun Sun, Xi Li 0001
IEEE Trans. Image Process.7
2021 Deep Attentive Video Summarization With Distribution Consistency Learning
abstract
This article studies supervised video summarization by formulating it into a sequence-to-sequence learning framework, in which the input and output are sequences of original video frames and their predicted importance scores, respectively. Two critical issues are addressed in this article: short-term contextual attention insufficiency and distribution inconsistency. The former lies in the insufficiency of capturing the short-term contextual attention information within the video sequence itself since the existing approaches focus a lot on the long-term encoder-decoder attention. The latter refers to the distributions of predicted importance score sequence and the ground-truth sequence is inconsistent, which may lead to a suboptimal solution. To better mitigate the first issue, we incorporate a self-attention mechanism in the encoder to highlight the important keyframes in a short-term context. The proposed approach alongside the encoder-decoder attention constitutes our deep attentive models for video summarization. For the second one, we propose a distribution consistency learning method by employing a simple yet effective regularization loss term, which seeks a consistent distribution for the two sequences. Our final approach is dubbed as Attentive and Distribution consistent video Summarization (ADSum). Extensive experiments on benchmark data sets demonstrate the superiority of the proposed ADSum approach against state-of-the-art approaches.
Zhong Ji, Yanwei Pang, Xi Li 0001, Jungong Han
IEEE Trans. Neural Networks Learn. Syst.4
2021 End-to-End Video Saliency Detection via a Deep Contextual Spatiotemporal Network
abstract
As an interesting and important problem in computer vision, learning-based video saliency detection aims to discover the visually interesting regions in a video sequence. Capturing the information within frame and between frame at different aspects (such as spatial contexts, motion information, temporal consistency across frames, and multiscale representation) is important for this task. A key issue is how to jointly model all these factors within a unified data-driven scheme in an end-to-end fashion. In this article, we propose an end-to-end spatiotemporal deep video saliency detection approach, which captures the information on spatial contexts and motion characteristics. Furthermore, it encodes the temporal consistency information across the consecutive frames by implementing a convolutional long short-term memory (Conv-LSTM) model. In addition, the multiscale saliency properties for each frame are adaptively integrated for final saliency prediction in a collaborative feature-pyramid way. Finally, the proposed deep learning approach unifies all the aforementioned parts into an end-to-end joint deep learning scheme. Experimental results demonstrate the effectiveness of our approach in comparison with the state-of-the-art approaches.
Lina Wei, Shanshan Zhao 0001, Omar El Farouk Bourahla, Xi Li 0001, Fei Wu 0001, Yueting Zhuang, Junwei Han 0001, Mingliang Xu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2020 BANet: Bidirectional Aggregation Network With Occlusion Handling for Panoptic Segmentation
abstract
Panoptic segmentation aims to perform instance segmentation for foreground instances and semantic segmentation for background stuff simultaneously. The typical top-down pipeline concentrates on two key issues: 1) how to effectively model the intrinsic interaction between semantic segmentation and instance segmentation, and 2) how to properly handle occlusion for panoptic segmentation. Intuitively, the complementarity between semantic segmentation and instance segmentation can be leveraged to improve the performance. Besides, we notice that using detection/mask scores is insufficient for resolving the occlusion problem. Motivated by these observations, we propose a novel deep panoptic segmentation scheme based on a bidirectional learning pipeline. Moreover, we introduce a plug-and-play occlusion handling algorithm to deal with the occlusion between different object instances. The experimental results on COCO panoptic benchmark validate the effectiveness of our proposed method. Codes will be released soon at https://github.com/Mooonside/BANet.
Guangchen Lin, Omar El Farouk Bourahla, Yiming Wu 0005, Junyi Feng, Mingliang Xu 0001, Xi Li 0001
CVPR9
2020 Graph-Guided Architecture Search for Real-Time Semantic Segmentation
abstract
Designing a lightweight semantic segmentation network often requires researchers to find a trade-off between performance and speed, which is always empirical due to the limited interpretability of neural networks. In order to release researchers from these tedious mechanical trials, we propose a Graph-guided Architecture Search (GAS) pipeline to automatically search real-time semantic segmentation networks. Unlike previous works that use a simplified search space and stack a repeatable cell to form a network, we introduce a novel search mechanism with a new search space where a lightweight model can be effectively explored through the cell-level diversity and latency oriented constraint. Specifically, to produce the cell-level diversity, the cell-sharing constraint is eliminated through the cell-independent manner. Then a graph convolution network (GCN) is seamlessly integrated as a communication mechanism between cells. Finally, a latency-oriented constraint is endowed into the search process to balance the speed and performance. Extensive experiments on Cityscapes and CamVid datasets demonstrate that GAS achieves the new state-of-the-art trade-off between accuracy and speed. In particular, on Cityscapes dataset, GAS achieves the new best performance of 73.5% mIoU with the speed of 108.4 FPS on Titan Xp.
Peiwen Lin, Peng Sun 0011, Sirui Xie, Xi Li 0001, Jianping Shi
CVPR5
2020 Ultra Fast Structure-Aware Deep Lane Detection
Zequn Qin, Xi Li 0001
ECCV (24)3
2020 Multi-Way Multi-View Deep Autoencoder for Image Feature Learning with Multi-Level Graph Regularization
abstract
Multi-view feature learning has garnered much attention recently since many real world data are comprised of different representations or views. How to explore the consensus structure and eliminate the inconsistency noise in different views remains a challenging problem in multi-view feature learning. In this paper, we propose a multi-way deep autoencoder for multi-view feature learning to explore the deep consensus structure and reconcile the efficiency of encoding process meanwhile. Through a multi-way encoding process, we embed the original data feature views to nonnegative representations of multiple levels which are structured hierarchically. Along the structure of embedded representations, we recover the diversity and important information layer by layer in the decoding process. The experiments on two image datasets show the superior performance of our method.
Sen Zhou, Xi Li 0001, Haoqi Zhu
ICASSP3
2020 Stacked Pooling for Boosting Scale Invariance of Crowd Counting
abstract
In this work, we take insight into the dense crowd counting problem by exploring the phenomenon of cross-scale visual similarity caused by perspective distortions. It is a quite common phenomenon in crowd scenarios, suggesting the crowd counting model to enable a good performance of scale invariance. Existing deep crowd counting approaches mainly focus on the multi-scale techniques over convolutional layers to capture scale-adaptive features, resulting in high computing costs. In this paper, we propose simple but effective pooling variants, i.e., multi-kernel pooling and stacked pooling, to take place of the vanilla pooling layers in convolutional neural networks (CNNs) for boosting the scale invariance. Our proposed pooling modules do not introduce extra parameters and can be easily implemented in practice. Empirical studies on two benchmark crowd counting datasets show that the proposed pooling modules beat the vanilla pooling layer in most experimental cases.
Siyu Huang, Xi Li 0001, Zhi-Qi Cheng, Zhongfei Zhang, Alex Hauptmann 0001
ICASSP2
2020 TextRay: Contour-based Geometric Modeling for Arbitrary-shaped Scene Text Detection
abstract
Arbitrary-shaped text detection is a challenging task due to the complex geometric layouts of texts such as large aspect ratios, various scales, random rotations and curve shapes. Most state-of-the-art methods solve this problem from bottom-up perspectives, seeking to model a text instance of complex geometric layouts with simple local units (e.g., local boxes or pixels) and generate detections with heuristic post-processings. In this work, we propose an arbitrary-shaped text detection method, namely TextRay, which conducts top-down contour-based geometric modeling and geometric parameter learning within a single-shot anchor-free framework. The geometric modeling is carried out under polar system with a bidirectional mapping scheme between shape space and parameter space, encoding complex geometric layouts into unified representations. For effective learning of the representations, we design a central-weighted training strategy and a content loss which builds propagation paths between geometric encodings and visual content. TextRay outputs simple polygon detections at one pass with only one NMS post-processing. Experiments on several benchmark datasets demonstrate the effectiveness of the proposed approach. The code is available at https://github.com/LianaWang/TextRay.
Fei Wu 0001, Xi Li 0001
ACM Multimedia4
2020 Human-Centric Clothing Segmentation via Deformable Semantic Locality-Preserving Network
abstract
In the fields of computer vision and graphics, clothing segmentation is a challenging and practical task which is typically implemented in a fine-grained semantic segmentation framework. Unlike the generic semantic segmentation task, clothing segmentation has some domain-specific properties such as diverse appearance variations, non-rigid geometry deformations, and small sample learning. To deal with these points, we propose a semantic locality-preserving segmentation model, which adaptively attaches an original clothing image with a semantically similar (e.g., appearance or pose) auxiliary exemplar by search. Through considering the interactions of the clothing image and its exemplar, more intrinsic knowledge about the locality manifold structures of clothing images is discovered to make the learning process of small sample problem more stable and tractable. Besides, we present a CNN model based on the deformable convolutions to extract the non-rigid geometry-aware features for clothing images. Furthermore, we apply our semantic locality-preserving segmentation model in both image and video cases, resulting in favorable clothing segmentation performance. Experimental results demonstrate the effectiveness of the proposed model against the state-of-the-art approaches.
Wei Ji 0008, Xi Li 0001, Fei Wu 0001, Yueting Zhuang
IEEE Trans. Circuits Syst. Video Technol.2
2020 Context-Aware Deep Spatiotemporal Network for Hand Pose Estimation From Depth Images
abstract
As a fundamental and challenging problem in computer vision, hand pose estimation aims to estimate the hand joint locations from depth images. Typically, the problems are modeled as learning a mapping function from images to hand joint coordinates in a data-driven manner. In this paper, we propose a context-aware deep spatiotemporal network, a novel method to jointly model the spatiotemporal properties for hand pose estimation. Our proposed network is able to learn the representations of the spatial information and the temporal structure from the image sequences. Moreover, by adopting the adaptive fusion method, the model is capable of dynamically weighting different predictions to lay emphasis on sufficient context. Our method is examined on two common benchmarks, the experimental results demonstrate that our proposed approach achieves the best or the second-best performance with the state-of-the-art methods and runs in 60 fps.
Yiming Wu 0005, Wei Ji 0008, Xi Li 0001, Gang Wang 0012, Jianwei Yin, Fei Wu 0001
IEEE Trans. Cybern.3
2020 Semantic Neighborhood-Aware Deep Facial Expression Recognition
abstract
Different from many other attributes, facial expression can change in a continuous way, and therefore, a slight semantic change of input should also lead to the output fluctuation limited in a small scale. This consistency is important. However, current Facial Expression Recognition (FER) datasets may have the extreme imbalance problem, as well as the lack of data and the excessive amounts of noise, hindering this consistency and leading to a performance decreasing when testing. In this paper, we not only consider the prediction accuracy on sample points, but also take the neighborhood smoothness of them into consideration, focusing on the stability of the output with respect to slight semantic perturbations of the input. A novel method is proposed to formulate semantic perturbation and select unreliable samples during training, reducing the bad effect of them. Experiments show the effectiveness of the proposed method and state-of-the-art results are reported, getting closer to an upper limit than the state-of-the-art methods by a factor of 30% in AffectNet, the largest in-the-wild FER database by now.
Yongjian Fu 0002, Xintian Wu, Xi Li 0001, Daxin Luo
IEEE Trans. Image Process.3
2020 Context-Aware Graph Label Propagation Network for Saliency Detection
abstract
Recently, a large number of existing methods for saliency detection have mainly focused on designing complex network architectures to aggregate powerful features from backbone networks. However, contextual information is not well utilized, which often causes false background regions and blurred object boundaries. Motivated by these issues, we propose an easyto-implement module that utilizes the edge-preserving ability of superpixels and the graph neural network to interact the context of superpixel nodes. In more detail, we first extract the features from the backbone network and obtain the superpixel information of images. This step is followed by superpixel pooling in which we transfer the irregular superpixel information to a structured feature representation. To propagate the information among the foreground and background regions, we use a graph neural network and self-attention layer to better evaluate the degree of saliency degree. Additionally, an affinity loss is proposed to regularize the affinity matrix to constrain the propagation path. Moreover, we extend our module to a multiscale structure with different numbers of superpixels. Experiments on five challenging datasets show that our approach can improve the performance of three baseline methods in terms of some popular evaluation metrics.
Wei Ji 0008, Xi Li 0001, Lina Wei, Fei Wu 0001, Yueting Zhuang
IEEE Trans. Image Process.2
2020 Adaptive Graph Representation Learning for Video Person Re-Identification
abstract
Recent years have witnessed the remarkable progress of applying deep learning models in video person re-identification (Re-ID). A key factor for video person Re-ID is to effectively construct discriminative and robust video feature representations for many complicated situations. Part-based approaches employ spatial and temporal attention to extract representative local features. While correlations between parts are ignored in the previous methods, to leverage the relations of different parts, we propose an innovative adaptive graph representation learning scheme for video person Re-ID, which enables the contextual interactions between relevant regional features. Specifically, we exploit the pose alignment connection and the feature affinity connection to construct an adaptive structure-aware adjacency graph, which models the intrinsic relations between graph nodes. We perform feature propagation on the adjacency graph to refine regional features iteratively, and the neighbor nodes' information is taken into account for part feature representation. To learn compact and discriminative representations, we further propose a novel temporal resolution-aware regularization, which enforces the consistency among different temporal resolutions for the same identities. We conduct extensive evaluations on four benchmarks, i.e. iLIDS-VID, PRID2011, MARS, and DukeMTMC-VideoReID, experimental results achieve the competitive performance which demonstrates the effectiveness of our proposed method. Code is available at https://github.com/weleen/AGRL.pytorch.
Yiming Wu 0005, Omar El Farouk Bourahla, Xi Li 0001, Fei Wu 0001, Qi Tian 0001
IEEE Trans. Image Process.3
2020 Learning Multi-Level Density Maps for Crowd Counting
abstract
People in crowd scenes often exhibit the characteristic of imbalanced distribution. On the one hand, people size varies largely due to the camera perspective. People far away from the camera look smaller and are likely to occlude each other, whereas people near to the camera look larger and are relatively sparse. On the other hand, the number of people also varies greatly in the same or different scenes. This article aims to develop a novel model that can accurately estimate the crowd count from a given scene with imbalanced people distribution. To this end, we have proposed an effective multi-level convolutional neural network (MLCNN) architecture that first adaptively learns multi-level density maps and then fuses them to predict the final output. Density map of each level focuses on dealing with people of certain sizes. As a result, the fusion of multi-level density maps is able to tackle the large variation in people size. In addition, we introduce a new loss function named balanced loss (BL) to impose relatively BL feedback during training, which helps further improve the performance of the proposed network. Furthermore, we introduce a new data set including 1111 images with a total of 49 061 head annotations. MLCNN is easy to train with only one end-to-end training stage. Experimental results demonstrate that our MLCNN achieves state-of-the-art performance. In particular, our MLCNN reaches a mean absolute error (MAE) of 242.4 on the UCF_CC_50 data set, which is 37.2 lower than the second-best result.
Xiaoheng Jiang, Li Zhang 0072, Pei Lv, Yibo Guo, Ruijie Zhu 0001, Yanwei Pang, Xi Li 0001, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Neural Networks Learn. Syst.8
2019 Spatio-Temporal Graph Routing for Skeleton-Based Action Recognition
abstract
With the representation effectiveness, skeleton-based human action recognition has received considerable research attention, and has a wide range of real applications. In this area, many existing methods typically rely on fixed physicalconnectivity skeleton structure for recognition, which is incapable of well capturing the intrinsic high-order correlations among skeleton joints. In this paper, we propose a novel spatio-temporal graph routing (STGR) scheme for skeletonbased action recognition, which adaptively learns the intrinsic high-order connectivity relationships for physicallyapart skeleton joints. Specifically, the scheme is composed of two components: spatial graph router (SGR) and temporal graph router (TGR). The SGR aims to discover the connectivity relationships among the joints based on sub-group clustering along the spatial dimension, while the TGR explores the structural information by measuring the correlation degrees between temporal joint node trajectories. The proposed scheme is naturally and seamlessly incorporated into the framework of graph convolutional networks (GCNs) to produce a set of skeleton-joint-connectivity graphs, which are further fed into the classification networks. Moreover, an insightful analysis on receptive field of graph node is provided to explain the necessity of our method. Experimental results on two benchmark datasets (NTU-RGB+D and Kinetics) demonstrate the effectiveness against the state-of-the-art.
Bin Li 0038, Xi Li 0001, Zhongfei Zhang, Fei Wu 0001
AAAI2
2019 Learning a Key-Value Memory Co-Attention Matching Network for Person Re-Identification
abstract
Person re-identification (Re-ID) is typically cast as the problem of semantic representation and alignment, which requires precisely discovering and modeling the inherent spatial structure information on person images. Motivated by this observation, we propose a Key-Value Memory Matching Network (KVM-MN) model that consists of key-value memory representation and key-value co-attention matching. The proposed KVM-MN model is capable of building an effective local-position-aware person representation that encodes the spatial feature information in the form of multi-head key-value memory. Furthermore, the proposed KVM-MN model makes use of multi-head co-attention to automatically learn a number of cross-person-matching patterns, resulting in more robust and interpretable matching results. Finally, we build a setwise learning mechanism that implements a more generalized query-to-gallery-image-set learning procedure. Experimental results demonstrate the effectiveness of the proposed model against the state-of-the-art.
Xi Li 0001, Zhongfei Zhang
AAAI2
2019 GroundNet: Monocular Ground Plane Normal Estimation with Geometric Consistency
abstract
We focus on estimating the 3D orientation of the ground plane from a single image. We formulate the problem as an inter-mingled multi-task prediction problem by jointly optimizing for pixel-wise surface normal direction, ground plane segmentation, and depth estimates. Specifically, our proposed model, GroundNet, first estimates the depth and surface normal in two separate streams, from which two ground plane normals are then computed deterministically. To leverage the geometric correlation between depth and normal, we propose to add a consistency loss on top of the computed ground plane normals. In addition, a ground segmentation stream is used to isolate the ground regions so that we can selectively back-propagate parameter updates through only the ground regions in the image. Our method achieves the top-ranked performance on ground plane normal estimation and horizon line detection on the real-world outdoor datasets of ApolloScape and KITTI, improving the performance of previous art by up to 17.7% relatively.
Yunze Man, Xinshuo Weng, Xi Li 0001, Kris Makoto Kitani
ACM Multimedia3
2019 State Distribution-Aware Sampling for Deep Q-Learning
Weichao Li 0003, Fuxian Huang, Xi Li 0001, Gang Pan 0001, Fei Wu 0001
Neural Process. Lett.3
2019 Multi-Task Structure-Aware Context Modeling for Robust Keypoint-Based Object Tracking
abstract
In the fields of computer vision and graphics, keypoint-based object tracking is a fundamental and challenging problem, which is typically formulated in a spatio-temporal context modeling framework. However, many existing keypoint trackers are incapable of effectively modeling and balancing the following three aspects in a simultaneous manner: temporal model coherence across frames, spatial model consistency within frames, and discriminative feature construction. To address this problem, we propose a robust keypoint tracker based on spatio-temporal multi-task structured output optimization driven by discriminative metric learning. Consequently, temporal model coherence is characterized by multi-task structured keypoint model learning over several adjacent frames; spatial model consistency is modeled by solving a geometric verification based structured learning problem; discriminative feature construction is enabled by metric learning to ensure the intra-class compactness and inter-class separability. To achieve the goal of effective object tracking, we jointly optimize the above three modules in a spatio-temporal multi-task learning scheme. Furthermore, we incorporate this joint learning scheme into both single-object and multi-object tracking scenarios, resulting in robust tracking results. Experiments over several challenging datasets have justified the effectiveness of our single-object and multi-object trackers against the state-of-the-art.
Xi Li 0001, Wei Ji 0008, Yiming Wu 0005, Fei Wu 0001, Ming-Hsuan Yang 0001, Dacheng Tao, Ian D. Reid 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 A Bilinear Ranking SVM for Knowledge Based Relation Prediction and Classification
abstract
As an important and challenging problem, knowledge representation and inference are typically carried out in a knowledge embedding framework over a multi-relational knowledge graph, and thus have a wide range of applications such as semantic retrieval and question answering. In this paper, we propose a bilinear learning framework which performs cross-entity knowledge relation analysis in the continuous vector space (derived from knowledge embedding). In the framework, we effectively model the intrinsic correlations among different types of knowledge relations within a max-margin multi-relational ranking scheme, which jointly optimizes the tasks of entity embedding and cross-entity relation prediction in terms of multi-relational structures of the knowledge graph. Specifically, we devise a bilinear scoring function that aims to evaluate the confidence degree of semantic relation prediction for entity pairs through a multi-relational learning-to-rank pipeline. In essence, the pipeline formulates the problem of relation prediction for entity pairs as that of learning relation-specific ranking functions by max-margin optimization. Experimental results demonstrate the effectiveness of the proposed framework on two common benchmark datasets.
Shengkang Yu, Xi Li 0001, Xueyi Zhao, Zhongfei Zhang, Fei Wu 0001, Jingdong Wang 0001, Yueting Zhuang, Xuelong Li 0001
IEEE Trans. Big Data2
2019 User-Ranking Video Summarization With Multi-Stage Spatio-Temporal Representation
abstract
Video summarization is a challenging task, mainly due to the difficulties in learning complicated semantic structural relations between videos and summaries. In this paper, we present a novel supervised video summarization scheme based on threestage deep neural networks. The scheme takes a divide-andconquer strategy to resolve the complicated task of 3D video summarization into a set of easy and flexible computational subtasks, and then to sequentially perform 2D CNNs, 1D CNNs, and LSTM to address the subtasks in an hierarchical fashion. The hierarchical modeling of spatio-temporal structure leads to high performance and efficiency. In addition, we propose a simple but effective user-ranking method to cope with the labeling subjectivity problem of user-created video summarization, leading to the labeling quality refinement for robust supervised learning. Experimental results show that our approach outperforms the state-of-the-art video summarization methods on two benchmark datasets.
Siyu Huang, Xi Li 0001, Zhongfei Zhang, Fei Wu 0001, Junwei Han 0001
IEEE Trans. Image Process.2
2019 Deep Group-Wise Fully Convolutional Network for Co-Saliency Detection With Graph Propagation
abstract
A key problem in co-saliency detection is how to effectively model the interactive relationship of a whole image group and the individual perspective of each image in a united data-driven manner. In this paper, we propose a group-wise deep co-saliency detection approach to address the co-saliency object discovery problem based on the fully convolutional network (FCN). The proposed approach captures the group-wise interaction information for group images by learning a semantics-aware image representation based on a convolutional neural network, which adaptively learns the group-wise features for co-saliency detection. Furthermore, the proposed approach discovers the collaborative and interactive relationships between group-wise feature representation and single image individual feature representation, and model this in a collaborative learning framework. Then, we set up a unified deep learning scheme to jointly optimize the process of group-wise feature representation learning and the collaborative learning, leading to more reliable and robust co-saliency detection results. Finally, we present a graph Laplacian regularized nonlinear regression model for saliency refinement. Experimental results demonstrate the effectiveness of our approach in comparison with the state-of-the-art approaches.
Lina Wei, Shanshan Zhao 0001, Omar El Farouk Bourahla, Xi Li 0001, Fei Wu 0001, Yueting Zhuang
IEEE Trans. Image Process.4
2019 Deep Q Learning Driven CT Pancreas Segmentation With Geometry-Aware U-Net
abstract
The segmentation of pancreas is important for medical image analysis, yet it faces great challenges of class imbalance, background distractions, and non-rigid geometrical features. To address these difficulties, we introduce a deep Q network (DQN) driven approach with deformable U-Net to accurately segment the pancreas by explicitly interacting with contextual information and extract anisotropic features from pancreas. The DQN-based model learns a context-adaptive localization policy to produce a visually tightened and precise localization bounding box of the pancreas. Furthermore, deformable U-Net captures geometry-aware information of pancreas by learning geometrically deformable filters for feature extraction. The experiments on NIH dataset validate the effectiveness of the proposed framework in pancreas segmentation.
Yunze Man, Yangsibo Huang, Junyi Feng, Xi Li 0001, Fei Wu 0001
IEEE Trans. Medical Imaging4
2019 Stylized Aesthetic QR Code
abstract
With the continued proliferation of smart mobile devices, the Quick Response (QR) code has become one of the most-used types of two-dimensional code in the world. Aiming at beautifying the visual-unpleasant appearance of QR codes, existing works have developed a series of techniques. However, these works still leave much to be desired, such as personalization, artistry, and robustness. To address these issues, in this paper, we propose a novel type of aesthetic QR codes, Stylized aEsthEtic (SEE) QR code , and a three-stage approach to automatically produce such robust style-oriented codes. Specifically, in the first stage, we propose a method to generate an optimized baseline aesthetic QR code, which reduces the visual contrast between the noise-like black/white modules and the blended image. In the second stage, to obtain an art style QR code, we tailor an appropriate neural style transformation network to endow the baseline aesthetic QR code with artistic elements. In the third stage, we design a module-based robustness-optimization mechanism to ensure the performance robust by balancing two competing terms: visual quality and readability. Extensive experiments demonstrate that the SEE QR code has high quality in terms of both visual appearance and robustness and also offers a greater variety of personalized choices to users.
Mingliang Xu 0001, Hao Su 0001, Xi Li 0001, Jing Liao 0001, Jianwei Niu 0002, Pei Lv, Bing Zhou 0003
IEEE Trans. Multim.4
2019 Editorial: Booming of Neural Networks and Learning Systems
abstract
As you open this January issue of the IEEE Transactions on Neural Networks and Learning Systems (TNNLS), I hope everyone enjoyed a great holiday season and is excited for the new year of 2019. I am very delighted and honored to report several key metrics of IEEE TNNLS to the community.
Akira Hirose 0001, Alessio Micheli, Artur S. d'Avila Garcez, Choon Ki Ahn, Gang Pan 0001, Hamid Reza Karimi, Jianbing Shen, José de Jesús Rubio, Lei Zhang 0005, Lingjia Liu 0001, Lorenzo Livi, Nishchal K. Verma, Pedro Antonio Gutiérrez, Qi Tian 0001, Qinglai Wei, Seiichi Ozawa, Stuart Harvey Rubin, Weineng Chen, Xi Li 0001, Xiaofeng Liao 0001, Youmin Zhang 0001, Zhen Ni, Haibo He
IEEE Trans. Neural Networks Learn. Syst.20
2018 FR-ANet: A Face Recognition Guided Facial Attribute Classification Network
abstract
In this paper, we study the problem of facial attribute learning. In particular, we propose a Face Recognition guided facial Attribute classification Network, called FR-ANet. All the attributes share low-level features, while high-level features are specially learned for attribute groups. Further, to utilize the identity information, high-level features are merged to perform face identity recognition. The experimental results on CelebA and LFWA datasets demonstrate the promise of the FR-ANet.
Jiajiong Cao, Yingming Li, Xi Li 0001, Zhongfei Zhang
AAAI3
2018 Multi-Channel Pyramid Person Matching Network for Person Re-Identification
abstract
In this work, we present a Multi-Channel deep convolutional Pyramid Person Matching Network (MC-PPMN) based on the combination of the semantic-components and the color-texture distributions to address the problem of person re-identification. In particular, we learn separate deep representations for semantic-components and color-texture distributions from two person images and then employ pyramid person matching network (PPMN) to obtain correspondence representations. These correspondence representations are fused to perform the re-identification task. Further, the proposed framework is optimized via a unified end-to-end deep learning scheme. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our approach against the state-of-the-art literature, especially on the rank-1 recognition rate.
Chaojie Mao, Yingming Li, Zhongfei Zhang, Xi Li 0001
AAAI5
2018 Geometry-Aware Scene Text Detection With Instance Transformation Network
abstract
Localizing text in the wild is challenging in the situations of complicated geometric layout of the targets like random orientation and large aspect ratio. In this paper, we propose a geometry-aware modeling approach tailored for scene text representation with an end-to-end learning scheme. In our approach, a novel Instance Transformation Network (ITN) is presented to learn the geometry-aware representation encoding the unique geometric configurations of scene text instances with in-network transformation embedding, resulting in a robust and elegant framework to detect words or text lines at one pass. An end-to-end multi-task learning strategy with transformation regression, text/non-text classification and coordinates regression is adopted in the ITN. Experiments on the benchmark datasets demonstrate the effectiveness of the proposed approach in detecting scene text in various geometric configurations.
Xi Li 0001, Xinchao Wang, Dacheng Tao
CVPR3
2018 Knowledge-Guided Agent-Tactic-Aware Learning for StarCraft Micromanagement
abstract
As an important and challenging problem in artificial intelligence (AI) game playing, StarCraft micromanagement involves a dynamically adversarial game playing process with complex multi-agent control within a large action space. In this paper, we propose a novel knowledge-guided agent-tactic-aware learning scheme, that is, opponent-guided tactic learning (OGTL), to cope with this micromanagement problem. In principle, the proposed scheme takes a two-stage cascaded learning strategy which is capable of not only transferring the human tactic knowledge from the human-made opponent agents to our AI agents but also improving the adversarial ability. With the power of reinforcement learning, such a knowledge-guided agent-tactic-aware scheme has the ability to guide the AI agents to achieve high winning-rate performances while accelerating the policy exploration process in a tactic-interpretable fashion. Experimental results demonstrate the effectiveness of the proposed scheme against the state-of-the-art approaches in several benchmark combat scenarios.
Yue Hu 0008, Xi Li 0001, Gang Pan 0001, Mingliang Xu 0001
IJCAI3
2018 Semantic Locality-Aware Deformable Network for Clothing Segmentation
abstract
Clothing segmentation is a challenging vision problem typically implemented within a fine-grained semantic segmentation framework. Different from conventional segmentation, clothing segmentation has some domain-specific properties such as texture richness, diverse appearance variations, non-rigid geometry deformations, and small sample learning. To deal with these points, we propose a semantic locality-aware segmentation model, which adaptively attaches an original clothing image with a semantically similar (e.g., appearance or pose) auxiliary exemplar by search. Through considering the interactions of the clothing image and its exemplar, more intrinsic knowledge about the locality manifold structures of clothing images is discovered to make the learning process of small sample problem more stable and tractable. Furthermore, we present a CNN model based on the deformable convolutions to extract the non-rigid geometry-aware features for clothing images. Experimental results demonstrate the effectiveness of the proposed model against the state-of-the-art approaches.
Wei Ji 0008, Xi Li 0001, Yueting Zhuang, Omar El Farouk Bourahla, Yixin Ji, Jiabao Cui
IJCAI2
2018 Progressive Blockwise Knowledge Distillation for Neural Network Acceleration
abstract
As an important and challenging problem in machine learning and computer vision, neural network acceleration essentially aims to enhance the computational efficiency without sacrificing the model accuracy too much. In this paper, we propose a progressive blockwise learning scheme for teacher-student model distillation at the subnetwork block level. The proposed scheme is able to distill the knowledge of the entire teacher network by locally extracting the knowledge of each block in terms of progressive blockwise function approximation. Furthermore, we propose a structure design criterion for the student subnetwork block, which is able to effectively preserve the original receptive field from the teacher network. Experimental results demonstrate the effectiveness of the proposed scheme against the state-of-the-art approaches.
Hui Wang 0107, Hanbin Zhao, Xi Li 0001, Xu Tan 0003
IJCAI3
2018 Deep Convolutional Neural Networks with Merge-and-Run Mappings
abstract
A deep residual network, built by stacking a sequence of residual blocks, is easy to train, because identity mappings skip residual branches and thus improve information flow. To further reduce the training difficulty, we present a simple network architecture, deep merge-and-run neural networks. The novelty lies in a modularized building block, merge-and-run block, which assembles residual branches in parallel through a merge-and-run mapping: average the inputs of these residual branches (Merge), and add the average to the output of each residual branch as the input of the subsequent residual branch (Run), respectively. We show that the merge-and-run mapping is a linear idempotent function in which the transformation matrix is idempotent, and thus improves information flow, making training easy. In comparison with residual networks, our networks enjoy compelling advantages: they contain much shorter paths and the width, i.e., the number of channels, is increased, and the time complexity remains unchanged. We evaluate the performance on the standard recognition tasks. Our approach demonstrates consistent improvements over ResNets with the comparable setup, and achieves competitive results (e.g., 3.06% testing error on CIFAR-10, 17.55% on CIFAR-100, 1.51% on SVHN).
Mingjie Li 0007, Depu Meng, Xi Li 0001, Zhaoxiang Zhang 0001, Yueting Zhuang, Zhuowen Tu, Jingdong Wang 0001
IJCAI4
2018 GNAS: A Greedy Neural Architecture Search Method for Multi-Attribute Learning
abstract
A key problem in deep multi-attribute learning is to effectively discover the inter-attribute correlation structures. Typically, the conventional deep multi-attribute learning approaches follow the pipeline of manually designing the network architectures based on task-specific expertise prior knowledge and careful network tunings, leading to the inflexibility for various complicated scenarios in practice. Motivated by addressing this problem, we propose an efficient greedy neural architecture search approach (GNAS) to automatically discover the optimal tree-like deep architecture for multi-attribute learning. In a greedy manner, GNAS divides the optimization of global architecture into the optimizations of individual connections step by step. By iteratively updating the local architectures, the global tree-like architecture gets converged where the bottom layers are shared across relevant attributes and the branches in top layers more encode attribute-specific features. Experiments on three benchmark multi-attribute datasets show the effectiveness and compactness of neural architectures derived by GNAS, and also demonstrate the efficiency of GNAS in searching neural architectures.
Siyu Huang, Xi Li 0001, Zhi-Qi Cheng, Zhongfei Zhang, Alex Hauptmann 0001
ACM Multimedia2
2018 Multimodal Deep Embedding via Hierarchical Grounded Compositional Semantics
abstract
For a number of important problems, isolated semantic representations of individual syntactic words or visual objects do not suffice, but instead a compositional semantic representation is required; for example, a literal phrase or a set of spatially concurrent objects. In this paper, we aim to harness the existing image-sentence databases to exploit the compositional nature of image-sentence data for multimodal deep embedding. In particular, we propose an approach called hierarchical-alike (bottom-up two layers) multimodal grounded compositional semantics (hiMoCS) learning. The proposed hiMoCS systemically captures the compositional semantic connotation of multimodal data in the setting of hierarchical-alike deep learning by modeling the inherent correlations between two modalities of collaboratively grounded semantics, such as the textual entity (with its describing attribute) and visual object, the phrase (e.g., subject-verb-object triplet), and spatially concurrent objects. We argue that hiMoCS is more appropriate to reflect the multimodal compositional semantics of the image and its narrative textual sentence, which are strongly coupled. We evaluate hiMoCS on the several benchmark data sets and show that the utilization of the hiMoCS (textual entities and visual objects, textual phrase, and spatially concurrent objects) achieves a much better performance than only using the flat grounded compositional semantics.
Yueting Zhuang, Jun Song 0004, Fei Wu 0001, Xi Li 0001, Zhongfei Zhang, Yong Rui
IEEE Trans. Circuits Syst. Video Technol.4
2018 Transductive Zero-Shot Learning With a Self-Training Dictionary Approach
abstract
As an important and challenging problem in computer vision, zero-shot learning (ZSL) aims at automatically recognizing the instances from unseen object classes without training data. To address this problem, ZSL is usually carried out in the following two aspects: 1) capturing the domain distribution connections between seen classes data and unseen classes data and 2) modeling the semantic interactions between the image feature space and the label embedding space. Motivated by these observations, we propose a bidirectional mapping-based semantic relationship modeling scheme that seeks for cross-modal knowledge transfer by simultaneously projecting the image features and label embeddings into a common latent space. Namely, we have a bidirectional connection relationship that takes place from the image feature space to the latent space as well as from the label embedding space to the latent space. To deal with the domain shift problem, we further present a transductive learning approach that formulates the class prediction problem in an iterative refining process, where the object classification capacity is progressively reinforced through bootstrapping-based model updating over highly reliable instances. Experimental results on four benchmark datasets (animal with attribute, Caltech-UCSD Bird2011, aPascal-aYahoo, and SUN) demonstrate the effectiveness of the proposed approach against the state-of-the-art approaches.
Yunlong Yu 0001, Zhong Ji, Xi Li 0001, Jichang Guo, Zhongfei Zhang, Haibin Ling, Fei Wu 0001
IEEE Trans. Cybern.3
2018 Body Structure Aware Deep Crowd Counting
abstract
Crowd counting is a challenging task, mainly due to the severe occlusions among dense crowds. This paper aims to take a broader view to address crowd counting from the perspective of semantic modeling. In essence, crowd counting is a task of pedestrian semantic analysis involving three key factors: pedestrians, heads, and their context structure. The information of different body parts is an important cue to help us judge whether there exists a person at a certain position. Existing methods usually perform crowd counting from the perspective of directly modeling the visual properties of either the whole body or the heads only, without explicitly capturing the composite body-part semantic structure information that is crucial for crowd counting. In our approach, we first formulate the key factors of crowd counting as semantic scene models. Then, we convert the crowd counting problem into a multi-task learning problem, such that the semantic scene models are turned into different sub-tasks. Finally, the deep convolutional neural networks are used to learn the sub-tasks in a unified scheme. Our approach encodes the semantic nature of crowd counting and provides a novel solution in terms of pedestrian semantic analysis. In experiments, our approach outperforms the state-of-the-art methods on four benchmark crowd counting data sets. The semantic structure information is demonstrated to be an effective cue in scene of crowd counting.
Siyu Huang, Xi Li 0001, Zhongfei Zhang, Fei Wu 0001, Shenghua Gao, Rongrong Ji, Junwei Han 0001
IEEE Trans. Image Process.2
2018 Deep Context-Sensitive Facial Landmark Detection With Tree-Structured Modeling
abstract
Facial landmark detection is typically cast as a point-wise regression problem that focuses on how to build an effective image-to-point mapping function. In this paper, we propose an end-to-end deep learning approach for contextually discriminative feature construction together with effective facial structure modeling. The proposed learning approach is able to predict more contextually discriminative facial landmarks by capturing their associated contextual information. Moreover, we present a tree model to characterize human face structure and a structural loss function to measure the deformation cost between the ground-truth and predicted tree model, which are further incorporated into the proposed learning approach and jointly optimized within a unified framework. The presented tree model is able to well characterize the spatial layout patterns of facial landmarks for capturing the facial structure information. Experimental results demonstrate the effectiveness of the proposed approach against the state-of-the-art over the MTFL and AFLW-full data sets.
Jiajian Zeng, Xi Li 0001, Debbah Abderrahmane Mahdi, Fei Wu 0001, Gang Wang 0012
IEEE Trans. Image Process.3
2018 Deep Air Learning: Interpolation, Prediction, and Feature Analysis of Fine-Grained Air Quality
abstract
The interpolation, prediction, and feature analysis of fine-gained air quality are three important topics in the area of urban air computing. The solutions to these topics can provide extremely useful information to support air pollution control, and consequently generate great societal and technical impacts. Most of the existing work solves the three problems separately by different models. In this paper, we propose a general and effective approach to solve the three problems in one model called the Deep Air Learning (DAL). The main idea of DAL lies in embedding feature selection and semi-supervised learning in different layers of the deep learning network. The proposed approach utilizes the information pertaining to the unlabeled spatio-temporal data to improve the performance of the interpolation and the prediction, and performs feature selection and association analysis to reveal the main relevant features to the variation of the air quality. We evaluate our approach with extensive experiments based on real data sources obtained in Beijing, China. Experiments show that DAL is superior to the peer models from the recent literature when solving the topics of interpolation, prediction, and feature analysis of fine-gained air quality.
Zhongang Qi, Tianchun Wang, Guojie Song, Weisong Hu, Xi Li 0001, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.5
2018 Weakly Supervised Object Detection via Object-Specific Pixel Gradient
abstract
Most existing object detection algorithms are trained based upon a set of fully annotated object regions or bounding boxes, which are typically labor-intensive. On the contrary, nowadays there is a significant amount of image-level annotations cheaply available on the Internet. It is hence a natural thought to explore such "weak" supervision to benefit the training of object detectors. In this paper, we propose a novel scheme to perform weakly supervised object localization, termed object-specific pixel gradient (OPG). The OPG is trained by using image-level annotations alone, which performs in an iterative manner to localize potential objects in a given image robustly and efficiently. In particular, we first extract an OPG map to reveal the contributions of individual pixels to a given object category, upon which an iterative mining scheme is further introduced to extract instances or components of this object. Moreover, a novel average and max pooling layer is introduced to improve the localization accuracy. In the task of weakly supervised object localization, the OPG achieves a state-of-the-art 44.5% top-5 error on ILSVRC 2013, which outperforms competing methods, including Oquab et al. and region-based convolutional neural networks on the Pascal VOC 2012, with gains of 2.6% and 2.3%, respectively. In the task of object detection, OPG achieves a comparable performance of 27.0% mean average precision on Pascal VOC 2007. In all experiments, the OPG only adopts the off-the-shelf pretrained CNN model, without using any object proposals. Therefore, it also significantly improves the detection speed, i.e., achieving three times faster compared with the state-of-the-art method.
Yunhang Shen, Rongrong Ji, Changhu Wang, Xi Li 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.4
2018 Identifying Objective and Subjective Words via Topic Modeling
abstract
It is observed that distinct words in a given document have either strong or weak ability in delivering facts (i.e., the objective sense) or expressing opinions (i.e., the subjective sense) depending on the topics they associate with. Motivated by the intuitive assumption that different words have varying degree of discriminative power in delivering the objective sense or the subjective sense with respect to their assigned topics, a model named as dentified bjective- ubjective latent Dirichlet allocation (LDA) ( osLDA) is proposed in this paper. In the osLDA model, the simple Pólya urn model adopted in traditional topic models is modified by incorporating it with a probabilistic generative process, in which the novel "Bag-of-Discriminative-Words" (BoDW) representation for the documents is obtained; each document has two different BoDW representations with regard to objective and subjective senses, respectively, which are employed in the joint objective and subjective classification instead of the traditional Bag-of-Topics representation. The experiments reported on documents and images demonstrate that: 1) the BoDW representation is more predictive than the traditional ones; 2) osLDA boosts the performance of topic modeling via the joint discovery of latent topics and the different objective and subjective power hidden in every word; and 3) osLDA has lower computational complexity than supervised LDA, especially under an increasing number of topics.
Hanqi Wang, Fei Wu 0001, Weiming Lu 0001, Yi Yang 0001, Xi Li 0001, Xuelong Li 0001, Yueting Zhuang
IEEE Trans. Neural Networks Learn. Syst.5
2017 Pyramid Person Matching Network for Person Re-identification
abstract
In this work, we present a deep convolutional pyramid person matching network (PPMN) with specially designed Pyramid Matching Module to address the problem of person re-identification. The architecture takes a pair of RGB images as input, and outputs a similiarity value indicating whether the two input images represent the same person or not. Based on deep convolutional neural networks, our approach first learns the discriminative semantic representation with the semantic-component-aware features for persons and then employs the Pyramid Matching Module to match the common semantic-components of persons, which is robust to the variation of spatial scales and misalignment of locations posed by viewpoint changes. The above two processes are jointly optimized via a unified end-to-end deep learning scheme. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our approach against the state-of-the-art approaches, especially on the rank-1 recognition rate.
Chaojie Mao, Yingming Li, Zhongfei Zhang, Xi Li 0001
ACML5
2017 Deeply-Learned Part-Aligned Representations for Person Re-identification
abstract
In this paper, we address the problem of person re-identification, which refers to associating the persons captured from different cameras. We propose a simple yet effective human part-aligned representation for handling the body part misalignment problem. Our approach decomposes the human body into regions (parts) which are discriminative for person matching, accordingly computes the representations over the regions, and aggregates the similarities computed between the corresponding regions of a pair of probe and gallery images as the overall matching score. Our formulation, inspired by attention models, is a deep neural network modeling the three steps together, which is learnt through minimizing the triplet loss function without requiring body part labeling information. Unlike most existing deep learning algorithms that learn a global or spatial partition-based local representation, our approach performs human body partition, and thus is more robust to pose changes and various human spatial distributions in the person bounding box. Our approach shows state-of-the-art results over standard datasets, Market-1501, CUHK03, CUHK01 and VIPeR.
Xi Li 0001, Yueting Zhuang, Jingdong Wang 0001
ICCV2
2017 Graph-theoretic spatiotemporal context modeling for video saliency detection
abstract
As an important and challenging problem in computer vision, video saliency detection is typically cast as a spatiotemporal context modeling problem over consecutive frames. As a result, a key issue in video saliency detection is how to effectively capture the intrinsical properties of atomic video structures as well as their associated contextual interactions along the spatial and temporal dimensions. Motivated by this observation, we propose a graph-theoretic video saliency detection approach based on adaptive video structure discovery, which is carried out within a spatiotemporal atomic graph. Through graph-based manifold propagation, the proposed approach is capable of effectively modeling the semantically contextual interactions among atomic video structures for saliency detection while preserving spatial smoothness and temporal consistency. Experiments demonstrate the effectiveness of the proposed approach over several benchmark datasets.
Lina Wei, Xi Li 0001, Fei Wu 0001, Jun Xiao 0001
ICIP3
2017 Boosted Zero-Shot Learning with Semantic Correlation Regularization
abstract
We study zero-shot learning (ZSL) as a transfer learning problem, and focus on the two key aspects of ZSL, model effectiveness and model adaptation. For effective modeling, we adopt the boosting strategy to learn a zero-shot classifier from weak models to a strong model. For adaptable knowledge transfer, we devise a Semantic Correlation Regularization (SCR) approach to regularize the boosted model to be consistent with the inter-class semantic correlations. With SCR embedded in the boosting objective, and with a self-controlled sample selection for learning robustness, we propose a unified framework, Boosted Zero-shot classification with Semantic Correlation Regularization (BZ-SCR). By balancing the SCR-regularized boosted model selection and the self-controlled sample selection, BZ-SCR is capable of capturing both discriminative and adaptable feature-to-class semantic alignments, while ensuring the reliability and adaptability of the learned samples. The experiments on two ZSL datasets show the superiority of BZ-SCR over the state-of-the-arts.
Te Pi, Xi Li 0001, Zhongfei Zhang
IJCAI2
2017 Group-wise Deep Co-saliency Detection
abstract
In this paper, we propose an end-to-end group-wise deep co-saliency detection approach to address the co-salient object discovery problem based on the fully convolutional network (FCN) with group input and group output. The proposed approach captures the group-wise interaction information for group images by learning a semantics-aware image representation based on a convolutional neural network, which adaptively learns the group-wise features for co-saliency detection. Furthermore, the proposed approach discovers the collaborative and interactive relationships between group-wise feature representation and single-image individual feature representation, and model this in a collaborative learning framework. Finally, we set up a unified end-to-end deep learning scheme to jointly optimize the process of group-wise feature representation learning and the collaborative learning, leading to more reliable and robust co-saliency detection results. Experimental results demonstrate the effectiveness of our approach in comparison with the state-of-the-art approaches.
Lina Wei, Shanshan Zhao 0001, Omar El Farouk Bourahla, Xi Li 0001, Fei Wu 0001
IJCAI4
2017 Deep Optical Flow Estimation Via Multi-Scale Correspondence Structure Learning
abstract
As an important and challenging problem in computer vision, learning based optical flow estimation aims to discover the intrinsic correspondence structure between two adjacent video frames through statistical learning. Therefore, a key issue to solve in this area is how to effectively model the multi-scale correspondence structure properties in an adaptive end-to-end learning fashion. Motivated by this observation, we propose an end-to-end multi-scale correspondence structure learning (MSCSL) approach for optical flow estimation. In principle, the proposed MSCSL approach is capable of effectively capturing the multi-scale inter-image-correlation correspondence structures within a multi-level feature space from deep learning. Moreover, the proposed MSCSL approach builds a spatial Conv-GRU neural network model to adaptively model the intrinsic dependency relationships among these multi-scale correspondence structures. Finally, the above procedures for correspondence structure learning and multi-scale dependency modeling are implemented in a unified end-to-end deep learning framework. Experimental results on several benchmark datasets demonstrate the effectiveness of the proposed approach.
Shanshan Zhao 0001, Xi Li 0001, Omar El Farouk Bourahla
IJCAI2
2017 KeyphraseDS: Automatic generation of survey by exploiting keyphrase information
Shansong Yang, Weiming Lu 0001, Dezhi Yang, Xi Li 0001, Baogang Wei
Neurocomputing4
2017 Joint entity-relation knowledge embedding via cost-sensitive learning
abstract
As a joint-optimization problem which simultaneously fulfills two different but correlated embedding tasks (i.e., entity embedding and relation embedding), knowledge embedding problem is solved in a joint embedding scheme. In this embedding scheme, we design a joint compatibility scoring function to quantitatively evaluate the relational facts with respect to entities and relations, and further incorporate the scoring function into the max-margin structure learning process that explicitly learns the embedding vectors of entities and relations using the context information of the knowledge base. By optimizing the joint problem, our design is capable of effectively capturing the intrinsic topological structures in the learned embedding spaces. Experimental results demonstrate the effectiveness of our embedding scheme in characterizing the semantic correlations among different relation units, and in relation prediction for knowledge inference.
Shengkang Yu, Xueyi Zhao, Xi Li 0001, Zhongfei Zhang
Frontiers Inf. Technol. Electron. Eng.3
2017 Regularized Deep Belief Network for Image Attribute Detection
abstract
In general, an image attribute is a human-nameable visual property that has a semantic connotation. Appropriate modeling of the intrinsic contextual correlations among attributes plays a fundamental role in attribute detection. In this paper, we consider image attribute detection from the perspective of regularized deep learning. In particular, we propose a regularized deep belief network (rDBN) to perform the image attribute detection task, which is composed of two parts: 1) a detection DBN (dDBN) that models the joint distribution of images and their corresponding attributes, which acts as an attribute detector and 2) a contextual restricted Boltzmann machine that explicitly models the correlations among attributes acting as a regularizer that restraints the output detection result given by the dDBN to meet the contextual prior of attributes. Furthermore, we propose an efficient fine-tuning scheme that can further optimize the performance of the dDBN by backpropagation. Experimental results show that the proposed rDBN obtains improvements over the state-of-the-art methods for attribute detection on the benchmark data sets.
Fei Wu 0001, Zhuhao Wang, Weiming Lu 0001, Xi Li 0001, Yi Yang 0001, Jiebo Luo 0001, Yueting Zhuang
IEEE Trans. Circuits Syst. Video Technol.4
2017 Data-Dependent Label Distribution Learning for Age Estimation
abstract
As an important and challenging problem in computer vision, face age estimation is typically cast as a classification or regression problem over a set of face samples with respect to several ordinal age labels, which have intrinsically cross-age correlations across adjacent age dimensions. As a result, such correlations usually lead to the age label ambiguities of the face samples. Namely, each face sample is associated with a latent label distribution that encodes the cross-age correlation information on label ambiguities. Motivated by this observation, we propose a totally data-driven label distribution learning approach to adaptively learn the latent label distributions. The proposed approach is capable of effectively discovering the intrinsic age distribution patterns for cross-age correlation analysis on the basis of the local context structures of face samples. Without any prior assumptions on the forms of label distribution learning, our approach is able to flexibly model the sample-specific context aware label distribution properties by solving a multi-task problem, which jointly optimizes the tasks of age-label distribution learning and age prediction for individuals. Experimental results demonstrate the effectiveness of our approach.
Zhouzhou He, Xi Li 0001, Zhongfei Zhang, Fei Wu 0001, Xin Geng 0001, Ming-Hsuan Yang 0001, Yueting Zhuang
IEEE Trans. Image Process.2
2017 Learning Bregman Distance Functions for Structural Learning to Rank
abstract
We study content-based learning to rank from the perspective of learning distance functions. Standardly, the two key issues of learning to rank, feature mappings and score functions, are usually modeled separately, and the learning is usually restricted to modeling a linear distance function such as the Mahalanobis distance. However, the modeling of feature mappings and score functions are mutually interacted, and the patterns underlying the data are probably complicated and nonlinear. Thus, as a general nonlinear distance family, the Bregman distance is a suitable distance function for learning to rank, due to its strong generalization ability for distance functions, and its nonlinearity for exploring the general patterns of data distributions. In this paper, we study learning to rank as a structural learning problem, and devise a Bregman distance function to build the ranking model based on structural SVM. To improve the model robustness to outliers, we develop a robust structural learning framework for the ranking model. The proposed model Robust Structural Bregman distance functions Learning to Rank (RSBLR) is a general and unified framework for learning distance functions to rank. The experiments of data ranking on real-world datasets show the superiority of this method to the state-of-the-art literature, as well as its robustness to the noisily labeled outliers.
Xi Li 0001, Te Pi, Zhongfei Zhang, Xueyi Zhao, Meng Wang 0001, Xuelong Li 0001, Philip S. Yu
IEEE Trans. Knowl. Data Eng.1
2016 Self-Paced Boost Learning for Classification
Te Pi, Xi Li 0001, Zhongfei Zhang, Deyu Meng, Fei Wu 0001, Jun Xiao 0001, Yueting Zhuang
IJCAI2
2016 Diverse Image Captioning via GroupTalk
Zhuhao Wang, Fei Wu 0001, Weiming Lu 0001, Jun Xiao 0001, Xi Li 0001, Yueting Zhuang
IJCAI5
2016 Semantics-Aware Deep Correspondence Structure Learning for Robust Person Re-Identification
Xi Li 0001, Zhongfei Zhang
IJCAI2
2016 Fusing ℝ Features and Local Features with Context-Aware Kernels for Action Recognition
Chunfeng Yuan, Baoxin Wu, Xi Li 0001, Weiming Hu 0004, Stephen J. Maybank, Fangshi Wang
Int. J. Comput. Vis.3
2016 Online Metric-Weighted Linear Representations for Robust Visual Tracking
abstract
In this paper, we propose a visual tracker based on a metric-weighted linear representation of appearance. In order to capture the interdependence of different feature dimensions, we develop two online distance metric learning methods using proximity comparison information and structured output learning. The learned metric is then incorporated into a linear representation of appearance. We show that online distance metric learning significantly improves the robustness of the tracker, especially on those sequences exhibiting drastic appearance changes. In order to bound growth in the number of training samples, we design a time-weighted reservoir sampling method. Moreover, we enable our tracker to automatically perform object identification during the process of object tracking, by introducing a collection of static template samples belonging to several object classes of interest. Object identification results for an entire video sequence are achieved by systematically combining the tracking information and visual recognition at each frame. Experimental results on challenging video sequences demonstrate the effectiveness of the method for both inter-frame tracking and object identification.
Xi Li 0001, Chunhua Shen, Anthony R. Dick, Zhongfei Zhang, Yueting Zhuang
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 Structure-Aware Slow Feature Analysis for Age Estimation
abstract
As an important and challenging problem in computer vision, face age estimation is typically cast as a classification or regression problem over a set of face samples. However, most existing efforts to age estimation usually cope with the face samples individually, which do not take full advantage of the temporal structure and contextual structure of the face samples. In this letter, we propose an age estimation approach named structure-aware slow feature analysis, which is capable of effectively capturing the structure of human faces in the aspects of time-related smoothness for progressive age variation as well as face-related attribute constraints for face age consistency. As a result, we present an iterative optimization scheme to effectively learn the slowly varying feature transformation. Experimental results demonstrate the effectiveness of our approach on the Morph dataset.
Zhouzhou He, Xi Li 0001, Zhongfei Zhang, Jun Xiao 0001
IEEE Signal Process. Lett.2
2016 Aspect Learning for Multimedia Summarization via Nonparametric Bayesian
abstract
Summarization is desirable for efficient comprehension of an increasingly vast amount of data. A summary of multiple documents is a concise description of the main topic. Generally speaking, a topic delivers various aspects. For example, the natural disaster topic is likely to imply the aspects of casualties and rescue. Therefore, a good summary is expected to cover all the informative aspects of a topic in order to enhance both diversity and coverage of the topic. However, for the real-world data, the profile of aspects in a given topic (e.g., the number of the aspects as well as their appropriate describing sentences or images) is hardly specified in advance. To address this problem, this paper proposes an approach to learn the hidden aspects in the topics via a nonparametric Bayesian model for multimedia summarization, namely, aspect learning for multimedia summarization via nonparametric Bayesian (ALSNB). More specifically, we introduce the priors of beta-Bernoulli process and Dirichlet process into the traditional dictionary learning. As a result, the proposed approach is able to adaptively identify the particular aspects of an individual topic. The experimental results on several datasets for text summarization and image summarization show the superiority of the proposed ALSNB over other methods.
Fei Wu 0001, Hanyin Fang, Xi Li 0001, Siliang Tang, Weiming Lu 0001, Yi Yang 0001, Wenwu Zhu 0001, Yueting Zhuang
IEEE Trans. Circuits Syst. Video Technol.3
2016 Learning A Superpixel-Driven Speed Function for Level Set Tracking
abstract
A key problem in level set tracking is to construct a discriminative speed function for effective contour evolution. In this paper, we propose a level set tracking method based on a discriminative speed function, which produces a superpixel-driven force for effective level set evolution. Based on kernel density estimation and metric learning, the speed function is capable of effectively encoding the discriminative information on object appearance within a feasible metric space. Furthermore, we introduce adaptive object shape modeling into the level set evolution process, which leads to the tracking robustness in complex scenarios. To ensure the efficiency of adaptive object shape modeling, we develop a simple but efficient weighted non-negative matrix factorization method that can online learn an object shape dictionary. Experimental results on a number of challenging video sequences demonstrate the effectiveness and robustness of the proposed tracking method.
Xi Li 0001, Weiming Hu 0004
IEEE Trans. Cybern.2
2016 Deep Learning Driven Visual Path Prediction From a Single Image
abstract
Capabilities of inference and prediction are the significant components of visual systems. Visual path prediction is an important and challenging task among them, with the goal to infer the future path of a visual object in a static scene. This task is complicated as it needs high-level semantic understandings of both the scenes and underlying motion patterns in video sequences. In practice, cluttered situations have also raised higher demands on the effectiveness and robustness of models. Motivated by these observations, we propose a deep learning framework, which simultaneously performs deep feature learning for visual representation in conjunction with spatiotemporal context modeling. After that, a unified path-planning scheme is proposed to make accurate path prediction based on the analytic results returned by the deep context models. The highly effective visual representation and deep context models ensure that our framework makes a deep semantic understanding of the scenes and motion patterns, consequently improving the performance on visual path prediction task. In experiments, we extensively evaluate the model's performance by constructing two large benchmark datasets from the adaptation of video tracking datasets. The qualitative and quantitative experimental results show that our approach outperforms the state-of-the-art approaches and owns a better generalization capability.
Siyu Huang, Xi Li 0001, Zhongfei Zhang, Zhouzhou He, Fei Wu 0001, Wei Liu 0005, Jinhui Tang 0001, Yueting Zhuang
IEEE Trans. Image Process.2
2016 DeepSaliency: Multi-Task Deep Neural Network Model for Salient Object Detection
abstract
A key problem in salient object detection is how to effectively model the semantic properties of salient objects in a data-driven manner. In this paper, we propose a multi-task deep saliency model based on a fully convolutional neural network with global input (whole raw images) and global output (whole saliency maps). In principle, the proposed saliency model takes a data-driven strategy for encoding the underlying saliency prior information, and then sets up a multi-task learning scheme for exploring the intrinsic correlations between saliency detection and semantic image segmentation. Through collaborative feature learning from such two correlated tasks, the shared fully convolutional layers produce effective features for object perception. Moreover, it is capable of capturing the semantic information on salient objects across different levels using the fully convolutional layers, which investigate the feature-sharing properties of salient object detection with a great reduction of feature redundancy. Finally, we present a graph Laplacian regularized nonlinear regression model for saliency refinement. Experimental results demonstrate the effectiveness of our approach in comparison with the state-of-the-art approaches.
Xi Li 0001, Lina Wei, Ming-Hsuan Yang 0001, Fei Wu 0001, Yueting Zhuang, Haibin Ling, Jingdong Wang 0001
IEEE Trans. Image Process.1
2016 Joint Multilabel Classification With Community-Aware Label Graph Learning
abstract
As an important and challenging problem in machine learning and computer vision, multilabel classification is typically implemented in a max-margin multilabel learning framework, where the inter-label separability is characterized by the sample-specific classification margins between labels. However, the conventional multilabel classification approaches are usually incapable of effectively exploring the intrinsic inter-label correlations as well as jointly modeling the interactions between inter-label correlations and multilabel classification. To address this issue, we propose a multilabel classification framework based on a joint learning approach called label graph learning (LGL) driven weighted Support Vector Machine (SVM). In principle, the joint learning approach explicitly models the inter-label correlations by LGL, which is jointly optimized with multilabel classification in a unified learning scheme. As a result, the learned label correlation graph well fits the multilabel classification task while effectively reflecting the underlying topological structures among labels. Moreover, the inter-label interactions are also influenced by label-specific sample communities (each community for the samples sharing a common label). Namely, if two labels have similar label-specific sample communities, they are likely to be correlated. Based on this observation, LGL is further regularized by the label Hypergraph Laplacian. Experimental results have demonstrated the effectiveness of our approach over several benchmark data sets.
Xi Li 0001, Xueyi Zhao, Zhongfei Zhang, Fei Wu 0001, Yueting Zhuang, Jingdong Wang 0001, Xuelong Li 0001
IEEE Trans. Image Process.1
2016 Scalable Linear Visual Feature Learning via Online Parallel Nonnegative Matrix Factorization
abstract
Visual feature learning, which aims to construct an effective feature representation for visual data, has a wide range of applications in computer vision. It is often posed as a problem of nonnegative matrix factorization (NMF), which constructs a linear representation for the data. Although NMF is typically parallelized for efficiency, traditional parallelization methods suffer from either an expensive computation or a high runtime memory usage. To alleviate this problem, we propose a parallel NMF method called alternating least square block decomposition (ALSD), which efficiently solves a set of conditionally independent optimization subproblems based on a highly parallelized fine-grained grid-based blockwise matrix decomposition. By assigning each block optimization subproblem to an individual computing node, ALSD can be effectively implemented in a MapReduce-based Hadoop framework. In order to cope with dynamically varying visual data, we further present an incremental version of ALSD, which is able to incrementally update the NMF solution with a low computational cost. Experimental results demonstrate the efficiency and scalability of the proposed methods as well as their applications to image clustering and image retrieval.
Xueyi Zhao, Xi Li 0001, Zhongfei Zhang, Chunhua Shen, Yueting Zhuang, Lixin Gao 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2015 Structured Embedding via Pairwise Relations and Long-Range Interactions in Knowledge Base
abstract
We consider the problem of embedding entities and relations of knowledge bases into low-dimensional continuous vector spaces (distributed representations). Unlike most existing approaches, which are primarily efficient for modelling pairwise relations between entities, we attempt to explicitly model both pairwise relations and long-range interactions between entities, by interpreting them as linear operators on the low-dimensional embeddings of the entities. Therefore, in this paper we introduces Path-Ranking to capture the long-range interactions of knowledge graph and at the same time preserve the pairwise relations of knowledge graph; we call it 'structured embedding via pairwise relation and long-range interactions' (referred to as SePLi). Comparing with the-state-of-the-art models, SePLi achieves better performances of embeddings.
Fei Wu 0001, Jun Song 0004, Yi Yang 0001, Xi Li 0001, Zhongfei Zhang, Yueting Zhuang
AAAI4
2015 Metric Learning Driven Multi-Task Structured Output Optimization for Robust Keypoint Tracking
abstract
As an important and challenging problem in computer vision and graphics, keypoint-based object tracking is typically formulated in a spatio-temporal statistical learning framework. However, most existing keypoint trackers are incapable of effectively modeling and balancing the following three aspects in a simultaneous manner: temporal model coherence across frames, spatial model consistency within frames, and discriminative feature construction. To address this issue, we propose a robust keypoint tracker based on spatio-temporal multi-task structured output optimization driven by discriminative metric learning. Consequently, temporal model coherence is characterized by multi-task structured keypoint model learning over several adjacent frames, while spatial model consistency is modeled by solving a geometric verification based structured learning problem. Discriminative feature construction is enabled by metric learning to ensure the intra-class compactness and inter-class separability. Finally, the above three modules are simultaneously optimized in a joint learning scheme. Experimental results have demonstrated the effectiveness of our tracker.
Xi Li 0001, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang
AAAI2
2015 3D Hand Pose Estimation Using Randomized Decision Forest with Segmentation Index Points
abstract
In this paper, we propose a real-time 3D hand pose estimation algorithm using the randomized decision forest framework. Our algorithm takes a depth image as input and generates a set of skeletal joints as output. Previous decision forest-based methods often give labels to all points in a point cloud at a very early stage and vote for the joint locations. By contrast, our algorithm only tracks a set of more flexible virtual landmark points, named segmentation index points (SIPs), before reaching the final decision at a leaf node. Roughly speaking, a SIP represents the centroid of a subset of skeletal joints, which are to be located at the leaves of the branch expanded from the SIP. Inspired by recent latent regression forest-based hand pose estimation framework (Tang et al. 2014), we integrate SIP into the framework with several important improvements: First, we devise a new forest growing strategy, whose decision is made using a randomized feature guided by SIPs. Second, we speed-up the training procedure since only SIPs, not the skeletal joints, are estimated at non-leaf nodes. Third, the experimental results on public benchmark datasets show clearly the advantage of the proposed algorithm over previous state-of-the-art methods, and our algorithm runs at 55.5 fps on a normal CPU without parallelism.
Peiyi Li 0001, Haibin Ling, Xi Li 0001, Chunyuan Liao
ICCV3
2015 Deep Compositional Cross-modal Learning to Rank via Local-Global Alignment
abstract
Cross-modal retrieval is a very hot research topic that is imperative to many applications involving multi-modal data. Discovering an appropriate representation for multi-modal data and learning a ranking function are essential to boost the cross-media retrieval. Motivated by the assumption that a compositional cross-modal semantic representation (pairs of images and text) is more attractive for cross-modal ranking, this paper exploits the existing image-text databases to optimize a ranking function for cross-modal retrieval, called deep compositional cross-modal learning to rank (C2MLR). In this paper, C2MLR considers learning a multi-modal embedding from the perspective of optimizing a pairwise ranking problem while enhancing both local alignment and global alignment. In particular, the local alignment (i.e., the alignment of visual objects and textual words) and the global alignment (i.e., the image-level and sentence-level alignment) are collaboratively utilized to learn the multi-modal embedding common space in a max-margin learning to rank manner. The experiments demonstrate the superiority of our proposed C2MLR due to its nature of multi-modal compositional embedding.
Xinyang Jiang, Fei Wu 0001, Xi Li 0001, Zhou Zhao 0001, Weiming Lu 0001, Siliang Tang, Yueting Zhuang
ACM Multimedia3
2015 Tracking news article evolution by dense subgraph learning
Shengkang Yu, Xi Li 0001, Xueyi Zhao, Zhongfei Zhang, Fei Wu 0001
Neurocomputing2
2015 Deep learning driven blockwise moving object detection with binary scene modeling
Xi Li 0001, Zhongfei Zhang, Fei Wu 0001
Neurocomputing2
2015 Dynamic spatio-temporal modeling for example-based human silhouette recovery
Xi Li 0001
Signal Process.2
2015 Multimedia Retrieval via Deep Learning to Rank
abstract
Many existing learning-to-rank approaches are incapable of effectively modeling the intrinsic interaction relationships between the feature-level and ranking-level components of a ranking model. To address this problem, we propose a novel joint learning-to-rank approach called Deep Latent Structural SVM (DL-SSVM), which jointly learns deep neural networks and latent structural SVM (connected by a set of latent feature grouping variables) to effectively model the interaction relationships at two levels (i.e., feature-level and ranking-level). To make the joint learning problem easier to optimize, we present an effective auxiliary variable-based alternating optimization approach with respect to deep neural network learning and structural latent SVM learning. Experimental results on several challenging datasets have demonstrated the effectiveness of the proposed learning to rank approach in real-world information retrieval.
Xueyi Zhao, Xi Li 0001, Zhongfei Zhang
IEEE Signal Process. Lett.2
2015 Cross-Modal Learning to Rank via Latent Joint Representation
abstract
Cross-modal ranking is a research topic that is imperative to many applications involving multimodal data. Discovering a joint representation for multimodal data and learning a ranking function are essential in order to boost the cross-media retrieval (i.e., image-query-text or text-query-image). In this paper, we propose an approach to discover the latent joint representation of pairs of multimodal data (e.g., pairs of an image query and a text document) via a conditional random field and structural learning in a listwise ranking manner. We call this approach cross-modal learning to rank via latent joint representation (CML²R). In CML²R, the correlations between multimodal data are captured in terms of their sharing hidden variables (e.g., topics), and a hidden-topic-driven discriminative ranking function is learned in a listwise ranking manner. The experiments show that the proposed approach achieves a good performance in cross-media retrieval and meanwhile has the capability to learn the discriminative representation of multimodal data.
Fei Wu 0001, Xinyang Jiang, Xi Li 0001, Siliang Tang, Weiming Lu 0001, Zhongfei Zhang, Yueting Zhuang
IEEE Trans. Image Process.3
2015 Joint Structural Learning to Rank with Deep Linear Feature Learning
abstract
Multimedia information retrieval usually involves two key modules including effective feature representation and ranking model construction. Most existing approaches are incapable of well modeling the inherent correlations and interactions between them, resulting in the loss of the latent consensus structure information. To alleviate this problem, we propose a learning to rank approach that simultaneously obtains a set of deep linear features and constructs structure-aware ranking models in a joint learning framework. Specifically, the deep linear feature learning corresponds to a series of matrix factorization tasks in a hierarchical manner, while the learning-to-rank part concentrates on building a ranking model that effectively encodes the intrinsic ranking information by structural SVM learning. Through a joint learning mechanism, the two parts are mutually reinforced in our approach, and meanwhile their underlying interaction relationships are implicitly reflected by solving an alternating optimization problem. Due to the intrinsic correlations among different queries (i.e., similar queries for similar ranking lists), we further formulate the learning-to-rank problem as a multi-task problem, which is associated with a set of mutually related query-specific learning-to-rank subproblems. For computational efficiency and scalability, we design a MapReduce-based parallelization approach to speed up the learning processes. Experimental results demonstrate the efficiency, effectiveness, and scalability of the proposed approach in multimedia information retrieval.
Xueyi Zhao, Xi Li 0001, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.2
2015 Structured Visual Feature Learning for Classification via Supervised Probabilistic Tensor Factorization
abstract
In this paper, structured visual feature learning aims at exploiting the intrinsic structural properties of mutually correlated multimedia collections (e.g., video frames or facial images) to learn a more effective feature representation for multimedia data classification. We pose structured visual feature learning as a problem of supervised tensor factorization (STF), which is capable of effectively learning multi-view visual features from structural tensorial multimedia data. In mathematics , STF is formulated as a joint optimization framework of probabilistic inference and$\epsilon $-insensitive support vector regression. As a result, the feature representation obtained by STF not only preserves the intrinsic multi-view structural information on tensorial multimedia data, but also includes the discriminative information derived from the max-margin learning process. Using the learned discriminative visual features, we conduct a set of multimedia classification experiments on several challenging datasets, including images and videos, which demonstrate the effectiveness of our method.
Xu Tan 0003, Fei Wu 0001, Xi Li 0001, Siliang Tang, Weiming Lu 0001, Yueting Zhuang
IEEE Trans. Multim.3
2014 Context Based Re-ranking for Object Retrieval
Yanzhi Chen, Anthony R. Dick, Xi Li 0001, Rhys Hill
ACCV (1)3
2014 Structural Bregman Distance Functions Learning to Rank with Self-Reinforcement
abstract
Learning to rank is an important task for many data mining applications. Essentially, the goal of learning to rank is to learn an appropriate similarity or distance metric to determine the relevance relationships among data points. However, most of the existing approaches for distance metric learning are limited in three aspects. First, they often assume a fixed form of distance metric for the entire input space. Second, the assumed distance functions are often computationally expensive or even intractable to learn for high dimensional data, such as Mahalanobis distance. Third, most of these approaches lack robustness to noisily labeled data, which is pervasive in many real-world applications. In this paper, we study learning to rank as a problem of distance metric learning to address the above three problems. We choose Bregman distance as the target distance function, due to its general functional form as a generalization of a wide class of distance functions, and its capacity of exploiting complicated nonlinear patterns underlying the data. Under the framework of structural SVM, we formulate the problem of learning Bregman distance functions for ranking as a QP problem by a nonparametric approach, and present an effective algorithm. Furthermore, we propose a self-reinforcement scheme that adaptively differentiates each data point in the role of learning to secure the robustness. We emphasize that the proposed method SBLR-S (Structural Bregman distance functions Learning to Rank with Self-reinforcement) is more general than the conventional distance metric learning approaches, and is able to handle high dimensional data as well as noisily labeled data. The experiments of data ranking on real-world datasets show the superiority of this method to the state-of-the-art literature.
Te Pi, Xi Li 0001, Zhongfei Zhang
ICDM2
2014 Learning Multimodal Neural Network with Ranking Examples
abstract
To support cross-modal information retrieval, cross-modal learning to rank approaches utilize ranking examples (e.g., an example may be a text query and its corresponding ranked images) to learn appropriate ranking (similarity) function. However, the fact that each modality is represented with intrinsically different low-level features hinders these approaches from better reducing the heterogeneity-gap between the modalities and thus giving satisfactory retrieval results. In this paper, we consider learning with neural networks, from the perspective of optimizing the listwise ranking loss of the cross-modal ranking examples. The proposed model, named Cross-Modal Ranking Neural Network (CMRNN), benefits from the advance of both neural networks on learning high-level semantics and learning to rank techniques on learning ranking function, such that the learned cross-modal ranking function is implicitly embedded in the learned high-level representation for data objects with different modalities (e.g., text and imagery) to perform cross-modal retrieval directly. We compare CMRNN to existing state-of-the-art cross-modal ranking methods on two datasets and show that it achieves a better performance.
Fei Wu 0001, Xi Li 0001, Yin Zhang 0006, Weiming Lu 0001, Yueting Zhuang
ACM Multimedia3
2014 Jointly Discovering Fine-grained and Coarse-grained Sentiments via Topic Modeling
abstract
The ever-increasing user-generated contents in social media and other web services make it highly desirable to discover opinions of users on all kinds of topics. Motivated by the assumption that individual word and paragraph in documents will deliver fine-grained (e.g., "laudatory", "annoyed" or "boring") and coarse-grained (e.g., positive, negative or neutral) sentiments about certain topics respectively, this paper focuses on a deeper thematic level to jointly disentangle fine-grained and coarse-grained opinions towards topics in terms of sentiment analysis, named as LDA with multi-grained sentiments (MgS-LDA). As a result, the proposed MgS-LDA not only discovers the topics in social media, but also identifies opinions about a given topic in terms of fine-grained and coarse-grained sentiment. Results of several experiments show that our proposed MgS-LDA achieves better performance on both sentimental classification and topic modeling than related methods.
Hanqi Wang, Fei Wu 0001, Xi Li 0001, Siliang Tang, Jian Shao 0001, Yueting Zhuang
ACM Multimedia3
2014 Multi-modal Mutual Topic Reinforce Modeling for Cross-media Retrieval
abstract
As an important and challenging problem in the multimedia area, multi-modal data understanding aims to explore the intrinsic semantic information across different modalities in a collaborative manner. To address this problem, a possible solution is to effectively and adaptively capture the common cross-modal semantic information by modeling the inherent correlations between the latent topics from different modalities. Motivated by this task, we propose a supervised multi-modal mutual topic reinforce modeling (M$^3$R) approach, which seeks to build a joint cross-modal probabilistic graphical model for discovering the mutually consistent semantic topics via appropriate interactions between model factors (e.g., categories, latent topics and observed multi-modal data). In principle, M$^3$R is capable of simultaneously accomplishing the following two learning tasks: 1) modality-specific (e.g., image-specific or text-specific ) latent topic learning; and 2) cross-modal mutual topic consistency learning. By investigating the cross-modal topic-related distribution information, M$^3$R encourages to disentangle the semantically consistent cross-modal topics (containing some common semantic information across different modalities). In other words, the semantically co-occurring cross-modal topics are reinforced by M$^3$R through adaptively passing the mutually reinforced messages to each other in the model-learning process. To further enhance the discriminative power of the learned latent topic representations, M$^3$R incorporates the auxiliary information (i.e., categories or labels) into the process of Bayesian modeling, which boosts the modeling capability of capturing the inter-class discriminative information. Experimental results over two benchmark datasets demonstrate the effectiveness of the proposed M$^3$R in cross-modal retrieval.
Fei Wu 0001, Jun Song 0004, Xi Li 0001, Yueting Zhuang
ACM Multimedia4
2014 Ranking consistency for image matching and object retrieval
Yanzhi Chen, Xi Li 0001, Anthony R. Dick, Rhys Hill
Pattern Recognit.2
2014 Modeling Geometric-Temporal Context With Directional Pyramid Co-Occurrence for Action Recognition
abstract
In this paper, we present a new geometric-temporal representation for visual action recognition based on local spatio-temporal features. First, we propose a modified covariance descriptor under the log-Euclidean Riemannian metric to represent the spatio-temporal cuboids detected in the video sequences. Compared with previously proposed covariance descriptors, our descriptor can be measured and clustered in Euclidian space. Second, to capture the geometric-temporal contextual information, we construct a directional pyramid co-occurrence matrix (DPCM) to describe the spatio-temporal distribution of the vector-quantized local feature descriptors extracted from a video. DPCM characterizes the co-occurrence statistics of local features as well as the spatio-temporal positional relationships among the concurrent features. These statistics provide strong descriptive power for action recognition. To use DPCM for action recognition, we propose a directional pyramid co-occurrence matching kernel to measure the similarity of videos. The proposed method achieves the state-of-the-art performance and improves on the recognition performance of the bag-of-visual-words (BOVWs) models by a large margin on six public data sets. For example, on the KTH data set, it achieves 98.78% accuracy while the BOVW approach only achieves 88.06%. On both Weizmann and UCF CIL data sets, the highest possible accuracy of 100% is achieved.
Chunfeng Yuan, Xi Li 0001, Weiming Hu 0004, Haibin Ling, Stephen J. Maybank
IEEE Trans. Image Process.2
2014 Context-Aware Hypergraph Construction for Robust Spectral Clustering
abstract
Spectral clustering is a powerful tool for unsupervised data analysis. In this paper, we propose a context-aware hypergraph similarity measure (CAHSM), which leads to robust spectral clustering in the case of noisy data. We construct three types of hypergraphs-the pairwise hypergraph, the k-nearest-neighbor (kNN) hypergraph, and the high-order over-clustering hypergraph. The pairwise hypergraph captures the pairwise similarity of data points; the kNNhypergraph captures the neighborhood of each point; and the clustering hypergraph encodes high-order contexts within the dataset. By combining the affinity information from these three hypergraphs, the CAHSM algorithm is able to explore the intrinsic topological information of the dataset. Therefore, data clustering using CAHSM tends to be more robust. Considering the intra-cluster compactness and the inter-cluster separability of vertices, we further design a discriminative hypergraph partitioning criterion (DHPC). Using both CAHSM and DHPC, a robust spectral clustering algorithm is developed. Theoretical analysis and experimental evaluation demonstrate the effectiveness and robustness of the proposed algorithm.
Xi Li 0001, Weiming Hu 0004, Chunhua Shen, Anthony R. Dick, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.1
2013 Learning Compact Binary Codes for Visual Tracking
abstract
A key problem in visual tracking is to represent the appearance of an object in a way that is robust to visual changes. To attain this robustness, increasingly complex models are used to capture appearance variations. However, such models can be difficult to maintain accurately and efficiently. In this paper, we propose a visual tracker in which objects are represented by compact and discriminative binary codes. This representation can be processed very efficiently, and is capable of effectively fusing information from multiple cues. An incremental discriminative learner is then used to construct an appearance model that optimally separates the object from its surrounds. Furthermore, we design a hyper graph propagation method to capture the contextual information on samples, which further improves the tracking accuracy. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracker.
Xi Li 0001, Chunhua Shen, Anthony R. Dick, Anton van den Hengel
CVPR1
2013 3D R Transform on Spatio-temporal Interest Points for Action Recognition
abstract
Spatio-temporal interest points serve as an elementary building block in many modern action recognition algorithms, and most of them exploit the local spatio-temporal volume features using a Bag of Visual Words (BOVW) representation. Such representation, however, ignores potentially valuable information about the global spatio-temporal distribution of interest points. In this paper, we propose a new global feature to capture the detailed geometrical distribution of interest points. It is calculated by using the R transform which is defined as an extended 3D discrete Radon transform, followed by applying a two-directional two-dimensional principal component analysis. Such R feature captures the geometrical information of the interest points and keeps invariant to geometry transformation and robust to noise. In addition, we propose a new fusion strategy to combine the R feature with the BOVW representation for further improving recognition accuracy. We utilize a context-aware fusion method to capture both the pairwise similarities and higher-order contextual interactions of the videos. Experimental results on several publicly available datasets demonstrate the effectiveness of the proposed approach for action recognition.
Chunfeng Yuan, Xi Li 0001, Weiming Hu 0004, Haibin Ling, Stephen J. Maybank
CVPR2
2013 Contextual Hypergraph Modeling for Salient Object Detection
abstract
Salient object detection aims to locate objects that capture human attention within images. Previous approaches often pose this as a problem of image contrast analysis. In this work, we model an image as a hyper graph that utilizes a set of hyper edges to capture the contextual properties of image pixels or regions. As a result, the problem of salient object detection becomes one of finding salient vertices and hyper edges in the hyper graph. The main advantage of hyper graph modeling is that it takes into account each pixel's (or region's) affinity with its neighborhood as well as its separation from image background. Furthermore, we propose an alternative approach based on center-versus-surround contextual contrast analysis, which performs salient object detection by optimizing a cost-sensitive support vector machine (SVM) objective function. Experimental results on four challenging datasets demonstrate the effectiveness of the proposed approaches against the state-of-the-art approaches to salient object detection.
Xi Li 0001, Yao Li 0003, Chunhua Shen, Anthony R. Dick, Anton van den Hengel
ICCV1
2013 Learning Hash Functions Using Column Generation
abstract
Fast nearest neighbor searching is becoming an increasingly important tool in solving many large-scale problems. Recently a number of approaches to learning data-dependent hash functions have been developed. In this work, we propose a column generation based method for learning data-dependent hash functions on the basis of proximity comparison information. Given a set of triplets that encode the pairwise proximity comparison information, our method learns hash functions that preserve the relative comparison relationships in the data as well as possible within the large-margin learning framework. The learning procedure is implemented using column generation and hence is named CGHash. At each iteration of the column generation procedure, the best hash function is selected. Unlike most other hashing methods, our method generalizes to new data points naturally; and has a training objective which is convex, thus ensuring that the global optimum can be identified. Experiments demonstrate that the proposed method learns compact binary codes and that its retrieval performance compares favorably with state-of-the-art methods when tested on a few benchmark datasets.
Xi Li 0001, Guosheng Lin, Chunhua Shen, Anton van den Hengel, Anthony R. Dick
ICML (1)1
2013 An Improved Hierarchical Dirichlet Process-Hidden Markov Model and Its Application to Trajectory Modeling and Retrieval
Weiming Hu 0004, Guodong Tian, Xi Li 0001, Stephen J. Maybank
Int. J. Comput. Vis.3
2013 Spatially aware feature selection and weighting for object retrieval
Yanzhi Chen, Anthony R. Dick, Xi Li 0001, Anton van den Hengel
Image Vis. Comput.3
2013 An Incremental DPMM-Based Method for Trajectory Clustering, Modeling, and Retrieval
abstract
Trajectory analysis is the basis for many applications, such as indexing of motion events in videos, activity recognition, and surveillance. In this paper, the Dirichlet process mixture model (DPMM) is applied to trajectory clustering, modeling, and retrieval. We propose an incremental version of a DPMM-based clustering algorithm and apply it to cluster trajectories. An appropriate number of trajectory clusters is determined automatically. When trajectories belonging to new clusters arrive, the new clusters can be identified online and added to the model without any retraining using the previous data. A time-sensitive Dirichlet process mixture model (tDPMM) is applied to each trajectory cluster for learning the trajectory pattern which represents the time-series characteristics of the trajectories in the cluster. Then, a parameterized index is constructed for each cluster. A novel likelihood estimation algorithm for the tDPMM is proposed, and a trajectory-based video retrieval model is developed. The tDPMM-based probabilistic matching method and the DPMM-based model growing method are combined to make the retrieval model scalable and adaptable. Experimental comparisons with state-of-the-art algorithms demonstrate the effectiveness of our algorithm.
Weiming Hu 0004, Xi Li 0001, Guodong Tian, Stephen J. Maybank, Zhongfei Zhang
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Incremental Learning of 3D-DCT Compact Representations for Robust Visual Tracking
abstract
Visual tracking usually requires an object appearance model that is robust to changing illumination, pose, and other factors encountered in video. Many recent trackers utilize appearance samples in previous frames to form the bases upon which the object appearance model is built. This approach has the following limitations: 1) The bases are data driven, so they can be easily corrupted, and 2) it is difficult to robustly update the bases in challenging situations. In this paper, we construct an appearance model using the 3D discrete cosine transform (3D-DCT). The 3D-DCT is based on a set of cosine basis functions which are determined by the dimensions of the 3D signal and thus independent of the input video data. In addition, the 3D-DCT can generate a compact energy spectrum whose high-frequency coefficients are sparse if the appearance samples are similar. By discarding these high-frequency coefficients, we simultaneously obtain a compact 3D-DCT-based object representation and a signal reconstruction-based similarity measure (reflecting the information loss from signal reconstruction). To efficiently update the object representation, we propose an incremental 3D-DCT algorithm which decomposes the 3D-DCT into successive operations of the 2D discrete cosine transform (2D-DCT) and 1D discrete cosine transform (1D-DCT) on the input video data. As a result, the incremental 3D-DCT algorithm only needs to compute the 2D-DCT for newly added frames as well as the 1D-DCT along the third dimension, which significantly reduces the computational complexity. Based on this incremental 3D-DCT algorithm, we design a discriminative criterion to evaluate the likelihood of a test sample belonging to the foreground object. We then embed the discriminative criterion into a particle filtering framework for object state inference over time. Experimental results demonstrate the effectiveness and robustness of the proposed tracker.
Xi Li 0001, Anthony R. Dick, Chunhua Shen, Anton van den Hengel, Hanzi Wang
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Visual Tracking With Spatio-Temporal Dempster-Shafer Information Fusion
abstract
A key problem in visual tracking is how to effectively combine spatio-temporal visual information from throughout a video to accurately estimate the state of an object. We address this problem by incorporating Dempster-Shafer (DS) information fusion into the tracking approach. To implement this fusion task, the entire image sequence is partitioned into spatially and temporally adjacent subsequences. A support vector machine (SVM) classifier is trained for object/nonobject classification on each of these subsequences, the outputs of which act as separate data sources. To combine the discriminative information from these classifiers, we further present a spatio-temporal weighted DS (STWDS) scheme. In addition, temporally adjacent sources are likely to share discriminative information on object/nonobject classification. To use such information, an adaptive SVM learning scheme is designed to transfer discriminative information across sources. Finally, the corresponding DS belief function of the STWDS scheme is embedded into a Bayesian tracking model. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracking approach.
Xi Li 0001, Anthony R. Dick, Chunhua Shen, Zhongfei Zhang, Anton van den Hengel, Hanzi Wang
IEEE Trans. Image Process.1
2013 A survey of appearance models in visual object tracking
abstract
Visual object tracking is a significant computer vision task which can be applied to many domains, such as visual surveillance, human computer interaction, and video compression. Despite extensive research on this topic, it still suffers from difficulties in handling complex object appearance changes caused by factors such as illumination variation, partial occlusion, shape deformation, and camera motion. Therefore, effective modeling of the 2D appearance of tracked objects is a key issue for the success of a visual tracker. In the literature, researchers have proposed a variety of 2D appearance models. To help readers swiftly learn the recent advances in 2D appearance models for visual object tracking, we contribute this survey, which provides a detailed review of the existing 2D appearance models. In particular, this survey takes a module-based architecture that enables readers to easily grasp the key points of visual object tracking. In this survey, we first decompose the problem of appearance modeling into two different processing stages: visual representation and statistical modeling. Then, different 2D appearance models are categorized and discussed with respect to their composition modules. Finally, we address several issues of interest as well as the remaining challenges for future research on this topic. The contributions of this survey are fourfold. First, we review the literature of visual representations according to their feature-construction mechanisms (i.e., local and global). Second, the existing statistical modeling schemes for tracking-by-detection are reviewed according to their model-construction mechanisms: generative, discriminative, and hybrid generative-discriminative. Third, each type of visual representations or statistical modeling techniques is analyzed and discussed from a theoretical or practical viewpoint. Fourth, the existing benchmark resources (e.g., source codes and video datasets) are examined in this survey.
Xi Li 0001, Weiming Hu 0004, Chunhua Shen, Zhongfei Zhang, Anthony R. Dick, Anton van den Hengel
ACM Trans. Intell. Syst. Technol.1
2012 Non-sparse linear representations for visual tracking with online reservoir metric learning
abstract
Most sparse linear representation-based trackers need to solve a computationally expensive li-regularized optimization problem. To address this problem, we propose a visual tracker based on non-sparse linear representations, which admit an efficient closed-form solution without sacrificing accuracy. Moreover, in order to capture the correlation information between different feature dimensions, we learn a Mahalanobis distance metric in an online fashion and incorporate the learned metric into the optimization problem for obtaining the linear representation. We show that online metric learning using proximity comparison significantly improves the robustness of the tracking, especially on those sequences exhibiting drastic appearance changes. Furthermore, in order to prevent the unbounded growth in the number of training samples for the metric learning, we design a time-weighted reservoir sampling method to maintain and update limited-sized foreground and background sample buffers for balancing sample diversity and adaptability. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracker.
Xi Li 0001, Chunhua Shen, Qinfeng Shi, Anthony R. Dick, Anton van den Hengel
CVPR1
2012 Adaptive human silhouette reconstruction based on the exploration of temporal information
abstract
Human silhouette reconstruction has a wide range of applications in motion analysis, object segmentation and tracking, etc. In this paper, we propose a human silhouette reconstruction method based on the exploration of temporal information. Given a test silhouette, the proposed method aims to find its reliable templates for reconstruction by using the intrinsic temporal relationship among different frames. To effectively obtain such templates, we propose an adaptive criterion based on the non-negative least square optimization. Experimental results on two challenging datasets demonstrate the effectiveness of our method.
Xi Li 0001, Tat-Jun Chin, David Suter
ICASSP2
2012 Superpixel-driven level set tracking
abstract
In this paper, we propose a superpixel-driven method for level set tracking. In particular, by taking a superpixel-based speed function, the level set evolution is accelerated greatly. We define a mutual information based speed function using a superpixel-unit as the underlying representation, which captures the correlation of a superpixel with object/background. In order to enhance the robustness of our method, a shape prior is incorporated to constrain the contour evolution. Experimental results on a number of challenging sequences demonstrate the effectiveness and robustness of our method.
Xi Li 0001, Tat-Jun Chin, David Suter
ICIP2
2012 Single and Multiple Object Tracking Using Log-Euclidean Riemannian Subspace and Block-Division Appearance Model
abstract
Object appearance modeling is crucial for tracking objects, especially in videos captured by nonstationary cameras and for reasoning about occlusions between multiple moving objects. Based on the log-euclidean Riemannian metric on symmetric positive definite matrices, we propose an incremental log-euclidean Riemannian subspace learning algorithm in which covariance matrices of image features are mapped into a vector space with the log-euclidean Riemannian metric. Based on the subspace learning algorithm, we develop a log-euclidean block-division appearance model which captures both the global and local spatial layout information about object appearances. Single object tracking and multi-object tracking with occlusion reasoning are then achieved by particle filtering-based Bayesian state inference. During tracking, incremental updating of the log-euclidean block-division appearance model captures changes in object appearance. For multi-object tracking, the appearance models of the objects can be updated even in the presence of occlusions. Experimental results demonstrate that the proposed tracking algorithm obtains more accurate results than six state-of-the-art tracking algorithms.
Weiming Hu 0004, Xi Li 0001, Wenhan Luo, Xiaoqin Zhang 0002, Stephen J. Maybank, Zhongfei Zhang
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Boosting Object Retrieval With Group Queries
abstract
Given a query image of an object, object retrieval aims to return all images from a corpus that depict the same object. Inevitably, the accuracy of the result depends strongly on the quality of the query image. Several measures have been taken to improve retrieval result quality, including the addition of a bounding box to the query, the mining of highly ranked results for more views of the object, and spatial consistency re-ranking. In this letter, we propose a discriminative criterion for improving result quality. This criterion lends itself to the addition of extra query data, and we show that multiple query images can be combined to produce enhanced results. Experiments compare the performance of the method to state-of-the-art in object retrieval, and show how performance is lifted by the inclusion of further query images.
Yanzhi Chen, Xi Li 0001, Anthony R. Dick, Anton van den Hengel
IEEE Signal Process. Lett.2
2011 Superpixel-based object class segmentation using conditional random fields
abstract
Object class segmentation (OCS) is a key issue in semantic scene labeling and understanding. Its general principle consists of naming object entities into scenes according to their intrinsic visual features as well as their dependencies. In this paper, we propose a novel superpixel-based framework for object class segmentation using conditional random fields (CRFs). The framework proceeds in two steps: (i) superpixel label estimate; and (ii) CRF label propagation. Step (i) is achieved using multi-scale boosted classifiers over superpixels and makes it possible to find coarse estimates of initial labels. Fine labeling is afterward achieved in Step (ii), using an anisotropic contrast sensitive pairwise function designed in order to characterize the intrinsic interaction potentials between objects according to 4-neighborhoods. Finally, a higher-order criterion is applied to enforce region label consistency of OCS. Experimental results demonstrate the effectiveness of the proposed framework.
Xi Li 0001, Hichem Sahbi
ICASSP1
2011 Efficient block-division model for robust multiple object tracking
abstract
Tracking multiple objects under occlusion is one of the most challenging issues in computer vision. Occlusion results in mistaken match when finding the most similar candidate. Adapting to the change of objects is essential for tracking as objects often undergo intrinsic changes, but noise is unavoidably introduced during updating of the object, and this further confuses the tracker. In order to address these problems, a block-division appearance model is introduced to efficiently handle occlusion. In this model, spatial information is introduced to avoid the mistaken match between object and candidate. Based on this model, a selective updating strategy is proposed to incrementally learn the change of the object, avoiding introducing noise when updating. At the same time occlusion is deduced by monitoring the variation of each block. Experimental results in various videos validate the effectiveness of our algorithm in tracking multiple objects under occlusion.
Wenhan Luo, Xiaoqin Zhang 0002, Yang Liu 0020, Xi Li 0001, Weiming Hu 0004, Wei Li 0034
ICASSP4
2011 Graph mode-based contextual kernels for robust SVM tracking
abstract
Visual tracking has been typically solved as a binary classification problem. Most existing trackers only consider the pairwise interactions between samples, and thereby ignore the higher-order contextual interactions, which may lead to the sensitivity to complicated factors such as noises, outliers, background clutters and so on. In this paper, we propose a visual tracker based on support vector machines (SVMs), for which a novel graph mode-based contextual kernel is designed to effectively capture the higher-order contextual information from samples. To do so, we first create a visual graph whose similarity matrix is determined by a baseline visual kernel. Second, a set of high-order contexts are discovered in the visual graph. The problem of discovering these high-order contexts is solved by seeking modes of the visual graph. Each graph mode corresponds to a vertex community termed as a high-order context. Third, we construct a contextual kernel that effectively captures the interaction information between the high-order contexts. Finally, this contextual kernel is embedded into SVMs for robust tracking. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracker.
Xi Li 0001, Anthony R. Dick, Hanzi Wang, Chunhua Shen, Anton van den Hengel
ICCV1
2011 Robust visual tracking via transfer learning
abstract
In this paper, we propose a boosting based tracking framework using transfer learning. To deal with complex appearance variations, the proposed tracking framework tries to utilize discriminative information from previous frames to conduct the tracking task in the current frame, and thus transfers some prior knowledge from the previous source data domain to the current target data domain, resulting in a high discriminative tracker for distinguishing the object from the background. The proposed tracking system has been tested on several challenging sequences. Experimental results demonstrate the effectiveness of the proposed tracking framework.
Wenhan Luo, Xi Li 0001, Wei Li 0034, Weiming Hu 0004
ICIP2
2011 Incremental Tensor Subspace Learning and Its Applications to Foreground Segmentation and Tracking
abstract
Appearance modeling is very important for background modeling and object tracking. Subspace learning-based algorithms have been used to model the appearances of objects or scenes. Current vector subspace-based algorithms cannot effectively represent spatial correlations between pixel values. Current tensor subspace-based algorithms construct an offline representation of image ensembles, and current online tensor subspace learning algorithms cannot be applied to background modeling and object tracking. In this paper, we propose an online tensor subspace learning algorithm which models appearance changes by incrementally learning a tensor subspace representation through adaptively updating the sample mean and an eigenbasis for each unfolding matrix of the tensor. The proposed incremental tensor subspace learning algorithm is applied to foreground segmentation and object tracking for grayscale and color image sequences. The new background models capture the intrinsic spatiotemporal characteristics of scenes. The new tracking algorithm captures the appearance characteristics of an object during tracking and uses a particle filter to estimate the optimal object state. Experimental evaluations against state-of-the-art algorithms demonstrate the promise and effectiveness of the proposed incremental tensor subspace learning algorithm, and its applications to foreground segmentation and object tracking.
Weiming Hu 0004, Xi Li 0001, Xiaoqin Zhang 0002, Xinchu Shi, Stephen J. Maybank, Zhongfei Zhang
Int. J. Comput. Vis.2
2011 Visual tracking via dynamic tensor analysis with mean update
Xiaoqin Zhang 0002, Xinchu Shi, Weiming Hu 0004, Xi Li 0001, Stephen J. Maybank
Neurocomputing4
2010 Context-Based Support Vector Machines for Interconnected Image Annotation
Hichem Sahbi, Xi Li 0001
ACCV (1)2
2010 Spatio-Temporal Proximity Distribution Kernels for Action Recognition
Chunfeng Yuan, Weiming Hu 0004, Hanzi Wang, Xi Li 0001, Nianhua Xie
ICASSP4
2010 Semi-supervised Trajectory Learning Using a Multi-Scale Key Point Based Trajectory Representation
abstract
Motion trajectories contain rich high-level semantic information such as object behaviors and gestures, which can be effectively captured by supervised trajectory learning. However, it is usually a tough task to obtain a large number of high-quality manually labeled samples in real applications. Thus, how to perform trajectory learning in small training sample size situations is an important research topic. In this paper, we propose a trajectory learning framework using graph-based semi-supervised transductive learning, which propagates training sample labels along a particular graph. Furthermore, a novel trajectory descriptor based on multi-scale key points is proposed to characterize the spatial structural information. Experimental results demonstrate effectiveness of our framework.
Yang Liu 0020, Xi Li 0001, Weiming Hu 0004
ICPR2
2010 Context dependent SVMs for interconnected image network annotation
abstract
The exponential growth of interconnected networks, such as Flickr, currently makes them the standard way to share and explore data where users put contents and refer to others. These interconnections create valuable information in order to enhance the performance of many tasks in information retrieval including ranking and annotation. We introduce in this paper a novel image annotation framework based on support vector machines (SVMs) and a new class of kernels referred to as context-dependent. The method goes beyond the naive use of the intrinsic low level features (such as color, texture, shape, etc.) and context-free kernels, in order to design a kernel function applicable to interconnected databases such as social networks. The main contribution of our method includes a variational framework which helps designing this function using both intrinsic features and the underlying contextual information. This function also converges to a positive definite fixed-point, usable for SVM training and other kernel methods. When plugged in SVMs, our context-dependent kernel consistently improves the performance of image annotation, compared to context-free kernels, on hundreds of thousands of Flickr images.
Hichem Sahbi, Xi Li 0001
ACM Multimedia2
2010 Linear discriminant analysis using rotational invariant L1 norm
Xi Li 0001, Weiming Hu 0004, Hanzi Wang, Zhongfei Zhang
Neurocomputing1
2010 Robust object tracking using a spatial pyramid heat kernel structural information representation
Xi Li 0001, Weiming Hu 0004, Hanzi Wang, Zhongfei Zhang
Neurocomputing1
2010 Heat Kernel Based Local Binary Pattern for Face Representation
abstract
Face classification has recently become a very hot research topic in computer vision and multimedia information processing. It has many potential applications, in which face representation is the most fundamental task. Most existing face representation methods perform poorly in capturing the intrinsic structural information of face appearance. To address this problem, we propose a novel multiscale heat kernel based face representation, for heat kernels perform well in characterizing the topological structural information of face appearance. Further, the local binary pattern (LBP) descriptor is incorporated into the multiscale heat kernel face representation for the purpose of capturing texture information of face appearance. As a result, we have the heat kernel based local binary pattern (HKLBP) descriptor. Finally, a Support Vector Machine (SVM) classifier is learned in theHKLBPfeature space for face classification. Experimental results demonstrate the effectiveness and superiority of our face classification framework.
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Hanzi Wang
IEEE Signal Process. Lett.1
2009 Spectral Graph Partitioning Based on a Random Walk Diffusion Similarity Measure
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Yang Liu 0020
ACCV (2)1
2009 Human Action Recognition under Log-Euclidean Riemannian Metric
Chunfeng Yuan, Weiming Hu 0004, Xi Li 0001, Stephen J. Maybank, Guan Luo
ACCV (1)3
2009 Human Action Recognition Using Pyramid Vocabulary Tree
Chunfeng Yuan, Xi Li 0001, Weiming Hu 0004, Hanzi Wang
ACCV (3)2
2009 Segment Model Based Vehicle Motion Analysis
abstract
Motion analysis is a very attractive research direction in computer vision field. In this paper, we propose a framework for analyzing real vehicle motion in visual traffic surveillance by using Segment Model (SM), which is a kind of probabilistic model. SM can grasp the underlying information of observation sequence by using segment distribution. It has been proved to be more precise than that of HMM. In the experiments, we compare our approach with the template matching method based on the Hausdorff distance and the state space method based on the Hidden Markov Model (HMM). The experimental results show the effectiveness of our approach.
Pengfei Zhu 0001, Weiming Hu 0004, Xi Li 0001, Li Li 0010
AVSS3
2009 Image spam filtering using Fourier-Mellin invariant features
abstract
Image spam is a new obfuscating method which spammers invented to more effectively bypass conventional text based spam filters. In this paper, a framework for filtering image spams by using the Fourier-Mellin invariant features is described. Fourier-Mellin features are robust for most kinds of image spam variations. A one-class classifier, the support vector data description (SVDD), is exploited to model the boundary of image spam class in the feature space without using information of legitimate emails. Experimental results demonstrate that our framework is effective for fighting image spam.
Haiqiang Zuo, Xi Li 0001, Ou Wu 0001, Weiming Hu 0004, Guan Luo
ICASSP2
2009 A Boosted Semi-supervised Learning Framework for Web Page Filtering
abstract
The World Wide Web provides great convenience for users to obtain information. However, there exists much harmful information on the Internet, such as pornographic content and prohibited drugs' information. Thus, how to filter harmful Web pages on the Internet is quite an important issue. In general, the problem of harmful Web page filtering is converted to that of Web page classification, which needs plenty of well labeled training samples. However, the cost of labeling a large set of Web pages is very expensive. To address this problem, we adopt a semi-supervised framework for Web page filtering. In this framework, each Web page is represented by bags of different features, extracted using its HTML structure. Then a semi-supervised learning strategy is taken for efficiently obtaining well labeled training samples. Finally, a boosting classifier is utilized for harmful Web page filtering. Experiments have demonstrated the effectiveness of our framework.
Zhu He, Xi Li 0001, Weiming Hu 0004
SMC2
2009 Adaptive Distributed Intrusion Detection Using Parametric Model
abstract
Due to the increasing demands for network security, distributed intrusion detection has become a hot research topic in computer science. However, the design and maintenance of the intrusion detection system (IDS) is still a challenging task due to its dynamic, scalability, and privacy properties. In this paper, we propose a distributed IDS framework which consists of the individual and global models. Specifically, the individual model for the local unit derives from Gaussian Mixture Model based on online Adaboost algorithm, while the global model is constructed through the PSO-SVM fusion algorithm. Experimental results demonstrate that our approach can achieve a good detection performance while being trained online and consuming little traffic to communicate between local units.
Weiming Hu 0004, Xiaoqin Zhang 0002, Xi Li 0001
Web Intelligence4
2008 Trajectory-Based Video Retrieval Using Dirichlet Process Mixture Models
abstract
In this paper, we present a trajectory-based video retrieval framework using Dirichlet process mixture models. The main contribution of this framework is four-fold. (1) We apply a Dirichlet process mixture model (DPMM) to unsupervised trajectory learning. DPMM is a countably infinite mixture model with its components growing by itself. (2) We employ a time-sensitive Dirichlet process mixture model (tDPMM) to learn trajectories ’ time-series characteristics. Furthermore, a novel likelihood estimation algorithm for tDPMM is proposed for the first time. (3) We develop a tDPMM-based probabilistic model matching scheme, which is empirically shown to be more error-tolerating and is able to deliver higher retrieval accuracy than the peer methods in the literature. (4) The framework has a nice scalability and adaptability in the sense that when new cluster data are presented, the framework automatically identifies the new cluster information without having to redo the training. Theoretic analysis and experimental evaluations against the state-of-the-art methods demonstrate the promise and effectiveness of the framework. 1
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Guan Luo
BMVC1
2008 Visual tracking via incremental Log-Euclidean Riemannian subspace learning
abstract
Recently, a novel Log-Euclidean Riemannian metric is proposed for statistics on symmetric positive definite (SPD) matrices. Under this metric, distances and Riemannian means take a much simpler form than the widely used affine-invariant Riemannian metric. Based on the Log-Euclidean Riemannian metric, we develop a tracking framework in this paper. In the framework, the covariance matrices of image features in the five modes are used to represent object appearance. Since a nonsingular covariance matrix is a SPD matrix lying on a connected Riemannian manifold, the Log-Euclidean Riemannian metric is used for statistics on the covariance matrices of image features. Further, we present an effective online Log-Euclidean Riemannian subspace learning algorithm which models the appearance changes of an object by incrementally learning a low-order Log-Euclidean eigenspace representation through adaptively updating the sample mean and eigenbasis. Tracking is then led by the Bayesian state inference framework in which a particle filter is used for propagating sample distributions over the time. Theoretic analysis and experimental evaluations demonstrate the promise and effectiveness of the proposed framework.
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Mingliang Zhu, Jian Cheng 0002
CVPR1
2008 Sequential particle swarm optimization for visual tracking
abstract
Visual tracking usually involves an optimization process for estimating the motion of an object from measured images in a video sequence. In this paper, a new evolutionary approach, PSO (particle swarm optimization), is adopted for visual tracking. Since the tracking process is a dynamic optimization problem which is simultaneously influenced by the object state and the time, we propose a sequential particle swarm optimization framework by incorporating the temporal continuity information into the traditional PSO algorithm. In addition, the parameters in PSO are changed adaptively according to the fitness values of particles and the predicted motion of the tracked object, leading to a favourable performance in tracking applications. Furthermore, we show theoretically that, in a Bayesian inference view, the sequential PSO framework is in essence a multilayer importance sampling based particle filter. Experimental results demonstrate that, compared with the state-of-the-art particle filter and its variation - the unscented particle filter, the proposed tracking algorithm is more robust and effective, especially when the object has an arbitrary motion or undergoes large appearance changes.
Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank, Xi Li 0001, Mingliang Zhu
CVPR4
2008 Robust Visual Tracking Based on an Effective Appearance Model
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002
ECCV (4)1
2008 Level set tracking with dynamical shape priors
abstract
Dynamical shape priors are curical for level set-based non- rigid object tracking with noise, occlusions or background clutter. In this paper, we propose a level set tracking framework using dynamical shape priors to capture contours changes of an object in a periodic action sequence. The framework consists of two stages - off-line training and on-line tracking. During the off-line training stage, a graph- based dominant set clustering (DSC) method is applied to learn a shape codebook with each codeword representing a certain shape mode. Then a codeword transition matrix is learnt to characterize the temporal correlations of contours of an object. During the on-line tracking stage, we fuse the knowledge of shape priors and current observations, and adopt maximum a posteriori (MAP) estimation to predict the current shape mode. The experimental results on synthetic and real video sequences demonstrate the effectiveness of our method.
Xi Li 0001, Weiming Hu 0004
ICIP2
2008 Multiclass spectral clustering based on discriminant analysis
abstract
Many existing spectral clustering algorithms share a conventional graph partitioning criterion: normalized cuts (NC). However, one problem with NC is that it poorly captures the graph¿s local marginal information which is very important to graph-based clustering. In this paper, we present a discriminant analysis based graph partitioning criterion (DAC), which is designed to effectively capture the graph¿s local marginal information characterized by the intra-class compactness and the inter-class separability. DAC preserves the intrinsic topological structures of the similarity graph on data points by constructing a k-nearest neighboring subgraph for each data point. Consequently, the clustering results generated by the DAC-based clustering algorithm (DACA) are robust to the outlier disturbance. Theoretic analysis and experimental evaluations demonstrate the promise and effectiveness of DACA.
Xi Li 0001, Zhongfei Zhang, Yanguo Wang, Weiming Hu 0004
ICPR1
2008 Boosted cannabis image recognition
abstract
With the large number of Web sites promoting the use of illicit drugs, it has become important to screen these sites for the protection of children on the Internet. Conventional keyword-based approaches are not sufficient because these Web sites often have lots of images and little meaningful words than prices. We propose an AdaBoost-based algorithm for cannabis image recognition. This is the first known attempt at computerized detection of illicit drug Web contents using images. The main technical contributions of our work are two-fold. First, we introduce a novel weak classifier which considers the inherently structural property or ldquoself-similarityrdquo of the cannabis plants. The self-correlation structural characteristics of cannabis can be used as a discriminative property for the purpose of cannabis image recognition. Second, we propose a rapid weak classifier finder, which can efficiently select discriminative weak classifiers from the weak classifier space with little degradation to the classification accuracy. Experiments on real world images have demonstrated improved performance of our method over other methods.
Nianhua Xie, Xi Li 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, James Z. Wang 0001
ICPR2
2008 SVD based Kalman particle filter for robust visual tracking
abstract
Object tracking is one of the most important tasks in computer vision. The unscented particle filter algorithm has been extensively used to tackle this problem and achieved a great success, because it uses the UKF (unscented Kalman filter) to generate a sophisticated proposal distributions which incorporates the newest observations into the state transition distribution and thus overcomes the sample impoverishment problem suffered by the particle filter. However, UKF often encounters the ill-conditioned problem when solving the square root of the covariance matrix in practice. In this paper, we propose a novel Kalman particle filter based on SVD (singular value decomposition), and apply it for visual tracking. Experimental results demonstrate that, compared with the particle filter and the unscented particle filter, the proposed algorithm is more robust in tracking performance.
Xiaoqin Zhang 0002, Weiming Hu 0004, Zixiang Zhao, Yanguo Wang, Xi Li 0001, Qingdi Wei
ICPR5
2008 Distributed detection of network intrusions based on a parametric model
abstract
With the increasing requirements of fast response and privacy protection, how to detect network intrusions in a distributed architecture becomes a hot research area in the development of modern information security systems. However, it is a challenge to build such a system, given the difficulties brought by the mixed-attribute property of network connection data and the constraints on network communication. In this paper, we present a framework for distributed detection of network intrusions based on a parametric model. The parametric model can explicitly reflect the distributions of different intrusion types and handle the mixed-attribute data naturally. Based on the model, we can generate an accurate global intrusion detector with a very low cost of communication among the distributed detection sites, and no sharing of original network data is needed. Experimental results demonstrate the advantages of the proposed framework in the distributed intrusion detection application.
Yanguo Wang, Xi Li 0001, Weiming Hu 0004
SMC2
2008 User oriented link function classification
abstract
Currently most link-related applications treat all links in the same web page to be identical. One link-related application usually requires one certain property of hyperlinks but actually not all links have this property or they have this property on different levels. Based on a study of how human users judge the links, the idea of the link function classification (LFC) is introduced in this paper. The link functions reflect the purpose that links are created by web page designers and the way they are used by viewers. Links in a certain function class imply one certain relationship between the adjacent pages, and thus they can be assumed to have similar properties. An algorithm is proposed to analyze the link functions based on both vision and structure features which simulates the reaction on the links of human users. Current applications can be enhanced by LFC with a more accurate modeling of the web graph. New mining methods can be also developed by making more and stronger assumptions on links within each function class due to the purer property set they share.
Mingliang Zhu, Weiming Hu 0004, Ou Wu 0001, Xi Li 0001, Xiaoqin Zhang 0002
WWW4
2007 Robust Visual Tracking Based on Incremental Tensor Subspace Learning
abstract
Most existing subspace analysis-based tracking algorithms utilize a flattened vector to represent a target, resulting in a high dimensional data learning problem. Recently, subspace analysis is incorporated into the multilinear framework which offline constructs a representation of image ensembles using high-order tensors. This reduces spatio-temporal redundancies substantially, whereas the computational and memory cost is high. In this paper, we present an effective online tensor subspace learning algorithm which models the appearance changes of a target by incrementally learning a low-order tensor eigenspace representation through adaptively updating the sample mean and eigenbasis. Tracking then is led by the state inference within the framework in which a particle filter is used for propagating sample distributions over the time. A novel likelihood function, based on the tensor reconstruction error norm, is developed to measure the similarity between the test image and the learned tensor subspace model during the tracking. Theoretic analysis and experimental evaluations against a state-of-the-art method demonstrate the promise and effectiveness of this algorithm.
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Guan Luo
ICCV1
2007 Graph Based Discriminative Learning for Robust and Efficient Object Tracking
abstract
Object tracking is viewed as a two-class 'one-versus-rest' classification problem, in which the sample distribution of the target is approximately Gausian while the background samples are often multimodal. Based on these special properties, we propose a graph embedding based discriminative learning method, in which the topology structures of graphs are carefully designed to reflect the properties of the sample distributions. This method can simultaneously learn the subspace of the target and its local discriminative structure against the background. Moreover, a heuristic negative sample selection scheme is adopted to make the classification more effective. In tracking procedure, the graph based learning is embedded into a Bayesian inference framework cascaded with hierarchical motion estimation, which significantly improves the accuracy and efficiency of the localization. Furthermore, an incremental updating technique for the graphs is developed to capture the changes in both appearance and illumination. Experimental results demonstrate that, compared with two state-of-the-art methods, the proposed tracking algorithm is more efficient and effective, especially in dynamically changing and clutter scenes.
Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank, Xi Li 0001
ICCV4
2007 Corner Detection of Contour Images using Spectral Clustering
abstract
Corner detection plays an important role in object recognition and motion analysis. In this paper, we propose a hierarchical corner detection framework based on spectral clustering (SC). The framework consists of three stages: contour smoothing, corner cell extraction and corner localization. In the contour smoothing stage, wavelet decomposition is imposed on the raw contour to reduce noise. In the corner cell extraction stage, several atomic corner cells are obtained by SC. In the corner localization stage, the corner points of each corner cell are located by the corner locator based on the kernel-weighted cosine curvature measure. Experimental results demonstrate the superiority of our framework.
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang
ICIP (3)1
2007 Customizable Instance-Driven Webpage Filtering Based on Semi-Supervised Learning
abstract
The World Wide Web has been growing rapidly in recent years, along with increasing needs for content-based Webpage filtering. But most existing filtering systems cannot easily satisfy the personalized filtering demands from different users at the same time. In this paper, a customizable instance-driven Webpage filtering strategy is proposed. For different users, different Webpage filters are produced by our system through mining the certain Webpage classes they focus on. A semi-supervised learning (SSL) approach is applied for obtaining a precise description of the Webpage class which a user wants to filter based on the small sized user instance set he or she provided. Subsequently, a feature selection step is performed and a Bayes classifier is created over the enlarged training set. Experimental results show the great stability and high performance of our proposed method, and it outperforms existing methods.
Mingliang Zhu, Weiming Hu 0004, Xi Li 0001, Ou Wu 0001
Web Intelligence3