EDBT 2026 Demo / reviewers in the wild / expert
Jinjun Wang
dblp:87/4858
· DBLP profile ↗
110ranked-venue papers
24as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 65 · 19 first-author · 13 since 2021Artificial intelligence and machine learning · 52 · 8 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical frequency adaptation for all-in-one image restoration
Yang Wu 0001, Ye Deng 0005, Siqi Hui, Yuhan Liu 0006, Kangyi Wu, Wenli Huang 0004, Jinjun Wang |
Knowl. Based Syst. | 7 |
| 2026 | Frequency-guided generalizable representation learning for cross-domain few-shot learning
Siqi Hui, Sanping Zhou, Ye Deng 0005, Wenli Huang 0004, Yang Wu 0001, Jinjun Wang |
Pattern Recognit. | 6 |
| 2025 | Event-Equalized Dense Video CaptioningabstractDense video captioning aims to localize and caption all events in arbitrary untrimmed videos. Although previous methods have achieved appealing results, they still face the issue of temporal bias, i.e, models tend to focus more on events with certain temporal characteristics. Specifically, 1) the temporal distribution of events in training datasets is uneven. Models trained on these datasets will pay less attention to out-of-distribution events. 2) long-duration events have more frame features than short ones and will attract more attention. To address this, we argue that events, with varying temporal characteristics, should be treated equally when it comes to dense video captioning. Intuitively, different events tend to have distinct visual differences due to varied camera views, backgrounds, or subjects. Inspired by that, we intend to utilize visual features to have an approximate perception of possible events and pay equal attention to them. In this paper, we introduce a simple but effective framework, called Event-Equalized Dense Video Captioning (E2DVC) to overcome the temporal bias and treat all possible events equally. Experimental results on ActivityNet Captions and YouCook2 dataset validate the effectiveness of the proposed methods and show State-of-the-art (SOTA) performance on dense video captioning. Kangyi Wu, Pengna Li, Jingwen Fu, Yang Wu 0001, Yuhan Liu 0006, Jinjun Wang, Sanping Zhou |
CVPR | 7 |
| 2025 | HomoMamba for Self-supervised Homography Estimation
Jiayong Zhong, Jinjun Wang, Kaifan Hou |
ICIC (6) | 2 |
| 2025 | Auxiliary Loss Reweighting for Image InpaintingabstractImage inpainting aims to reconstruct missing regions in corrupted images with semantically consistent content. While modern methods employ perceptual and style losses to enhance inpainting quality by supervising deep feature representations, two key challenges persist: (i) existing approaches necessitate time-consuming grid searches to determine optimal loss weights, and (ii) heterogeneous auxiliary loss terms are assigned fixed weights, limiting their adaptive contributions. To address these limitations, we propose a framework featuring dynamically weighted auxiliary losses and an automated weight adaptation mechanism. Specifically, we introduce Tunable Perceptual Loss (TPL) and Tunable Style Loss (TSL), which generalize traditional perceptual and style losses by incorporating tunable weights that independently scale distinct loss components according to their auxiliary potential. These are optimized via our Adaptive Weight Adjustment (AWA) algorithm, which dynamically reweights TPL and TSL during training by prioritizing loss terms that maximally improve inpainting performance. Empirical evaluations on public datasets demonstrate that our framework enhances state-of-the-art inpainting performance while eliminating manual weight tuning. Wenli Huang 0004, Siqi Hui, Ye Deng 0005, Xiaomeng Xin, Yang Wu 0001, Jinjun Wang |
IECON | 6 |
| 2025 | Event-Frame Temporal Information Fusion for Visual Object TrackingabstractIn recent years, researchers have explored integrating event-based data into visual tracking, often using event representations such as voxel grids or event frames. These methods have demonstrated the potential of event data to improve tracking accuracy in dynamic scenes. However, most existing methods fail to effectively utilize the temporal continuity and inter-frame relationships within event streams. By treating event streams as discrete data, these methods overlook the rich contextual and temporal information within the events, limiting their performance in scenarios where temporal dynamics are critical. To overcome these limitations, this paper proposes a novel framework that better leverages the temporal relationships within event streams while integrating complementary features from frame-based data. This method includes two key modules: the Multimodal Temporal Transformer Module (MTTM) and the Feature Transformation Module (FTM). These modules aim to address the core challenge of integrating event-based and frame-based data for robust tracking, leveraging the strengths of both modalities while overcoming their inherent limitations. Moreover, our framework supports multiple event representation formats, including voxel grids and event frames. This flexibility allows our approach to adapt to different computational requirements and application scenarios. Extensive experiments show that our method achieves state-of-the-art performance on benchmark datasets, demonstrating its ability to handle challenging scenarios such as fast object motion, occlusion, and dynamic environments. Kaifan Hou, Jinjun Wang, Jiayong Zhong |
IJCNN | 3 |
| 2025 | Meta channel masking for cross-domain few-shot image classificationabstractCross-domain Few-shot Learning (CD-FSL) aims to address the challenges of FSL where significant domain gaps exist between source and target image datasets. Unlike many existing CD-FSL methods that utilize an auxiliary target dataset with a few labeled target images to enhance model generalization, our approach directly tackles the limitations imposed by the reliance on source-specific knowledge. We observe that models trained on unbalanced datasets tend to overfit to source-specific features, which, while effective in the source domain, generalize poorly to the target image domain. To address this, we introduce a novel dropout-based framework named Meta Channel Masking (MCM). This framework attenuates the learning of model channels on the source domain by dynamically masking source feature channels during training. In contrast to traditional dropout techniques that manually set masking probabilities based on statistical assumptions about the source data, our MCM framework employs a meta-learning process that automatically adjusts channel mask probabilities. This adjustment is informed by auxiliary target data, effectively minimizing few-shot loss on the auxiliary target dataset and thereby enhancing the model’s generalization capabilities in the target domain. Our extensive experiments across various image classification benchmark datasets demonstrate that our framework outperforms state-of-the-art methods. Siqi Hui, Sanping Zhou, Ye Deng 0005, Pengna Li, Jinjun Wang |
Neurocomputing | 5 |
| 2025 | An Open-Set Domain Adaptation Framework for Hyperspectral Image Classification With Pixel-Aware Weighting and Decoupled AlignmentabstractRecent studies have shown that deep domain adaptation techniques perform excellently in cross-domain hyperspectral image classification. However, these methods typically assume that the source domain and the target domain share the same class set, while in practice, the target domain may include unknown classes, and direct alignment can result in negative transfer. Moreover, in hyperspectral image classification based on deep learning, using the label of the central pixel to represent the label of the image patch may lead to feature bias due to the uncertainty of the labels of neighboring pixels, thereby reducing the generalization performance of the model. To address this, this paper proposes an open-set domain adaptation framework, including a Pixel-Aware Weight Learning (PAWL) module and a Decoupled Dual Alignment (DDA) strategy. The PAWL module effectively reduces the feature bias caused by inconsistency in neighboring pixel labels by analyzing the uncertainty of neighboring pixel labels and utilizing adaptive weight learning, thereby improving recognition performance in open-set environments. The DDA strategy decouples the features of the source domain and target domain into known and unknown classes and aligns them separately to mitigate negative transfer. Experiments on two cross-scene hyperspectral datasets validated the effectiveness of the method. Zhaokui Li, Mingtai Qi, Yan Wang 0087, Xuewei Gong, Cuiwei Liu, Jinjun Wang |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | Semi-independent Convolution for Image InpaintingabstractIn typical image inpainting tasks, the locations and shapes of damaged or masked areas are often random and irregular. Vanilla convolutions, commonly employed in learning-based inpainting models, treat all spatial features as valid and share parameters across different regions. This approach can struggle with irregular damage patterns, leading to inpainted results that may suffer from color discrepancies and blurriness. In this paper, we introduce a novel operator known as Semi-Independent Convolution (SIConv) to tackle this challenge. The proposed SIConv, on top of the regular convolution with shared weights, also introduces dynamic terms that assign their own independent weights to each part of the image, and the overall computation is formulated as a shared convolution parameter with an additional term to describe the local structure. Qualitative and quantitative experiments demonstrate that our method outperforms the state-of-the-art, yielding clearer, more coherent, and visually convincing inpainting results. Wenli Huang 0004, Ye Deng 0005, Xiaomeng Xin, Jinbao He, Jinjun Wang |
IECON | 6 |
| 2024 | Pruning CNN based on Combinational Filter DeletionabstractNetwork pruning is a figurative model compression technique designed to lighten and accelerate neural network models. Most existing pruning methods prioritize the selection of filters by their importance or apply regularization based on the properties of individual filters, neglecting the internal connections within combinations of multiple filters. This work introduces a pruning method termed Combinational Filter Deletion (CFD), which incorporates a straightforward yet effective evaluation metric based on the diversity of filter combination distributions to reveal the characteristics inherent to multiple filter interactions. CFD enables the exploration of an expanded search space, offering a greater array of choices and leveraging the intrinsic information of conventional layers. Moreover, this method is both general and non-exclusive, capable of enhancing the efficacy of other single-filter-based pruning techniques. Xiujie Wang, Shuai Sui, Jinjun Wang |
IECON | 6 |
| 2024 | Semantic-aware Representation Learning for Homography EstimationabstractHomography estimation is the task of determining the transformation from an image pair. Our approach focuses on employing detector-free feature matching methods to address this issue. Previous work has underscored the importance of incorporating semantic information, however there still lacks an efficient way to utilize semantic information. Previous methods suffer from treating the semantics as a pre-processing, causing the utilization of semantics overly coarse-grained and lack adaptability when dealing with different tasks. In our work, we seek another way to use the semantic information, that is semantic-aware feature representation learning framework. Based on this, we propose SRMatcher, a new detector-free feature matching method, which encourages the network to learn integrated semantic feature representation. Specifically, to capture precise and rich semantics, we leverage the capabilities of recently popularized vision foundation models (VFMs) trained on extensive datasets. Then, a cross-images Semantic-aware Fusion Block (SFB) is proposed to integrate its fine-grained semantic features into the feature representation space. In this way, by reducing errors stemming from semantic inconsistencies in matching pairs, our proposed SRMatcher is able to deliver more accurate and realistic outcomes. Extensive experiments show that SRMatcher surpasses solid baselines and attains SOTA results on multiple real-world datasets. Compared to the previous SOTA approach GeoFormer, SRMatcher increases the area under the cumulative curve (AUC) by about 11% on HPatches. Additionally, the SRMatcher could serve as a plug-and-play framework for other matching methods like LoFTR, yielding substantial precision improvement. Yuhan Liu 0006, Qianxin Huang, Siqi Hui, Jingwen Fu, Sanping Zhou, Kangyi Wu, Pengna Li, Jinjun Wang |
ACM Multimedia | 8 |
| 2024 | Gradient-guided channel masking for cross-domain few-shot learning
Siqi Hui, Sanping Zhou, Ye Deng 0005, Yang Wu 0001, Jinjun Wang |
Knowl. Based Syst. | 5 |
| 2024 | Sparse self-attention transformer for image inpainting
Wenli Huang 0004, Ye Deng 0005, Siqi Hui, Yang Wu 0001, Sanping Zhou, Jinjun Wang |
Pattern Recognit. | 6 |
| 2024 | Bidirectional feature learning network for RGB-D salient object detectionabstractRGB-D salient object detection aims to perform the pixel-wise localization of salient objects from both RGB and depth images, whose challenge mainly comes from how to learn complementary features from each modality. Existing works often use increasingly large models for performance enhancement, which need large memory and time consumption in practice. In this paper, we propose a simple yet effective B idirectional F eature L earning Net work (BFLNet) for RGB-D salient object detection under limited memory and time conditions. To achieve accurate performance with lightweight backbone networks , an effective B idirectional F eature F usion (BFF) module is designed to merge features from both RGB and depth streams, in which the cross-modal fusions and cross-scale fusions are jointly conducted to fuse the immediate features in multiple scales and multiple modals. What is more, a simple D ual C onsistency L oss (DCL) function is designed to prompt cross-modal fusion by keeping the consistency between cross-modal target predictions. Extensive experiments on four benchmark datasets demonstrate that our method has achieved the state-of-the-art performance with high efficiency in RGB-D salient object detection. Code will be available at https://github.com/nightsky-nostar/BFLNet . Ye Niu, Sanping Zhou, Yonghao Dong, Le Wang 0003, Jinjun Wang, Nanning Zheng 0001 |
Pattern Recognit. | 5 |
| 2024 | Attentive Contextual Attention for Cloud RemovalabstractCloud cover can significantly hinder the use of remote sensing images for Earth observation, prompting urgent advancements in cloud removal technology. Recently, deep learning strategies, especially convolutional neural networks (CNNs) with attention mechanisms, have shown strong potential in restoring cloud-obscured areas. These methods utilize convolution to extract intricate local features and attention mechanisms to gather long-range information, improving the overall comprehension of the scene. However, a common drawback of these approaches is that the resulting images often suffer from blurriness, artifacts, and inconsistencies. This is partly because attention mechanisms apply weights to all features based on generalized similarity scores, which can inadvertently introduce noise and irrelevant details from cloud-covered areas. To overcome this limitation and better capture relevant distant context, we introduce a novel approach named attentive contextual attention (AC-Attention). This method enhances conventional attention mechanisms by dynamically learning data-driven attentive selection scores, enabling it to filter out noise and irrelevant features effectively. By integrating the AC-Attention module into the DSen2-CR cloud removal framework, we significantly improve the model’s ability to capture essential distant information, leading to more effective cloud removal. Our extensive evaluation of various datasets shows that our method outperforms existing ones regarding image reconstruction quality. In addition, we conducted ablation studies by integrating AC-Attention into multiple existing methods and widely used network architectures. These studies demonstrate the effectiveness and adaptability of AC-Attention and reveal its ability to focus on relevant features, thereby improving the overall performance of the networks. The code is available athttps://github.com/huangwenwenlili/ACA-CRNet. Wenli Huang 0004, Ye Deng 0005, Yang Wu 0001, Jinjun Wang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | CR-former: Single-Image Cloud Removal With Focused Taylor AttentionabstractCloud removal aims to restore high-quality images from cloud-contaminated captures, which is essential in remote sensing applications. Effectively modeling the long-range relationships between image features is key to achieving high-quality cloud-free images. While self-attention mechanisms excel at modeling long-distance relationships, their computational complexity scales quadratically with image resolution, limiting their applicability to high-resolution remote sensing images. Current cloud removal methods have mitigated this issue by restricting the global receptive field to smaller regions or adopting channel attention to model long-range relationships. However, these methods either compromise pixel-level long-range dependencies or lose spatial information, potentially leading to structural inconsistencies in restored images. In this work, we propose the focused Taylor attention (FT-Attention), which captures pixel-level long-range relationships without limiting the spatial extent of attention and achieves the$\mathcal {O}(N)$computational complexity, where N represents the image resolution. Specifically, we utilize Taylor series expansions to reduce the computational complexity of the attention mechanism from$\mathcal {O}(N^{2})$to$\mathcal {O}(N)$, enabling efficient capture of pixel relationships directly in high-resolution images. Additionally, to fully leverage the informative pixel, we develop a new normalization function for the query and key, which produces more distinguishable attention weights, enhancing focus on important features. Building on FT-Attention, we design a U-net style network, termed the CR-former, specifically for cloud removal. Extensive experimental results on representative cloud removal datasets demonstrate the superior performance of our CR-former. The code is available athttps://github.com/wuyang2691/CR-former. Yang Wu 0001, Ye Deng 0005, Sanping Zhou, Yuhan Liu 0006, Wenli Huang 0004, Jinjun Wang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Single-Shot and Multi-Shot Feature Learning for Multi-Object TrackingabstractMulti-Object Tracking (MOT) remains a vital component of intelligent video analysis, which aims to locate targets and maintain a consistent identity for each target throughout a video sequence. Existing works usually learn a discriminative feature representation, such as motion and appearance, to associate the detections across frames, which are easily affected by mutual occlusion and background clutter in practice. In this paper, we propose a simple yet effective two-stage feature learning paradigm to jointly learn single-shot and multi-shot features for different targets, so as to achieve robust data association in the tracking process. For the detections without being associated, we design a novel single-shot feature learning module to extract discriminative features of each detection, which can efficiently associate targets between adjacent frames. For the tracklets being lost several frames, we design a novel multi-shot feature learning module to extract discriminative features of each tracklet, which can accurately refind these lost targets after a long period. Once equipped with a simple data association logic, the resulting VisualTracker can perform robust MOT based on the single-shot and multi-shot feature representations. Extensive experimental results demonstrate that our method has achieved significant improvements on MOT17 and MOT20 datasets while reaching state-of-the-art performance on DanceTrack dataset. Sanping Zhou, Le Wang 0003, Jinjun Wang, Nanning Zheng 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Inverse Adversarial Diversity Learning for Network EnsembleabstractNetwork ensemble aims to obtain better results by aggregating the predictions of multiple weak networks, in which how to keep the diversity of different networks plays a critical role in the training process. Many existing approaches keep this kind of diversity either by simply using different network initializations or data partitions, which often requires repeated attempts to pursue a relatively high performance. In this article, we propose a novel inverse adversarial diversity learning (IADL) method to learn a simple yet effective ensemble regime, which can be easily implemented in the following two steps. First, we take each weak network as a generator and design a discriminator to judge the difference between the features extracted by different weak networks. Second, we present an inverse adversarial diversity constraint to push the discriminator to cheat generators that all the resulting features of the same image are too similar to distinguish each other. As a result, diverse features will be extracted by these weak networks through a min-max optimization. What is more, our method can be applied to a variety of tasks, such as image classification and image retrieval, by applying a multitask learning objective function to train all these weak networks in an end-to-end manner. We conduct extensive experiments on the CIFAR-10, CIFAR-100, CUB200-2011, and CARS196 datasets, in which the results show that our method significantly outperforms most of the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Le Wang 0003, Xingyu Wan, Siqi Hui, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Pseudo Labels Refinement with Intra-Camera Similarity for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) aims to retrieve person images across cameras without any identity labels. Most clustering-based methods roughly divide image features into clusters and neglect the feature distribution noise caused by domain shifts among different cameras, leading to inevitable performance degradation. To address this challenge, we propose a novel label refinement framework with clustering intra-camera similarity. Intra-camera feature distribution pays more attention to the appearance of pedestrians and labels are more reliable. We conduct intra-camera training to get local clusters in each camera, respectively, and refine inter-camera clusters with local results. We hence train the Re-ID model with refined reliable pseudo labels in a self-paced way. Extensive experiments demonstrate that the proposed method surpasses state-of-the-art performance. Code is available at https://github.com/leeBooMla/ICSR. Pengna Li, Kangyi Wu, Sanping Zhou, Qianxin Huang, Jinjun Wang |
ICIP | 5 |
| 2023 | Context Adaptive Network for Image InpaintingabstractIn a typical image inpainting task, the location and shape of the damaged or masked area is often random and irregular. The vanilla convolutions widely used in learning-based inpainting models treat all spatial features as valid and share parameters across regions, making it difficult for them to cope with those irregular damages, and models tend to produce inpainting results with color discrepancy and blurriness. In this paper, we propose a novel Context Adaptive Network (CANet) to address this issue. The main idea of the proposed CANet is able to generate different weights depending on the miscellaneous input, which may help to complement images with multiple broken forms in a flexible way. Specifically, the proposed CANet has two novel context adaptive modules, namely, the context adaptive block (CAB) and the cross-scale contextual attention (CSCA), which utilize attention mechanisms to cope with diverse content breakdowns. The proposed CAB, during the forward propagation, uses an adaptive term to determine the importance between adaptive term and convolution kernel, so as to dynamically balance features based on the degree of breakage (confidence level or soft mask), and the overall calculation is formulated as a classic convolution implementation with an additional attention term to describe local structure. Besides, the proposed CSCA, not only takes advantage of the contextual attention module, but also considers cross-scale information transfer to generate reasonable features for damaged areas, thus alleviating the inefficiency of the long-range modeling capability of convolutional neural networks. Qualitative and quantitative experiments show that our method performs better than state-of-the-arts, producing clearer, more coherent and visually plausible inpainting results. The code can be found at github.com/dengyecode/CANet_image_inpainting. Ye Deng 0005, Siqi Hui, Sanping Zhou, Wenli Huang 0004, Jinjun Wang |
IEEE Trans. Image Process. | 5 |
| 2023 | Milestones in Autonomous Driving and Intelligent Vehicles - Part I: Control, Computing System Design, Communication, HD Map, Testing, and Human BehaviorsabstractInterest in autonomous driving (AD) and intelligent vehicles (IVs) is growing at a rapid pace due to the convenience, safety, and economic benefits. Although a number of surveys have reviewed research achievements in this field, they are still limited in specific tasks and lack systematic summaries and research directions in the future. Our work is divided into three independent articles and the first part is a survey of surveys (SoS) for total technologies of AD and IVs that involves the history, summarizes the milestones, and provides the perspectives, ethics, and future research directions. This is the second part (Part I for this technical survey) to review the development of control, computing system design, communication, high-definition map (HD map), testing, and human behaviors in IVs. In addition, the third part (Part II for this technical survey) is to review the perception and planning sections. The objective of this article is to involve all the sections of AD, summarize the latest technical milestones, and guide abecedarians to quickly understand the development of AD and IVs. Combining the SoS and Part II, we anticipate that this work will bring novel and diverse insights to researchers and abecedarians, and serve as a bridge between past and future. Long Chen 0005, Yuchen Li 0004, Chao Huang 0006, Yang Xing 0002, Daxin Tian, Li Li 0013, Zhongxu Hu, Siyu Teng, Chen Lv 0001, Jinjun Wang, Dongpu Cao, Nanning Zheng 0001, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 10 |
| 2023 | Milestones in Autonomous Driving and Intelligent Vehicles - Part II: Perception and PlanningabstractA growing interest in autonomous driving (AD) and intelligent vehicles (IVs) is fueled by their promise for enhanced safety, efficiency, and economic benefits. While previous surveys have captured progress in this field, a comprehensive and forward-looking summary is needed. Our work fills this gap through three distinct articles. The first part, a “survey of surveys” (SoS), outlines the history, surveys, ethics, and future directions of AD and IV technologies. The second part, “Milestones in AD and IVs Part I: Control, Computing System Design, Communication, high-definition map (HD map), Testing, and Human Behaviors” delves into the development of control, computing system, communication, HD map, testing, and human behaviors in IVs. This part, the third part, reviews perception and planning in the context of IVs. Aiming to provide a comprehensive overview of the latest advancements in AD and IVs, this work caters to both newcomers and seasoned researchers. By integrating the SoS and Part I, we offer unique insights and strive to serve as a bridge between past achievements and future possibilities in this dynamic field. Long Chen 0005, Siyu Teng, Bai Li 0002, Xiaoxiang Na, Yuchen Li 0004, Jinjun Wang, Dongpu Cao, Nanning Zheng 0001, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2022 | Hourglass Attention Network for Image Inpainting
Ye Deng 0005, Siqi Hui, Rongye Meng, Sanping Zhou, Jinjun Wang |
ECCV (18) | 5 |
| 2022 | Learning Knowledge Graph Embedding with Batch Circle LossabstractKnowledge Graph Embedding (KGE) is the process to learn low-dimension representations for entities and relations in knowledge graphs. It is a critical component in Knowledge Graph (KG) for link prediction and knowledge discovery. Many works focus on designing proper score function for KGE, while the study of loss function has attracted relatively less attention. In this paper, we focus on improving the loss function when learning KGE. Specifically, we find that the frequently used margin-based loss in KGE models seeks to maximize the gap between the true facts score fpand the false facts score fnand only cares about the relative order of scores. Since its optimization objective is fp- fn= m, increasing fpis equivalent to decreasing fn. Its optimization objective creates an ambiguous convergence status which impairs the separability of positive and negative facts in embedding space. Inspired by the circle loss that offers a more flexible optimization manner with definite convergence targets and is widely used in computer vision tasks, we further extend it into the KGE model with the presented Batch Circle Loss (BCL). BCL allows multiple positives to be considered per anchor (h, r) (or (r, t)) in addition to multiple negatives (as opposed to a single positive sample as used before in KGE models). By comparing with other approaches, the obtained KGE models using our proposed loss function and training method shows superior performance. Yang Wu 0001, Wenli Huang 0004, Siqi Hui, Jinjun Wang |
IJCNN | 4 |
| 2022 | T-former: An Efficient Transformer for Image InpaintingabstractBenefiting from powerful convolutional neural networks (CNNs), learning-based image inpainting methods have made significant breakthroughs over the years. However, some nature of CNNs (e.g. local prior, spatially shared parameters) limit the performance in the face of broken images with diverse and complex forms. Recently, a class of attention-based network architectures, called transformer, has shown significant performance on natural language processing fields and high-level vision tasks. Compared with CNNs, attention operators are better at long-range modeling and have dynamic weights, but their computational complexity is quadratic in spatial resolution, and thus less suitable for applications involving higher resolution images, such as image inpainting. In this paper, we design a novel attention linearly related to the resolution according to Taylor expansion. And based on this attention, a network called T-former is designed for image inpainting. Experiments on several benchmark datasets demonstrate that our proposed method achieves state-of-the-art accuracy while maintaining a relatively low number of parameters and computational complexity. Ye Deng 0005, Siqi Hui, Sanping Zhou, Deyu Meng, Jinjun Wang |
ACM Multimedia | 5 |
| 2022 | A Novel Hybrid Level Set Model for Non-Rigid Object Contour TrackingabstractMost existing trackers use bounding boxes for object tracking. However, the background contained in the bounding box inevitably decreases the accuracy of the target model, which affects the performance of the tracker and is particularly pronounced for non-rigid objects. To address the above issue, this paper proposes a novel hybrid level set model, which can robustly address the issue of topology changing, occlusions and abrupt motion in non-rigid object tracking by accurately tracking the object contour. In particular, an appearance model is first obtained by repeatedly training and relabeling the initial labeled frame using competing one-class SVMs. Then, by integrating the trained appearance model, an edge detector and image spatial information into the level set model, a new hybrid level set model is presented, which accurately locates the object contour and feeds back to the competing one-class SVMs to update the appearance model of the next frame. In addition, a motion model is defined to predict the accurate location of the object when occlusion and abrupt motion occur in the next frame. Finally, the experimental results on state-of-the-art benchmarks demonstrate the feasibility and effectiveness of the proposed model and the superiority of the proposed method over existing trackers in terms of accuracy and robustness. Yiming Qian, Sanping Zhou, Jinjun Wang, Yee-Hong Yang |
IEEE Trans. Image Process. | 5 |
| 2022 | Multinetwork Collaborative Feature Learning for Semisupervised Person ReidentificationabstractPerson reidentification (Re-ID) aims at matching images of the same identity captured from the disjoint camera views, which remains a very challenging problem due to the large cross-view appearance variations. In practice, the mainstream methods usually learn a discriminative feature representation using a deep neural network, which needs a large number of labeled samples in the training process. In this article, we design a simple yet effective multinetwork collaborative feature learning (MCFL) framework to alleviate the data annotation requirement for person Re-ID, which can confidently estimate the pseudolabels of unlabeled sample pairs and consistently learn the discriminative features of input images. To keep the precision of pseudolabels, we further build a novel self-paced collaborative regularizer to extensively exchange the weight information of unlabeled sample pairs between different networks. Once the pseudolabels are correctly estimated, we take the corresponding sample pairs into the training process, which is beneficial to learn more discriminative features for person Re-ID. Extensive experimental results on the Market1501, DukeMTMC, and CUHK03 data sets have shown that our method outperforms most of the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Deyu Meng, Le Wang 0003, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Learning Generic Feature Representations with Adversarial Regularization for Person Re-IdentificationabstractMany existing person re-identification (Re-ID) methods can achieve human-level accuracy on a single dataset, while most of them can be poorly generalized to other datasets. This is mainly caused by different data distributions between different domains. In this paper, we propose a novel adversarial regularization method to address this issue. Specifically, the features extracted from different datasets will be constrained and focused to follow a more similar distribution during the training process. As a result, our method can learn a feature representation with better inter-domain invariance, which will improve the generalization ability of the resulting model. Besides, our method is flexible and can be combined with any feature learning network. Extensive experiments on both Market1501 and DukeMTMC-reID datasets have demonstrated the effectiveness of our method. Qindong Zhang, Sanping Zhou, Jinjun Wang |
ICIP | 3 |
| 2021 | Learning Contextual Transformer Network for Image InpaintingabstractFully Convolutional Networks with attention modules have been proven effective for learning-based image inpainting. While many existing approaches could produce visually reasonable results, the generated images often show blurry textures or distorted structures around corrupted areas. The main reason is due to the fact that convolutional neural networks have limited capacity for modeling contextual information with long range dependencies. Although the attention mechanism can alleviate this problem to some extent, existing attention modules tend to emphasize similarities between the corrupted and the uncorrupted regions while ignoring the dependencies from within each of them. Hence, this paper proposes the Contextual Transformer Network (CTN) which not only learns relationships between the corrupted and the uncorrupted regions but also exploits their respective internal closeness. Besides, instead of a fully convolutional network, in our CTN, we stack several transformer blocks to replace convolution layers to better model the long range dependencies. Finally, by dividing the image into patches of different sizes, we propose a multi-scale multi-head attention module to better model the affinity among various image regions. Experiments on several benchmark datasets demonstrate superior performance by our proposed approach. Ye Deng 0005, Siqi Hui, Sanping Zhou, Deyu Meng, Jinjun Wang |
ACM Multimedia | 5 |
| 2021 | Multiple Object Tracking by Trajectory Map Regression with Temporal Priors EmbeddingabstractPrevailing Multiple Object Tracking (MOT) works following the Tracking-by-Detection (TBD) paradigm pay most attention to either object detection in a first step or data association in a second step. In this paper, we approach the MOT problem from a different perspective by directly obtaining the embedded spatial-temporal information of trajectories from raw video data. For the purpose we propose a joint trajectory locating and attributes encoding framework for real-time, on-line MOT. We firstly introduce a trajectory attribute representation scheme designed for each tracked target (instead of object) where the extracted Trajectory Map (TM) encodes the spatial-temporal attributes of a trajectory across a window of consecutive video frames. Next we present a Temporal Priors Embedding (TPE) methodology to infer these attributes with a logical reasoning strategy based on long-term feature dynamics. The proposed MOT framework projects multiple attributes of tracked targets, e.g., presence, enter/exit, location, scale, motion, etc. into a continuous TM to perform one-shot regression for real-time MOT. Experimental results show that, our proposed video-based method runs at 33 FPS and is more accurate and robust as compared to the detection-based tracking methods and a few other State-of-the- Art (SOTA) approaches on MOT16/17/20 benchmarks. Xingyu Wan, Sanping Zhou, Jinjun Wang, Rongye Meng |
ACM Multimedia | 3 |
| 2021 | Single-Image super-resolution - When model adaptation matters
Yudong Liang, Radu Timofte, Jinjun Wang, Sanping Zhou, Yihong Gong, Nanning Zheng 0001 |
Pattern Recognit. | 3 |
| 2021 | Tracking Beyond Detection: Learning a Global Response Map for End-to-End Multi-Object TrackingabstractMost of the existing Multi-Object Tracking (MOT) approaches follow the Tracking-by-Detection and Data Association paradigm, in which objects are firstly detected and then associated in the tracking process. In recent years, deep neural network has been utilized to obtain more discriminative appearance features for cross-frame association, and noticeable performance improvement has been reported. On the other hand, the Tracking-by-Detection framework is yet not completely end-to-end, which leads to huge computation and limited performance especially in the inference (tracking) process. To address this problem, we present an effective end-to-end deep learning framework which can directly take image-sequence/video as input and output the located and tracked objects of learned types. Specifically, a novel global response network is learned to project multiple objects in the image-sequence/video into a continuous response map, and the trajectory of each tracked object can then be easily picked out. The overall process is similar to how a detector inputs an image and outputs the bounding boxes of each detected object. Experimental results based on the MOT16 and MOT17 benchmarks show that our proposed on-line tracker achieves state-of-the-art performance on several tracking metrics. Xingyu Wan, Jiakai Cao, Sanping Zhou, Jinjun Wang, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Hierarchical and Interactive Refinement Network for Edge-Preserving Salient Object DetectionabstractSalient object detection has undergone a very rapid development with the blooming of Deep Neural Network (DNN), which is usually taken as an important preprocessing procedure in various computer vision tasks. However, the down-sampling operations, such as pooling and striding, always make the final predictions blurred at edges, which has seriously degenerated the performance of salient object detection. In this paper, we propose a simple yet effective approach, i.e., Hierarchical and Interactive Refinement Network (HIRN), to preserve the edge structures in detecting salient objects. In particular, a novel multi-stage and dual-path network structure is designed to estimate the salient edges and regions from the low-level and high-level feature maps, respectively. As a result, the predicted regions will become more accurate by enhancing the weak responses at edges, while the predicted edges will become more semantic by suppressing the false positives in background. Once the salient maps of edges and regions are obtained at the output layers, a novel edge-guided inference algorithm is introduced to further filter the resulting regions along the predicted edges. Extensive experiments on several benchmark datasets have been conducted, in which the results show that our method significantly outperforms a variety of state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Le Wang 0003, Jimuyang Zhang, Fei Wang 0037, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Image Inpainting Using Parallel NetworkabstractDue to the lack of contextual information and the difficulty to directly learn the distribution of a complete image, existing image inpainting methods always use a two-stages approach to make plausible prediction for missing pixels in a coarse-to-fine manner. In this paper, we propose a novel inpainting method with two parallel pipelines. The first pipeline is a standard image completion path that takes the corrupted image as input and outputs the predicted complete image. The second pipeline exists only during the training phase that inputs a complementary image of the corrupted one and still outputs the same complete image. The two pipelines operate simultaneously, and they share identical encoder and most parameters in the decoder. Furthermore, inspired by VAE, random Gaussian noise are added to the features not only to improve the robustness of the model but also to enable generating diverse and plausible results. We evaluated our model on several public datasets and demonstrated that the proposed method outperforms several state-of-the-arts approaches. Ye Deng 0005, Jinjun Wang |
ICIP | 2 |
| 2020 | Meta Corrupted Pixels Mining for Medical Image Segmentation
Sanping Zhou, Chaowei Fang, Le Wang 0003, Jinjun Wang |
MICCAI (1) | 5 |
| 2020 | Temporal Aggregation with Clip-level Attention for Video-based Person Re-identificationabstractVideo-based person re-identification (Re-ID) methods can extract richer features than image-based ones from short video clips. The existing methods usually apply simple strategies, such as average/max pooling, to obtain the tracklet-level features, which has been proved hard to aggregate the information from all video frames. In this paper, we propose a simple yet effective Temporal Aggregation with Clip-level Attention Network (TACAN) to solve the temporal aggregation problem in a hierarchal way. Specifically, a tracklet is firstly broken into different numbers of clips, through a two-stage temporal aggregation network we can get the tracklet-level feature representation. A novel min-max loss is introduced to learn both a clip-level attention extractor and a clip-level feature representer in the training process. Afterwards, the resulting clip-level weights are further taken to average the clip-level features, which can generate a robust tracklet-level feature representation at the testing stage. Experimental results on four benchmark datasets, including the MARS, iLIDS-VID, PRID-2011 and DukeMTMC-VideoReID, show that our TACAN has achieved significant improvements as compared with the state-of-the-art approaches. Mengliu Li, Jinjun Wang, Wenpeng Li, Yongli Sun |
WACV | 3 |
| 2020 | Tracking Persons-of-Interest via Unsupervised Representation Adaptation
Jia-Bin Huang 0001, Jongwoo Lim, Yihong Gong, Jinjun Wang, Narendra Ahuja, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 5 |
| 2020 | Hierarchical U-Shape Attention Network for Salient Object DetectionabstractSalient object detection aims at locating the most conspicuous objects in natural images, which usually acts as a very important pre-processing procedure in many computer vision tasks. In this paper, we propose a simple yet effective Hierarchical U-shape Attention Network (HUAN) to learn a robust mapping function for salient object detection. Firstly, a novel attention mechanism is formulated to improve the well-known U-shape network [1], in which the memory consumption can be extensively reduced and the mask quality can be significantly improved by the resulting U-shape Attention Network (UAN). Secondly, a novel hierarchical structure is constructed to well bridge the low-level and high-level feature representations between different UANs, in which both the intra-network and inter-network connections are considered to explore the salient patterns from a local to global view. Thirdly, a novel Mask Fusion Network (MFN) is designed to fuse the intermediate prediction results, so as to generate a salient mask which is in higher-quality than any of those inputs. Our HUAN can be trained together with any backbone network in an end-to-end manner, and high-quality masks can be finally learned to represent the salient objects. Extensive experimental results on several benchmark datasets show that our method significantly outperforms most of the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Jimuyang Zhang, Le Wang 0003, Shaoyi Du, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Discriminative Feature Learning With Consistent Attention Regularization for Person Re-IdentificationabstractPerson re-identification (Re-ID) has undergone a rapid development with the blooming of deep neural network. Most methods are very easily affected by target misalignment and background clutter in the training process. In this paper, we propose a simple yet effective feedforward attention network to address the two mentioned problems, in which a novel consistent attention regularizer and an improved triplet loss are designed to learn foreground attentive features for person Re-ID. Specifically, the consistent attention regularizer aims to keep the deduced foreground masks similar from the low-level, mid-level and high-level feature maps. As a result, the network will focus on the foreground regions at the lower layers, which is benefit to learn discriminative features from the foreground regions at the higher layers. Last but not least, the improved triplet loss is introduced to enhance the feature learning capability, which can jointly minimize the intra-class distance and maximize the inter-class distance in each triplet unit. Experimental results on the Market1501, DukeMTMC-reID and CUHK03 datasets have shown that our method outperforms most of the state-of-the-art approaches. Sanping Zhou, Fei Wang 0037, Zeyi Huang, Jinjun Wang |
ICCV | 4 |
| 2019 | Deep Self-Paced Learning for Semi-Supervised Person Re-Identification Using Multi-View Self-Paced ClusteringabstractSemi-supervised person re-identification (Re-ID) is an extension of the existing popular Re-ID research, which only uses a small portion of labeled data, while the majority of the training samples are unlabeled. This paper approaches the problem by constructing a set of heterogeneous convolutional neural networks (CNNs) fine-tuned by utilizing the labeled training samples, and then propagating the labels to the unlabeled portion for further fine-tuning the overall system in a self-paced manner. In this work, a novel self-paced multi-view clustering is presented to generate pseudo labels for unlabeled training samples, which combines multiple heterogeneous CNNs features to cluster. In our clustering method, we introduce a self-paced regularizer to select reliable samples for fine-tuning each CNNs by minimizing ranking loss and identification loss. Specifically, we select a small portion of unlabeled training data when multiple CNNs are weak. With CNNs become stronger, more and more unlabeled samples are selected. Pseudo label estimation and CNNs training are improved simultaneously, which optimize alternatively until all the unlabeled training samples are selected. In our framework, both the optimization of multiple CNNs training and multi-view clustering on unlabeled training samples are self-paced optimizing procedure. Extensive experiments have been conducted on two large-scale Re-ID datasets to demonstrate the superiority of the proposed method. Xiaomeng Xin, Xindi Wu, Yuechen Wang, Jinjun Wang |
ICIP | 4 |
| 2019 | Semi-supervised person re-identification using multi-view clustering
Xiaomeng Xin, Jinjun Wang, Ruji Xie, Sanping Zhou, Wenli Huang 0004, Nanning Zheng 0001 |
Pattern Recognit. | 2 |
| 2019 | Normalized Non-Negative Sparse Encoder for Fast Image RepresentationabstractImage representation based on sparse coding generalizes the bag of words model. Although it reduces the reconstruction error for local features to achieve the state-of-the-art image classification performance, the large computational cost hinders the application of sparse coding-based image features. In this paper, we propose approximating a sparse code using the output of a simple neural network. The resulting parameter learning model for the neural network automatically incorporates non-negative and shift-invariant constraints, leading to an efficient normalized non-negative sparse coding (N3SC) sparse encoder. Without the use of the traditional iterative process to solve the sparse coding objective, the sparse encoder directly “converts” each local feature into a sparse code. We also introduce a method for training the encoder based on the auto-encoder method. In addition, we formally propose the corresponding sparse coding scheme called N3SC, which enforces both the non-negative constraint and the shift-invariant constraint in addition to the traditional sparse coding criteria. As demonstrated by several experiments, the obtained N3SC encoder requires only 3%-10% of the processing time for image feature extraction compared with the standard sparse coding scheme. At the same time, the features extracted using the exact solutions of the N3SC coding scheme and the N3SC encoder offer superior image classification accuracy compared to the accuracy of many existing sparse coding-based representations. Shizhou Zhang, Jinjun Wang, Weiwei Shi 0003, Yihong Gong, Yong Xia 0001, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Discriminative Feature Learning With Foreground Attention for Person Re-IdentificationabstractThe performance of person re-identification (Re-ID) has been seriously affected by the large cross-view appearance variations caused by mutual occlusions and background clutter. Hence, learning a feature representation that can adaptively emphasize the foreground persons becomes very critical to solve the person Re-ID problem. In this paper, we propose a simple yet effective foreground attentive neural network (FANN) to learn a discriminative feature representation for person Re-ID, which can adaptively enhance the positive side of foreground and weaken the negative side of background. Specifically, a novel foreground attentive subnetwork is designed to drive the network’s attention, in which a decoder network is used to reconstruct the binary mask by using a novel local regression loss function, and an encoder network is regularized by the decoder network to focus its attention on the foreground persons. The resulting feature maps of encoder network are further fed into the body part subnetwork and feature fusion subnetwork to learn discriminative features. Besides, a novel symmetric triplet loss function is introduced to supervise feature learning, in which the intra-class distance is minimized and the inter-class distance is maximized in each triplet unit, simultaneously. Training our FANN in a multi-task learning framework, a discriminative feature representation can be learned to find out the matched reference to each probe among various candidates in the gallery. Extensive experimental results on several public benchmark datasets are evaluated, which have shown clear improvements of our method over the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Deyu Meng, Yudong Liang, Yihong Gong, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Multi-Object Tracking Using Online Metric Learning with Long Short-Term MemoryabstractThe capacity to model temporal dependency by Recurrent Neural Networks (RNNs) makes it a plausible selection for the multi-object tracking (MOT) problem. Due to the nonlinear transformations and the unique memory mechanism, Long Short-Term Memory (LSTM) can consider a window of history when learning discriminative features, which suggests that the LSTM is suitable for state estimation of target objects as they move around. This paper focuses on association based MOT, and we propose a novel Siamese LSTM Network to interpret both temporal and spatial components nonlinearly by learning the feature of trajectories, and outputs the similarity score of two trajectories for data association. In addition, we also introduce an online metric learning scheme to update the state estimation of each trajectory dynamically. Experimental evaluation on MOT16 benchmark shows that the proposed method achieves competitive performance compared with other state-of-the-art works. Xingyu Wan, Jinjun Wang, Zhifeng Kong, Shunming Deng |
ICIP | 2 |
| 2018 | Continuous Action Recognition and Segmentation in Untrimmed VideosabstractRecognizing continuous human action is a fundamental task in many real-world computer vision applications including video surveillance, video retrieval, and human-computer interaction, etc. It requires to recognize each action performed as well as their segmentation boundaries in a continuous sequence. In previous works, great progress has been reported for single action recognition, by using deep convolutional networks. In order to further improve the performance for continuous action recognition, in this paper, we introduce a discriminative approach consisting of three modules. The first feature extraction module uses a two stream Convolutional Neural Network to capture the appearance and the short-term motion information from the raw video input. Based on the obtained features, the second classification module performs spatial and temporal recognition and then fuses the two scores from respective feature stream. In the final segmentation module, a semi-Markov Conditional Field model, capable of handling long-term action interactions, is built to partition the action sequence. As can be seen in the experimental results, our approach obtains state-of-the-art performance on public datasets including 50Salads, Breakfast, and MERL Shopping. We have also visualized the continuous actions segmentation results for more insightful discussion in the paper. Ruibin Bai, Sanping Zhou, Xueji Zhao, Jinjun Wang |
ICPR | 6 |
| 2018 | Learning solutions to two dimensional electromagnetic equations using LS-SVM
Jinjun Wang, Ziku Wu, Guofeng Li |
Neurocomputing | 2 |
| 2018 | Face alignment recurrent network
Qiqi Hou, Jinjun Wang, Ruibin Bai, Sanping Zhou, Yihong Gong |
Pattern Recognit. | 2 |
| 2018 | Deep ranking model by large adaptive margin learning for person re-identification
Sanping Zhou, Jinjun Wang, Qiqi Hou |
Pattern Recognit. | 3 |
| 2018 | Deep self-paced learning for person re-identification
Sanping Zhou, Jinjun Wang, Deyu Meng, Xiaomeng Xin, Yihong Gong, Nanning Zheng 0001 |
Pattern Recognit. | 2 |
| 2018 | Large Margin Learning in Set-to-Set Similarity Comparison for Person ReidentificationabstractPerson reidentification aims at matching images of the same person across disjoint camera views, which is a challenging problem in multimedia analysis, multimedia editing, and content-based media retrieval communities. The major challenge lies in how to preserve similarity of the same person across video footages with large appearance variations, while discriminating different individuals. To address this problem, conventional methods usually consider the pairwise similarity between persons by only measuring the point-to-point distance. In this paper, we propose using a deep learning technique to model a novel set-to-set (S2S) distance, in which the underline objective focuses on preserving the compactness of intraclass samples for each camera view, while maximizing the margin between the intraclass set and interclass set. The S2S distance metric consists of three terms, namely, the class-identity term, the relative distance term, and the regularization term. The class-identity term keeps the intraclass samples within each camera view gathering together, the relative distance term maximizes the distance between the intraclass class set and interclass set across different camera views, and the regularization term smoothes the parameters of the deep convolutional neural network. As a result, the final learned deep model can effectively find out the matched target to the probe object among various candidates in the video gallery by learning discriminative and stable feature representations. Using the CUHK01, CUHK03, PRID2011, and Market1501 benchmark datasets, we extensively conducted comparative evaluations to demonstrate the advantages of our method over the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Qiqi Hou, Yihong Gong, Nanning Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Improving CNN Performance Accuracies With Min-Max ObjectiveabstractWe propose a novel method for improving performance accuracies of convolutional neural network (CNN) without the need to increase the network complexity. We accomplish the goal by applying the proposed Min-Max objective to a layer below the output layer of a CNN model in the course of training. The Min-Max objective explicitly ensures that the feature maps learned by a CNN model have the minimum within-manifold distance for each object manifold and the maximum between-manifold distances among different object manifolds. The Min-Max objective is general and able to be applied to different CNNs with insignificant increases in computation cost. Moreover, an incremental minibatch training procedure is also proposed in conjunction with the Min-Max objective to enable the handling of large-scale training data. Comprehensive experimental evaluations on several benchmark data sets with both the image classification and face verification tasks reveal that employing the proposed Min-Max objective in the training process can remarkably improve performance accuracies of a CNN model in comparison with the same model trained without using this objective. Weiwei Shi 0003, Yihong Gong, Jinjun Wang, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Point to Set Similarity Based Deep Feature Learning for Person Re-IdentificationabstractPerson re-identification (Re-ID) remains a challenging problem due to significant appearance changes caused by variations in view angle, background clutter, illumination condition and mutual occlusion. To address these issues, conventional methods usually focus on proposing robust feature representation or learning metric transformation based on pairwise similarity, using Fisher-type criterion. The recent development in deep learning based approaches address the two processes in a joint fashion and have achieved promising progress. One of the key issues for deep learning based person Re-ID is the selection of proper similarity comparison criteria, and the performance of learned features using existing criterion based on pairwise similarity is still limited, because only P2P distances are mostly considered. In this paper, we present a novel person Re-ID method based on P2S similarity comparison. The P2S metric can jointly minimize the intra-class distance and maximize the inter-class distance, while back-propagating the gradient to optimize parameters of the deep model. By utilizing our proposed P2S metric, the learned deep model can effectively distinguish different persons by learning discriminative and stable feature representations. Comprehensive experimental evaluations on 3DPeS, CUHK01, PRID2011 and Market1501 datasets demonstrate the advantages of our method over the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Yihong Gong, Nanning Zheng 0001 |
CVPR | 2 |
| 2017 | Single Image Super-Resolution with a Parameter Economic Residual-Like Convolutional Neural Network
Ze Yang 0003, Yudong Liang, Jinjun Wang |
MMM (1) | 4 |
| 2017 | Part-aware trajectories association across non-overlapping uncalibrated cameras
De Cheng, Yihong Gong, Jinjun Wang, Qiqi Hou, Nanning Zheng 0001 |
Neurocomputing | 3 |
| 2017 | Combining local and global hypotheses in deep neural network for multi-label image classification
Qinghua Yu, Jinjun Wang, Shizhou Zhang, Yihong Gong, Jizhong Zhao |
Neurocomputing | 2 |
| 2017 | Correntropy-based level set method for medical image segmentation and bias correction
Sanping Zhou, Jinjun Wang, Yihong Gong |
Neurocomputing | 2 |
| 2017 | Constructing Deep Sparse Coding Network for image classification
Shizhou Zhang, Jinjun Wang, Yihong Gong, Nanning Zheng 0001 |
Pattern Recognit. | 2 |
| 2016 | Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss FunctionabstractPerson re-identification across cameras remains a very challenging problem, especially when there are no overlapping fields of view between cameras. In this paper, we present a novel multi-channel parts-based convolutional neural network (CNN) model under the triplet framework for person re-identification. Specifically, the proposed CNN model consists of multiple channels to jointly learn both the global full-body and local body-parts features of the input persons. The CNN model is trained by an improved triplet loss function that serves to pull the instances of the same person closer, and at the same time push the instances belonging to different persons farther from each other in the learned feature space. Extensive comparative evaluations demonstrate that our proposed method significantly outperforms many state-of-the-art approaches, including both traditional and deep network-based ones, on the challenging i-LIDS, VIPeR, PRID2011 and CUHK01 datasets. De Cheng, Yihong Gong, Sanping Zhou, Jinjun Wang, Nanning Zheng 0001 |
CVPR | 4 |
| 2016 | Image Quality Assessment Using Similar Scene as Reference
Yudong Liang, Jinjun Wang, Xingyu Wan, Yihong Gong, Nanning Zheng 0001 |
ECCV (5) | 2 |
| 2016 | Tracking Persons-of-Interest via Adaptive Discriminative Features
Yihong Gong, Jia-Bin Huang 0001, Jongwoo Lim, Jinjun Wang, Narendra Ahuja, Ming-Hsuan Yang 0001 |
ECCV (5) | 5 |
| 2016 | Application of fuzzy enhancement in moving object detection and trackingabstractIn this paper, we propose an improved fuzzy enhancement algorithm for the moving object detection and tracking. A new membership function is proposed based on the theory of fuzzy sets, which improves the traditional Pal-King fuzzy enhancement algorithm and overcomes the loss of gray information after the processing of enhancement. Since a gray image with multiple targets need more than one crossover point (threshold) for image segmentation, a method for multi-threshold segmentation based on Otsu algorithm is proposed. This method can get multiple thresholds of the image accurately in a short time. After the processing of fuzzy enhancement the moving object is detected and tracked according to the image centroid. Experimental results show that the proposed algorithm can achieve moving object detection and tracking accurately and quickly. Zhengguang Xu, Jinjun Wang |
ICARCV | 3 |
| 2016 | Improving CNN Performance with Min-Max Objective
Weiwei Shi 0003, Yihong Gong, Jinjun Wang |
IJCAI | 3 |
| 2016 | Improving DCNN Performance with Sparse Category-Selective Objective Function
Shizhou Zhang, Yihong Gong, Jinjun Wang |
IJCAI | 3 |
| 2016 | Incorporating image priors with deep convolutional neural networks for image super-resolution
Yudong Liang, Jinjun Wang, Sanping Zhou, Yihong Gong, Nanning Zheng 0001 |
Neurocomputing | 2 |
| 2016 | Active contour model based on local and global intensity information for medical image segmentation
Sanping Zhou, Jinjun Wang, Yudong Liang, Yihong Gong |
Neurocomputing | 2 |
| 2015 | Facial landmark detection via cascade multi-channel convolutional neural networkabstractThis paper presents a novel cascade multi-channel convolutional neural networks(CMC-CNN) approach for face alignment. Several CNN are jointly used for the finally output. In our method, each stage CNN takes the local region around the landmarks as input, and each local patches does convolution separately, which can lead network to learn local high-level features. Then a fully connected layer is put to learn global information from these local features. Our methods has achieves the state-of-the-art results when tested on the 300 Face in-the-Wild(300-W) dataset. Qiqi Hou, Jinjun Wang, Lele Cheng, Yihong Gong |
ICIP | 2 |
| 2015 | Incorporating image degeneration modeling with multitask learning for image super-resolutionabstractLearning the non-linear image upscaling process has previously been considered as a simple regression process, where various models have been utilized to describe the correlations between high-resolution (HR) and low-resolution (LR) images/patches. In this paper, we present a multitask learning framework based on deep neural network for image super-resolution, where we jointly consider the image super-resolution process and the image degeneration process. By sharing parameters between the two highly relevant tasks, the proposed framework could effectively improve the obtained neural network based mapping model between HR and LR image patches. Experimental results have demonstrated clear visual improvement and high computational efficiency, especially with large magnification factors. Yudong Liang, Jinjun Wang, Shizhou Zhang, Yihong Gong |
ICIP | 2 |
| 2015 | Multi-cue Normalized Non-Negative Sparse Encoder for image classificationabstractRecently, the sparse coding based image representation has achieved state-of-the-art recognition results on many benchmarks. In this paper, we propose Multi-cue Normalized Non-Negative Sparse Encoder (MN3SE) which enforces both the non-negative constraint and the shift-invariant constraint on top of the traditional sparse coding criteria, and takes multi-cue to further boost the performance. The former constraint reduces information loose by the negative coefficients and improves the coding stability, and the latter allows the sparseness to be self-adaptive to the local feature. The proposed coding scheme is then approximated by an neural network based encoder for speed-up. More importantly, the multi-layer neural network architecture allows us to apply a multi-task learning strategy to fuse information from multi-cue. Specifically, we take one type of descriptor, such as SIFT as the input, and enforce the learned encoder to produce sparse code that can reconstruct not only SIFT but also other types of descriptors such as color moments. In this way, we could achieve not only 10 to 33 times speed up for sparse-coding, the multi-cue enforced learning strategy gives the image feature extracted by MN3SE superior image classification accuracy. Shizhou Zhang, Jinjun Wang, Yudong Liang, Yihong Gong, Nanning Zheng 0001 |
ICME | 2 |
| 2015 | Deep Self-Organizing Map for visual classificationabstractWe proposed a Deep Self-Organizing Map (DSOM) algorithm which is completely different from the existing multi-layers SOM algorithms, such as SOINN. It consists of layers of alternating self-organizing map and sampling operator. The self-organizing layer is made up of certain numbers of SOMs, with each map only looking at a local region block on its input. The winning neuron's index value from every SOM in self-organizing layer is then organized in the sampling layer to generate another 2D map, which could then be fed to a second self-organizing layer. In this way, local information is gathered together, forming more global information in higher layers. The construction method of the DSOM is unique and will be introduced in this paper. Experiments were carried out to discuss how the DSOM architecture parameters affect the performance. We evaluate our proposed DSOM on MNIST and CASIA-HWDB1.1 dataset. Experimental results show that DSOM outperforms the original supervised SOM by 7:17% on MNIST and 7:25% on CASIA-HWDB1.1. Jinjun Wang, Yihong Gong |
IJCNN | 2 |
| 2015 | Robust Deep Auto-encoder for Occluded Face RecognitionabstractOcclusions by sunglasses, scarf, hats, beard, shadow etc, can significantly reduce the performance of face recognition systems. Although there exists a rich literature of researches focusing on face recognition with illuminations, poses and facial expression variations, there is very limited work reported for occlusion robust face recognition. In this paper, we present a method to restore occluded facial regions using deep learning technique to improve face recognition performance. Inspired by SSDA for facial occlusion removal with known occlusion type and explicit occlusion location detection from a preprocessing step, this paper further introduces Double Channel SSDA (DC-SSDA) which requires no prior knowledge of the types and the locations of occlusions. Experimental results based on CMU-PIE face database have showed that, the proposed method is robust to a variety of occlusion types and locations, and the restored faces could yield significant recognition performance improvements over occluded ones. Lele Cheng, Jinjun Wang, Yihong Gong, Qiqi Hou |
ACM Multimedia | 2 |
| 2015 | Training mixture of weighted SVM for object detection using EM algorithm
De Cheng, Jinjun Wang, Xing Wei 0001, Yihong Gong |
Neurocomputing | 2 |
| 2015 | Summarizing surveillance videos with local-patch-learning-based abnormality detection, blob sequence optimization, and type-based synopsis
Weiyao Lin, Jiwen Lu, Bing Zhou 0003, Jinjun Wang, Yu Zhou 0015 |
Neurocomputing | 5 |
| 2015 | Visual tracking based on online sparse feature learning
Zelun Wang, Jinjun Wang, Yihong Gong |
Image Vis. Comput. | 2 |
| 2015 | Discriminative and generative vocabulary tree: With application to vein image authentication and recognition
Jinjun Wang, Jing Xiao 0006, Weiyao Lin, Chuanfei Luo |
Image Vis. Comput. | 1 |
| 2015 | Multi-target tracking by learning local-to-global trajectory models
Jinjun Wang, Zelun Wang, Yihong Gong, Yuehu Liu |
Pattern Recognit. | 2 |
| 2015 | Online Multi-Target Tracking With Unified Handling of Complex ScenariosabstractComplex scenarios, including miss detections, occlusions, false detections, and trajectory terminations, make the data association challenging. In this paper, we propose an online tracking-by-detection method to track multiple targets with unified handling of aforementioned complex scenarios, where current detection responses are linked to the previous trajectories. We introduce a dummy node to each trajectory to allow it to temporally disappear. If a trajectory fails to find its matching detection, it will be linked to its corresponding dummy node until the emergence of its matching detection. Source nodes are also incorporated to account for the entrance of new targets. The standard Hungarian algorithm, extended by the dummy nodes, can be exploited to solve the online data association implicitly in a global manner, although it is formulated between two consecutive frames. Moreover, as dummy nodes tend to accumulate in a fake or disappeared trajectory while they only occasionally appear in a real trajectory, we can deal with false detections and trajectory terminations by simply checking the number of consecutive dummy nodes. Our approach works on a single, uncalibrated camera, and requires neither scene prior knowledge nor explicit occlusion reasoning, running at 132 frames/s on the PETS09-S2L1 benchmark sequence. The experimental results validate the effectiveness of the dummy nodes in complex scenarios and show that our proposed approach is robust against false detections and miss detections. Quantitative comparisons with other methods on five benchmark sequences demonstrate that we can achieve comparable results with the most existing offline methods and better results than other online algorithms. Huaizu Jiang, Jinjun Wang, Yihong Gong, Na Rong, Zhenhua Chai, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Low Computation Face Verification Using Class Center AnalysisabstractDespite the existence of many state-of-the-art face verification systems, the use of complex features and/or high order recognition models in these systems limits their application in devices with low computation power or low latency requirement. In this paper, we approach the problem by performing verification using simple linear distance model. We introduce a novel probability-based distance metric learning algorithm called Class Center Analysis (CCA) to improve the matching performance in a transformed space. CCA generalizes the classic Neighborhood Components Analysis (NCA) from two aspects. First NCA often leads to distributed clusters, while CCA produces more concentrative clusters, And second, NCA sometimes gives over-fitted distance transformation model, while CCA has better generalization ability. With CCA, our system is able to directly project the difference between face image pair into a real-valued score as their similarity, using only simple matrix-vector operation, and thus consuming very low computation. Our comprehensive experimental evaluation show that, CCA outperforms several other benchmark algorithms in verification accuracy. We have also built the CCA algorithm into a mobile application that uses face image for user authentication. Xinzi Zhang, Jinjun Wang, Yihong Gong, Shizhou Zhang |
ICPR | 2 |
| 2014 | Representing And Recognizing Motion Trajectories: A Tube And Droplet ApproachabstractThis paper addresses the problem of representing and recognizing motion trajectories. We first propose to derive scene-related equipotential lines for points in a motion trajectory and concatenate them to construct a 3D tube for representing the trajectory. Based on this 3D tube, a droplet-based method is further proposed which derives a "water droplet" from the 3D tube and recognizes trajectory activities accordingly. Our proposed 3D tube can effectively embed both motion and scene-related information of a motion trajectory while the proposed droplet- based method can suitably catch the characteristics of the 3D tube for activity recognition. Experimental results demonstrate the effectiveness of our approach. Weiyao Lin, Hang Su 0006, Jianxin Wu 0001, Jinjun Wang, Yu Zhou 0015 |
ACM Multimedia | 5 |
| 2014 | Image parsing by loopy dynamic programming
Shizhou Zhang, Jinjun Wang, Yihong Gong, Xinzi Zhang, Xuguang Lan |
Neurocomputing | 2 |
| 2013 | Human behavior segmentation and recognition using Continuous Linear Dynamic SystemabstractRecognizing continuous action composition in human behavior is an important and yet challenging problem. In this paper we tackle the task by developing both reliable image features and classification algorithms. For image features, we introduce the Embedded Optical Flow (EOF) feature based on embedding optical flow using Locality-constrained Linear Coding with weighted average pooling. The EOF feature is histogram-like but presents excellent linear separability. For classification, we propose the Continuous Linear Dynamic System (CLDS) framework that consists of two sets of Linear Dynamic System (LDS) models, one to model the dynamics of individual actions and the other to model the transition between actions. The inference process estimates the best decomposition of the whole sequence into continuous alternating between human actions and action transitions. In this way, both action type and action boundary can be accurately recognized. Extensive experiments demonstrate the effectiveness and efficiency of the proposed EOF feature and CLDS algorithm. Jinjun Wang, Jing Xiao 0006 |
WACV | 1 |
| 2012 | Substructure and boundary modeling for continuous action recognitionabstractThis paper introduces a probabilistic graphical model for continuous action recognition with two novel components: substructure transition model and discriminative boundary model. The first component encodes the sparse and global temporal transition prior between action primitives in state-space model to handle the large spatial-temporal variations within an action class. The second component enforces the action duration constraint in a discriminative way to locate the transition boundaries between actions more accurately. The two components are integrated into a unified graphical structure to enable effective training and inference. Our comprehensive experimental results on both public and in-house datasets show that, with the capability to incorporate additional information that had not been explicitly or efficiently modeled by previous methods, our proposed algorithm achieved significantly improved performance for continuous action recognition. Jinjun Wang, Jing Xiao 0006, Kai-Hsiang Lin, Thomas S. Huang |
CVPR | 2 |
| 2012 | Discriminative and generative vocabulary tree for vein image recognition
Jinjun Wang, Jing Xiao 0006 |
ICPR | 1 |
| 2012 | Resolution-invariant coding for continuous image super-resolution
Jinjun Wang, Shenghuo Zhu |
Neurocomputing | 1 |
| 2012 | Discovering Image Semantics in Codebook Derivative SpaceabstractThe sparse coding based approaches for image recognition have recently shown improved performance than traditional bag-of-features technique. Due to high dimensionality of the image descriptor space, existing systems usually require very large codebook size to minimize coding error in order to get satisfactory accuracy. While most research efforts try to address the problem by constructing a relatively smaller codebook with stronger discriminative power, in this paper, we introduce an alternative solution by enhancing the quality of coding. Particularly, we apply the idea similar to Fisher kernel to the coding framework, where we use the image-dependent codebook derivative to represent the image. The proposed idea is generic across multiple coding criteria, and in this paper, it is applied to enhance the locality-constraint linear coding (LLC). Experiments show that, the extracted new feature, called “LLC+,” achieved significantly improved accuracy on several challenging datasets even with a small codebook of 1/20 the reported size used by LLC. This obviously adds to LLC+ the modeling accuracy, processing speed and codebook training advantages. Jinjun Wang, Yihong Gong |
IEEE Trans. Multim. | 1 |
| 2011 | Learning semantic embedding at a large scaleabstractA key problem in image annotation is to learn the underlying semantics. However, finding such semantic embeddings is a challenge task and often requires large amount of tagging information. In this paper, we propose to utilize multi-modality cues by incorporating visual and textual information as embedded objects. The paper further presents a multi-task learning framework that simultaneously learns the approximation of two semantic embeddings with efficient multi-stage convex relaxation technique. The experiments show that the proposed method presents very promising performance in both memory usage and training time for large-scale dataset, as well as image classification accuracy. Min-Hsuan Tsai, Jinjun Wang, Tong Zhang 0005, Yihong Gong, Thomas S. Huang |
ICIP | 2 |
| 2011 | Special edition on semi-supervised learning for visual content analysis and understanding
Jian Cheng 0001, Jinjun Wang, Shuqiang Jiang, Zhi-Hua Zhou, Edwin R. Hancock |
Pattern Recognit. | 2 |
| 2010 | Locality-constrained Linear Coding for image classificationabstractThe traditional SPM approach based on bag-of-features (BoF) requires nonlinear classifiers to achieve good image classification performance. This paper presents a simple but effective coding scheme called Locality-constrained Linear Coding (LLC) in place of the VQ coding in traditional SPM. LLC utilizes the locality constraints to project each descriptor into its local-coordinate system, and the projected coordinates are integrated by max pooling to generate the final representation. With linear classifier, the proposed approach performs remarkably better than the traditional nonlinear SPM, achieving state-of-the-art performance on several benchmarks. Compared with the sparse coding strategy [22], the objective function used by LLC has an analytical solution. In addition, the paper proposes a fast approximated LLC method by first performing a K-nearest-neighbor search and then solving a constrained least square fitting problem, bearing computational complexity of O(M + K2). Hence even with very large codebooks, our system can still process multiple frames per second. This efficiency significantly adds to the practical values of LLC for real applications. Jinjun Wang, Jianchao Yang, Kai Yu 0001, Fengjun Lv, Thomas S. Huang, Yihong Gong |
CVPR | 1 |
| 2010 | Real-time driving danger-level prediction
Jinjun Wang, Wei Xu 0007, Yihong Gong |
Eng. Appl. Artif. Intell. | 1 |
| 2010 | Resolution enhancement based on learning the sparse association of image patches
Jinjun Wang, Shenghuo Zhu, Yihong Gong |
Pattern Recognit. Lett. | 1 |
| 2010 | Driving Safety Monitoring Using Semisupervised Learning on Time Series DataabstractThis paper introduces a dangerous-driving warning system that uses statistical modeling to predict driving risks. The major challenge of the research is how to discover the safe/dangerous driving patterns from a sparsely labeled training data set. This paper proposes a semisupervised learning method to utilize both the labeled and the unlabeled data, as well as their interdependence to build a proper danger-level function. In addition, the learned function adopts a continuous parametric form, which is more suitable in modeling the continuous safe/dangerous-driving state transitions in a practical dangerous-driving warning system. Our comprehensive experimental evaluations reveal that, in comparison with driving danger-level estimation using classification-based methods, such as the hidden Markov model (HMM) or the conditional random field algorithm, the proposed method requires less training time and achieved higher prediction accuracy. Jinjun Wang, Shenghuo Zhu, Yihong Gong |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2009 | Resolution-Invariant Image Representation and its applicationsabstractWe present a resolution-invariant image representation (RIIR) framework in this paper. The RIIR framework includes the methods of building a set of multi-resolution bases from training images, estimating the optimal sparse resolution-invariant representation of any image, and reconstructing the missing patches of any resolution level. As the proposed RIIR framework has many potential resolution enhancement applications, we discuss three novel image magnification applications in this paper. In the first application, we apply the RIIR framework to perform Multi-Scale Image Magnification where we also introduced a training strategy to built a compact RIIR set. In the second application, the RIIR framework is extended to conduct Continuous Image Scaling where a new base at any resolution level can be generated using existing RIIR set on the fly. In the third application, we further apply the RIIR framework onto Content-Base Automatic Zooming applications. The experimental results show that in all these applications, our RIIR based method outperforms existing methods in various aspects. Jinjun Wang, Shenghuo Zhu, Yihong Gong |
CVPR | 1 |
| 2009 | Normalizing multi-subject variation for drivers' emotion recognitionabstractThe paper attempts the recognition of multiple drivers' emotional state from physiological signals. The major challenge of the research is the severe inter-subject variation such that it is extreme difficult to build a general model for multiple drivers. In this paper, we focus on discovering an optimal feature mapping by utilizing the additional attribute from the drivers. Two models are reported, specifically an auxiliary dimension model and a factorization model. Experimental results show that the proposed method outperform existing algorithms used for emotional state recognition. Jinjun Wang, Yihong Gong |
ICME | 1 |
| 2009 | Resolution-Invariant Image Representation for Content-Based ZoomingabstractThis paper presents a novel Resolution-Invariant Image Representation (RIIR) framework, and applies it for Content-Based Zooming (CBZ) applications. We explain how to generate a multi-resolution bases set, from which the learned image representation can be resolution-invariant. This provides the key technology to support the continues image up-scaling task for the CBZ applications, which existing example-based resolution enhancement approaches cannot handel, or simply 2-D image interpolation algorithm cannot give satisfactory image quality for. We discuss two clustering based methods to construct the bases set. Experimental results show that, both the two methods give good image quality, and the proposed RIIR framework outperforms existing methods in various aspects. Jinjun Wang, Shenghuo Zhu, Yihong Gong |
ICME | 1 |
| 2009 | Visual-Quality Optimizing Super ResolutionabstractAbstract In this paper, we propose a robust image super‐resolution (SR) algorithm that aims to maximize the overall visual quality of SR results. We consider a good SR algorithm to be fidelity preserving, image detail enhancing and smooth. Accordingly, we define perception‐based measures for these visual qualities. Based on these quality measures, we formulate image SR as an optimization problem aiming to maximize the overall quality. Since the quality measures are quadratic, the optimization can be solved efficiently. Experiments on a large image set and subjective user study demonstrate the effectiveness of the perception‐based quality measures and the robustness and efficiency of the presented method. Feng Liu 0015, Jinjun Wang, Shenghuo Zhu, Michael Gleicher, Yihong Gong |
Comput. Graph. Forum | 2 |
| 2008 | Fast image super-resolution using Connected Component enhancementabstractThe paper focuses on reconstructing the discontinuity between homogenous color regions in an interpolated image to improve its perceptual quality. A low-resolution input image is firstly interpolated and then decomposed into several patches. Each patch is then segmented into multiple homogenous regions using connected component analysis technique. Then a spatial-filter is applied to enhance the color/intensity transition between neighboring components. The designed spatial-filter combines the advantages of both bilateral-filtering and unsharp masking methods, with high computational efficiency. The proposed method can be used for image/video super-resolution applications. Experimental results are promising. Jinjun Wang, Yihong Gong |
ICME | 1 |
| 2008 | Recognition of multiple drivers' emotional stateabstractThe paper attempted the recognition of multiple driverspsila emotional state from physiological signals. The major challenge of the research is due to the severe inter-driver variation such that the features of different emotional state are high correlated, and it is found that simple decorrelation method cannot normalize the features well to achieve acceptable classification accuracy. Hence, in this paper, we propose to apply a latent variable to represent the hidden attribute of individual driver and use statistical training. In addition, we applied temporal constraints for the inference process to improve the recognition accuracy. Experimental results show that the proposed method outperform existing algorithms used for emotional state recognition. Jinjun Wang, Yihong Gong |
ICPR | 1 |
| 2008 | Noisy video super-resolutionabstractLow-quality videos often not only have limited resolution, but also suffer from noise. Directly up-sampling a video without considering noise could deteriorate its visual quality due to magnifying noise. This paper addresses this problem with a unified framework that achieves simultaneous de-noising and super-resolution. This framework formulates noisy video super-resolution as an optimization problem, aiming to maximize the visual quality of the result. We consider a good quality result to be fidelity-preserving, detailpreserving and smooth. Accordingly, we propose measures for these qualities in the scenario of de-noising and superresolution. The experiments on a variety of noisy videos demonstrate the effectiveness of the presented algorithm. Feng Liu 0015, Jinjun Wang, Shenghuo Zhu, Michael Gleicher, Yihong Gong |
ACM Multimedia | 2 |
| 2008 | Automatic composition of broadcast sports video
Jinjun Wang, Changsheng Xu, Chng Eng Siong, Hanqing Lu, Qi Tian 0002 |
Multim. Syst. | 1 |
| 2008 | A Novel Framework for Semantic Annotation and Personalized Retrieval of Sports VideoabstractSports video annotation is important for sports video semantic analysis such as event detection and personalization. In this paper, we propose a novel approach for sports video semantic annotation and personalized retrieval. Different from the state of the art sports video analysis methods which heavily rely on audio/visual features, the proposed approach incorporates web-casting text into sports video analysis. Compared with previous approaches, the contributions of our approach include the following. 1) The event detection accuracy is significantly improved due to the incorporation of web-casting text analysis. 2) The proposed approach is able to detect exact event boundary and extract event semantics that are very difficult or impossible to be handled by previous approaches. 3) The proposed method is able to create personalized summary from both general and specific point of view related to particular game, event, player or team according to user's preference. We present the framework of our approach and details of text analysis, video analysis, text/video alignment, and personalized retrieval. The experimental results on event boundary detection in sports video are encouraging and comparable to the manually selected events. The evaluation on personalized retrieval is effective in helping meet users' expectations. Changsheng Xu, Jinjun Wang, Hanqing Lu, Yifan Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2007 | Efficient Video Object Segmentation by Graph-CutabstractSegmentation of video objects from background is a popular computer vision topic and has many important applications. Most existing methods are either computationally expensive or requiring manual initialization, static cameras, and/or rigid scenes. In a previous work, we proposed a joint spatio-temporal linear regression algorithm to automatically cluster the sparse edge/corner pixels in each video frame and obtain two motion models for the object and background respectively. To label the rest pixels for object segmentation, in this paper, we propose to model the Optical-Flow residual error, color intensity residual error and temporal label consistency features, as well as color/edge orientation consistency constrains, in a graph, and apply the Graph-Cut algorithm to minimize the energy of the graph to obtain an optimal segmentation of the two motion layers boundaries. Finally the object layer is identified from the two using simple heuristics. Experimental segmentation result with videos taken by webcams is promising. Jinjun Wang, Wei Xu 0007, Shenghuo Zhu, Yihong Gong |
ICME | 1 |
| 2007 | Generation of Personalized Music Sports Video Using Multimodal CuesabstractIn this paper, we propose a novel automatic approach for personalized music sports video generation. Two research challenges are addressed, specifically the semantic sports video content extraction and the automatic music video composition. For the first challenge, we propose to use multimodal (audio, video, and text) feature analysis and alignment to detect the semantics of events in broadcast sports video. For the second challenge, we introduce the video-centric and music-centric music video composition schemes and proposed a dynamic-programming based algorithm to perform fully or semi-automatic generation of personalized music sports video. The experimental results and user evaluations are promising and show that our systems generated music sports video is comparable to professionally generated ones. Our proposed system greatly facilitates the music sports video editing task for both professionals and amateurs Jinjun Wang, Chng Eng Siong, Changsheng Xu, Hanqing Lu, Qi Tian 0002 |
IEEE Trans. Multim. | 1 |
| 2006 | Fully and Semi-Automatic Music Sports Video CompositionabstractVideo composition is important for music video production. In this paper we propose an automatic method to assist the music sports video composition operation. Our approach is based on dynamic programming algorithm which finds a set of video shots that best matches the music. The method by default is fully-automatic, and users specification could be inserted to control the composition process, making it a semiautomatic system. This research has obvious importance to reduce manual processing, and enables the generation of high quality personalized music sports video. The proposed method is generic and fast. The experimental results are satisfactory Jinjun Wang, Chng Eng Siong, Changsheng Xu |
ICME | 1 |
| 2006 | Identify Sports Video Shots with "Happy" or "Sad" EmotionsabstractSemantic video content extraction and selection are critical steps in sports video analysis and editing. The identification of video segments can be from various semantic perspectives, e.g. certain event, player or emotional state. In this paper, we examined the possibility of automatically identifying shots with "happy" or "sad" emotion from broadcast sports video. Our proposed model first performs the sports highlight extraction to obtain candidate shots that possibly contain emotion information and then classifies these shots into either "happy" or "sad" emotion groups using hidden Markov model based method. The final experimental results are satisfactory Jinjun Wang, Chng Eng Siong, Changsheng Xu, Hanqing Lu, Xiaofeng Tong |
ICME | 1 |
| 2006 | Live sports event detection based on broadcast video and web-casting textabstractEvent detection is essential for sports video summarization, indexing and retrieval and extensive research efforts have been devoted to this area. However, the previous approaches are heavily relying on video content itself and require the whole video content for event detection. Due to the semantic gap between low-level features and high-level events, it is difficult to come up with a generic framework to achieve a high accuracy of event detection. In addition, the dynamic structures from different sports domains further complicate the analysis and impede the implementation of live event detection systems. In this paper, we present a novel approach for event detection from the live sports game using web-casting text and broadcast video. Web-casting text is a text broadcast source for sports game and can be live captured from the web. Incorporating web-casting text into sports video analysis significantly improves the event detection accuracy. Compared with previous approaches, the proposed approach is able to: (1) detect live event only based on the partial content captured from the web and TV; (2) extract detailed event semantics and detect exact event boundary, which are very difficult or impossible to be handled by previous approaches; and (3) create personalized summary related to certain event, player or team according to user's preference. We present the framework of our approach and details of text analysis, video analysis and text/video alignment. We conducted experiments on both live games and recorded games. The results are encouraging and comparable to the manually detected events. We also give scenarios to illustrate how to apply the proposed solution to professional and consumer services. Changsheng Xu, Jinjun Wang, Kong-Wah Wan, Ling-Yu Duan |
ACM Multimedia | 2 |
| 2005 | Soccer replay detection using scene transition structure analysisabstractReplay scene detection is a useful technique for content based sports video analysis. Most current researchers try to find suitable visual and/or compressed domain features to detect the replay scene from a broadcast video. We present a novel approach using context information from the concurrence of replay and other types of shots to detect the replay scenes. We first perform a shot classification and then a scene transition structure analysis on the generated shot label sequence to extract the replay scene. The proposed model is computationally fast and some promising results were obtained. Jinjun Wang, Chng Eng Siong, Changsheng Xu |
ICASSP (2) | 1 |
| 2005 | Periodicity Detection of Local MotionabstractPeriodicity is useful for compact representation of periodic motion and a reasonable selection of a proper temporal scale for periodic motion analysis. In this paper, we concern the periodicity detection of local motion within an interesting region and present an approach to automatically detect the motion periodicity inherent to local motion under complex condition. The task is challenging as local motion is usually buried in clutters with global motion and noises. Most exist ing methods have assumed a static camera and a labeled moving object region. We instead apply robust local motion estimation and an object localization method to extract the object motion. The object motion is characterized by the confidence based motion probability map and the motion vectors obtained by global motion compensation. The autocorrelation series of motion energy is then carried out to locate local maximum points. With the set of indices of local maximum points, we can estimate the basic periodicity through a least-square fitting. This method has been applied to swimming videos and got encouraging results. Xiaofeng Tong, Ling-Yu Duan, Changsheng Xu, Qi Tian 0002, Hanqing Lu, Jinjun Wang, Jesse S. Jin |
ICME | 6 |
| 2005 | Automatic generation of personalized music sports videoabstractIn this paper, we propose a novel automatic approach for personalized music sports video generation. Two research challenges, semantic sports video content selection and automatic video composition, are addressed. For the first challenge, we propose to use multi-modal (audio, video and text) feature analysis and alignment to detect the semantic of events in sports video. For the second challenge, we propose video-centric and music-centric music video composition schemes to automatically generate personalized music sports video based on user's preference. The experimental results and user evaluations are promising and show that our system's generated music sports video is comparable to manually generated ones. The proposed approach greatly facilitates the automatic music sports video generation for both professionals and amateurs. Jinjun Wang, Changsheng Xu, Chng Eng Siong, Ling-Yu Duan, Kong-Wah Wan, Qi Tian 0002 |
ACM Multimedia | 1 |
| 2004 | Event detection based on non-broadcast sports video
Jinjun Wang, Changsheng Xu, Chng Eng Siong, Xinguo Yu, Qi Tian 0002 |
ICIP | 1 |
| 2004 | Sports highlight detection from keyword sequences using HMMabstractSports video highlight detection is a popular topic. A multi-layer sport event detection framework is described. In the mid-level of this framework, visual and audio keywords are created from low-level features and the original video is converted into a keyword sequence. In the high-level, the temporal pattern of keyword sequences is analyzed by an HMM classifier. The creation of visual and audio keywords can help to bridge the gap between low-level features and high-level semantics. The use of the HMM classifier can automatically find the temporal change character of the event instead of rule based heuristic modeling to map certain keyword sequences into events. Experiments using our model on soccer games produced some promising results. Jinjun Wang, Changsheng Xu, Chng Eng Siong, Qi Tian 0002 |
ICME | 1 |
| 2004 | Automatic replay generation for soccer video broadcastingabstractWhile most current approaches for sports video analysis are based on broadcast video, in this paper, we present a novel approach for highlight detection and automatic replay generation for soccer videos taken by the main camera. This research is important as current soccer highlight detection and replay generation from a live game is a labor-intensive process. A robust multi-level, multi-model event detection framework is proposed to detect the event and event boundaries from the video taken by the main camera. This framework explores the possible analysis cues, using a mid-level representation to bridge the gap between low-level features and high-level events. The event detection results and mid-level representation are used to generate replays which are automatically inserted into the video. Experimental results are promising and found to be comparable with those generated by broadcast professionals. Jinjun Wang, Changsheng Xu, Chng Eng Siong, Kong-Wah Wan, Qi Tian 0002 |
ACM Multimedia | 1 |