EDBT 2026 Demo / reviewers in the wild / expert
Jun Chen 0001
dblp:85/5901-1
· DBLP profile ↗
87ranked-venue papers
0as first author
34since 2021 · last 2026
0000-0003-1376-0167ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 19 since 2021Artificial intelligence and machine learning · 23 · 14 since 2021Computer networks · 4 · 3 since 2021Databases, data management, data science and information retrieval · 4Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Pedestrian Detection with Uncertain ModalityabstractExisting cross-modal pedestrian detection (CMPD) employs complementary information from RGB and thermal-infrared (TIR) modalities to detect pedestrians in 24h-surveillance systems. RGB captures rich pedestrian details under daylight, while TIR excels at night. However, TIR focuses primarily on the person's silhouette, neglecting critical texture details essential for detection. While the near-infrared (NIR) captures texture under low-light conditions, which effectively alleviates performance issues of RGB and detail loss in TIR, thereby reducing missed detections. To this end, we construct a new Triplet RGB–NIR–TIR (TRNT) dataset, comprising 8,281 pixel-aligned image triplets, establishing a comprehensive foundation for algorithmic research. However, due to the variable nature of real-world scenarios, imaging devices may not always capture all three modalities simultaneously. This results in input data with unpredictable combinations of modal types, which challenge existing CMPD methods that fail to extract robust pedestrian information under arbitrary input combinations, leading to significant performance degradation. To address these challenges, we propose the Adaptive Uncertainty-aware Network (AUNet) for accurately discriminating modal availability and fully utilizing the available information under uncertain inputs. Specifically, we introduce Unified Modality Validation Refinement (UMVR), which includes an uncertainty-aware router to validate modal availability and a semantic refinement to ensure the reliability of information within the modality. Furthermore, we design a Modality-Aware Interaction (MAI) module to adaptively activate or deactivate its internal interaction mechanisms per UMVR output, enabling effective complementary information fusion from available modalities. AUNet enables accurate modality validation and robust inference without fixed modality pairings, facilitating the effective fusion of RGB, NIR, and TIR information across diverse inputs. Qian Bie, Xiao Wang 0029, Bin Yang 0026, Zhixi Yu, Jun Chen 0001, Xin Xu 0007 |
AAAI | 5 |
| 2025 | Balancing Privacy and Performance: A Many-in-One Approach for Image AnonymizationabstractThe effective utilization of data through Deep Neural Networks (DNNs) has profoundly influenced various aspects of society. The growing demand for high-quality, particularly personalized, data has spurred research efforts to prevent data leakage and protect privacy in recent years. Early privacy-preserving methods primarily relied on instance-wise modifications, such as erasing or obfuscating essential features for de-identification. However, this approach highlights an inherent trade-off: minimal modification offers insufficient privacy protection, while excessive modification significantly degrades task performance. In this paper, we propose a novel Recombining for Obfuscation (FRO) approach to address this trade-off. Unlike existing methods that generate one anonymized instance by perturbing the original data on a one-to-one basis, our FRO approach generates an anonymized instance by reassembling mixed ID-related features from multiple original data sources on a many-in-one basis. Instead of introducing additional noise for de-identification, our approach leverages the existing non-polluted features from other instances to anonymize data. Extensive experiments on identity identification tasks demonstrate that FRO outperforms previous state-of-the-art methods, not only in utility performance but also in visual anonymization. Xuemei Jia, Jiawei Du 0002, Hui Wei 0004, Ruinian Xue, Zheng Wang 0007, Hongyuan Zhu 0002, Jun Chen 0001 |
AAAI | 7 |
| 2025 | Unbiased Prototype Consistency Learning for Multi-Modal and Multi-Task Object Re-IdentificationabstractIn object re-identification (ReID) task, both cross-modal and multi-modal retrieval methods have achieved notable progress. However, existing approaches are designed for specific modality and category (person or vehicle) retrieval task, lacking generalizability to others. Acquiring multiple task-specific models would result in wasteful allocation of both training and deployment resources. To address the practical requirements for unified retrieval, we introduce Multi-Modal and Multi-Task object ReID ($\rm {M^3T}$-ReID). The $\rm {M^3T}$-ReID task aims to utilize a unified model to simultaneously achieve retrieval tasks across different modalities and different categories. Specifically,
to tackle the challenges of modality distibution divergence and category semantics discrepancy posed in $\rm {M^3T}$-ReID, we design a novel Unbiased Prototype Consistency Learning (UPCL) framework, which consists of two main modules: Unbiased Prototypes-guided Modality Enhancement (UPME) and Cluster Prototype Consistency Regularization (CPCR).
UPME leverages modality-unbiased prototypes to simultaneously enhance cross-modal shared features and multi-modal fused features. Additionally, CPCR regulates discriminative semantics learning with category-consistent information through prototypes clustering.
Under the collaborative operation of these two modules, our model can simultaneously learn robust cross-modal shared feature and multi-modal fused feature spaces, while also exhibiting strong category-discriminative capabilities. Extensive experiments on multi-modal datasets RGBNT201 and RGBNT100 demonstrates our UPCL framework showcasing exceptional performance for $\rm {M^3T}$-ReID. The code is available at https://github.com/ZhouZhongao/UPCL. Zhongao Zhou, Bin Yang 0026, Wenke Huang 0003, Jun Chen 0001, Mang Ye |
NeurIPS | 4 |
| 2025 | MonOri: Orientation-Guided PnP for Monocular 3-D Object DetectionabstractMonocular 3-D object detection is a challenging task in the field of autonomous driving and has made great progress. However, current monocular image methods tend to incorporate additional information such as pseudolabels to improve algorithm performance while overlooking the geometric relationship between the object's keypoints, resulting in low performance for occluded object detection. To address this issue, we find that introducing the orientation information of objects in the 3-D detection pipeline can help improve the detection performance of occluded objects. An orientation-guided perspective-n-point (PnP) for monocular 3-D object detection method named MonOri is presented in this article, which uses object's orientation to guide keypoints' optimization. Considering the existence of different deformation objects in the scene, we design the feature aggregation detection module (FADM), which consists of the feature focus fusion module (FFFM) and CondConv detection module (CCDM). First, FFFM can highlight signals from irregularly occluded objects, effectively modeling features of elongated and small-sized objects. This module enhances the model's ability to recognize elongated and small-sized objects in complex scenes. Then, the CCDM is designed to improve the network's ability to estimate object keypoints' location regression under occlusion conditions and minimize the network computational overhead. Finally, considering that the unoccluded portions of occluded objects are closely related to the orientation of the objects, an orientation-guided keypoints' selection module (OGKSM) is proposed to enhance the accuracy of objected optimization for keypoint positions and spatial location inference of the object. Experimental results indicate that the MonOri method achieves competitive results; it is also demonstrated that the orientation information is introduced in the PnP algorithm to estimate the object's spatial position that can mitigate the impact of occlusion on object detection, thus improving the recognition rate of occluded objects. Our code is available at https://github.com/DL-YHD/MonOri. Hongdou Yao, Jun Chen 0001, Zheng Wang 0007, Yansheng Qiu, Xiao Wang 0029, Yimin wang, Xiaoyu Chai, Chenglong Cao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Shallow-Deep Collaborative Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identification (US-VI-ReID) centers on learning a cross-modality retrieval model without labels, reducing the reliance on expensive cross-modality manual annotation. Previous US-VI-ReID works gravitate toward learning cross-modality information with the deep features extracted from the ultimate layer. Nevertheless, interfered by the multiple discrepancies, solely relying on deep features is insufficient for accurately learning modality-invariant features, resulting in negative optimization. The shallow feature from the shallow layers contains nuanced detail information, which is critical for effective cross-modality learning but is dis- regarded regrettably by the existing methods. To address the above issues, we design a Shallow-Deep Collaborative Learning (SDCL) framework based on the transformer with shallow-deep contrastive learning, incorporating Collaborative Neighbor Learning (CNL) and Collaborative Ranking Association (CRA) module. Specifically, CNL unveils the intrinsic homogeneous and heterogeneous collaboration which are harnessed for neighbor alignment, enhancing the robustness in a dynamic manner. Furthermore, CRA associates the cross-modality labels with the ranking association between shallow and deep features, furnishing valuable supervision for cross-modality learning. Extensive experiments validate the superiority of our method, even outperforming certain supervised counterparts. Bin Yang 0026, Jun Chen 0001, Mang Ye |
CVPR | 2 |
| 2024 | GRAformer: A gated residual attention transformer for multivariate time series forecasting
Chengcao Yang, Bin Yang 0026, Jun Chen 0001 |
Neurocomputing | 4 |
| 2024 | Data compensation and feature fusion for sketch based person retrieval
Jun Chen 0001, Mithun Mukherjee 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | SATCount: A scale-aware transformer-based class-agnostic counting framework
Bin Yang 0026, Chao Liang 0001, Jun Chen 0001 |
Neural Networks | 5 |
| 2024 | Federated Learning With Long-Tailed Data via Representation Unification and Classifier RectificationabstractPrevalent federated learning commonly develops under the assumption that the ideal global class distributions are balanced. In contrast, real-world data typically follows the long-tailed class distribution, where models struggle to classify samples from tail classes. In this paper, we alleviate the issue under the long-tailed data, dissecting the into two aspects: the distorted feature space and the biased classifier. Specifically, we propose the Representation Unification and Classifier Rectification (RUCR), which leverages global unified prototypes to shape the feature space and calibrate the classifier. RUCR aggregates local prototypes (class-wise mean features) extracted by the global model to obtain global unified prototypes. It calibrates the feature space by pulling features within the same class towards corresponding global unified prototypes and pushing the other classes away. Moreover, RUCR utilizes global prototypes to reduce the classifier bias via prototypical mix-up. It generates a balanced virtual feature set by arbitrarily fusing global unified prototypes and local features. The classifier re-training is then conducted on the balanced virtual feature set to rectify the decision boundary and thus alleviate the shifts. Empirical results on CIFAR-10-LT, CIFAR-100-LT, and Tiny-Imagenet-LT datasets validate the superior performance of our proposed method. Wenke Huang 0003, Yuxia Liu, Mang Ye, Jun Chen 0001, Bo Du 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Dual Consistency-Constrained Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identification (US-VI-ReID) aims at learning a cross-modality matching model under unsupervised conditions, which is an extremely important task for practical nighttime surveillance to retrieve a specific identity. Previous advanced US-VI-ReID works mainly focus on associating the positive cross-modality identities to optimize the feature extractor by off-line manners, inevitably resulting in error accumulation of incorrect off-line cross-modality associations in each training epoch due to the intra-modality and inter-modality discrepancies. They ignore the direct cross-modality feature interaction in the training process, i.e., the on-line representation learning and updating. Worse still, existing interaction methods are also susceptible to inter-modality differences, leading to unreliable heterogeneous neighborhood learning. To address the above issues, we propose a dual consistency-constrained learning framework (DCCL) simultaneously incorporating off-line cross-modality label refinement and on-line feature interaction learning. The basic idea is that the relations between cross-modality instance-instance and instance-identity should be consistent. More specifically, DCCL constructs an instance memory, an identity memory, and a domain memory for each modality. At the beginning of each training epoch, DCCL explores the off-line consistency of cross-modality instance-instance and instance-identity similarities to refine the reliable cross-modality identities. During the training, DCCL finds credible homogeneous and heterogeneous neighborhoods with on-line consistency between query-instance similarity and query-instance domain probability similarities for feature interaction in one batch, enhancing the robustness against intra-modality and inter-modality variations. Extensive experiments validate that our method significantly outperforms existing works, and even surpasses some supervised counterparts. The source code is available athttps://github.com/yangbincv/DCCL. Bin Yang 0026, Jun Chen 0001, Cuiqun Chen, Mang Ye |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Occlusion-Aware Plane-Constraints for Monocular 3D Object DetectionabstractThe task of 3D object detection poses a significant challenge for 3D scene understanding and is primarily employed in the fields of robot control and autonomous driving. Monocular-based 3D detection methods are more cost-effective and practical than stereo-based or LiDAR-based methods. Monocular image 3D detection methods have garnered considerable attention from researchers. However, the impact of occlusion scenarios of the objects on the keypoints prediction is often overlooked. To address this issue, the present paper proposes a novel 3D monocular object detection method named MonOAPC, which is equipped with occlusion-aware plane-constraints. This method can adaptively utilize partial keypoints to infer the plane location of the object based on the level of occlusion and highlights that the introduction of plane constraints is advantageous for the 3D detection task. First, the plane information of the object in 3D space is beneficial to optimize the keypoints regression, and considering that each plane holds different significance, an adaptive plane location inference module is proposed to enhance the keypoints location regression. Second, a novel co-depth estimation module is proposed to jointly estimate the object’s spatial location through various depth inference methods, thereby improving the generalization of the object depth estimation. Furthermore, the paper demonstrates that the accuracy of 3D object detection can be indirectly improved by introducing plane information to promote keypoints regression, and that plane information is effective for monocular 3D detection. The experimental outcomes show that the MonOAPC method can attain competitive results. Hongdou Yao, Jun Chen 0001, Zheng Wang 0007, Xiao Wang 0029, Xiaoyu Chai, Yansheng Qiu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | An Adaptive-Guidance GAN for Accurate Face Reenactment
Xiaoyu Chai, Jun Chen 0001, Dongshu Xu, Hongdou Yao |
CGI | 2 |
| 2023 | Multi-image 3D Face Reconstruction via an Adaptive Aggregation Network
Xiaoyu Chai, Jun Chen 0001, Dongshu Xu, Hongdou Yao, Zheng Wang 0007, Chia-Wen Lin |
CGI | 2 |
| 2023 | Learning Degradation for Real-World Face Super-Resolution
Jun Chen 0001, Dongshu Xu, Chao Liang 0001, Zhen Han 0002 |
CGI | 2 |
| 2023 | Top-K Visual Tokens Transformer: Selecting Tokens for Visible-Infrared Person Re-IdentificationabstractVisible modality and infrared modality person re-identification (VI-ReID) is an extremely important and challenging task. Existing works mainly focus on reducing the modality gap with Convolutional Neural Networks (CNN). However, the features extracted by CNN may contain useless identity-irrelevant information, which inevitably reduces the discrimination of features. To address this issue, this paper introduces a Top-K Visual Tokens Transformer (TVTR) framework which utilizes a top-k visual tokens selection module to accurately select top-k discriminative visual patches for reducing the distraction of identity-irrelevant information and learning discriminative features. Furthermore, a global-local circle loss is developed to optimize the TVTR for achieving cross-modality positive concentration and negative separation properties. The experimental results on SYSU-MM01 and RegDB datasets demonstrate the superiority of our method. The source code will be released. Bin Yang 0026, Jun Chen 0001, Mang Ye |
ICASSP | 2 |
| 2023 | Towards Grand Unified Representation Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised learning visible-infrared person re-identification (USL-VI-ReID) is an extremely important and challenging task, which can alleviate the issue of expensive cross-modality annotations. Existing works focus on handling the cross-modality discrepancy under unsupervised conditions. However, they ignore the fact that USL-VI-ReID is a cross-modality retrieval task with the hierarchical discrepancy, i.e., camera variation and modality discrepancy, resulting in clustering inconsistencies and ambiguous cross-modality label association. To address these issues, we propose a hierarchical framework to learn grand unified representation (GUR) for USL-VI-ReID. The grand unified representation lies in two aspects: 1) GUR adopts a bottom-up domain learning strategy with a cross-memory association embedding module to explore the information of hierarchical domains, i.e., intra-camera, inter-camera, and inter-modality domains, learning a unified and robust representation against hierarchical discrepancy. 2) To unify the identities of the two modalities, we develop a cross-modality label unification module that constructs a cross-modality affinity matrix as a bridge for propagating labels between two modalities. Then, we utilize the homogeneous structure matrix to smooth the propagated labels, ensuring that the label structure within one modality remains unchanged. Extensive experiments demonstrate that our GUR framework significantly outperforms existing USL-VI-ReID methods, and even surpasses some supervised counterparts. Bin Yang 0026, Jun Chen 0001, Mang Ye |
ICCV | 2 |
| 2023 | Transferring fashion to surveillance with weak labels
Zheng He 0001, Chao Liang 0001, Jun Chen 0001, Chia-Wen Lin, Dapeng Tao |
Neural Comput. Appl. | 4 |
| 2023 | Vertex points are not enough: Monocular 3D object detection via intra- and inter-plane constraints
Hongdou Yao, Jun Chen 0001, Zheng Wang 0007, Xiao Wang 0029, Xiaoyu Chai, Yansheng Qiu |
Neural Networks | 2 |
| 2023 | Translation, Association and Augmentation: Learning Cross-Modality Re-Identification From Single-Modality AnnotationabstractDaytime visible modality (RGB) and night-time infrared (IR) modality person re-identification (VI-ReID) is a challenging cross-modality pedestrian retrieval problem. However, training a cross-modality ReID model requires plenty of cross-modality (visible-infrared) identity labels that are more expensive than single-modality person ReID. To alleviate this issue, this paper studies unsupervised domain adaptive visible infrared person re-identification (UDA-VI-ReID) task without the reliance on any cross-modality annotation. To transfer learned knowledge from the labelled visible source domain to the unlabelled visible-infrared target domain, we propose a Translation, Association and Augmentation (TAA) framework. Specifically, the modality translator is firstly utilized to transfer visible image to infrared image, formulating generated visible-infrared image pairs for cross-modality supervised training. A Robust Association and Mutual Learning (RAML) module is then designed to exploit the underlying relations between visible and infrared modalities for label noise modeling. Moreover, a Translation Supervision and Feature Augmentation (TSFA) module is designed to enhance the discriminability by enriching the supervision with feature augmentation and modality translation. The extensive experimental results demonstrate that our method significantly outperforms current state-of-the-art unsupervised methods under various settings, and even surpasses some supervised counterparts, providing a powerful baseline for UDA-VI-ReID. Bin Yang 0026, Jun Chen 0001, Xianzheng Ma, Mang Ye |
IEEE Trans. Image Process. | 2 |
| 2022 | REMOTE: Reinforced Motion Transformation Network for Semi-supervised 2D Pose Estimation in VideosabstractExisting approaches for 2D pose estimation in videos often require a large number of dense annotations, which are costly and labor intensive to acquire. In this paper, we propose a semi-supervised REinforced MOtion Transformation nEtwork (REMOTE) to leverage a few labeled frames and temporal pose variations in videos, which enables effective learning of 2D pose estimation in sparsely annotated videos. Specifically, we introduce a Motion Transformer (MT) module to perform cross frame reconstruction, aiming to learn motion dynamic knowledge in videos. Besides, a novel reinforcement learning-based Frame Selection Agent (FSA) is designed within our framework, which is able to harness informative frame pairs on the fly to enhance the pose estimator under our cross reconstruction mechanism. We conduct extensive experiments that show the efficacy of our proposed REMOTE framework. Xianzheng Ma, Hossein Rahmani 0001, Zhipeng Fan 0001, Bin Yang 0026, Jun Chen 0001, Jun Liu 0036 |
AAAI | 5 |
| 2022 | Face Super-Resolution with Better Semantics and More Efficient Guidance
Jun Chen 0001, Zheng Wang 0007, Chao Liang 0001, Zhen Han 0002, Chia-Wen Lin |
CGI | 2 |
| 2022 | Augmented Dual-Contrastive Aggregation Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractVisible infrared person re-identification (VI-ReID) aims at searching out the corresponding infrared (visible) images from a gallery set captured by other spectrum cameras. Recent works mainly focus on supervised VI-ReID methods that require plenty of cross-modality (visible-infrared) identity labels which are more expensive than the annotations in single-modality person ReID. For the unsupervised learning visible infrared re-identification (USL-VI-ReID), the large cross-modality discrepancies lead to difficulties in generating reliable cross-modality labels and learning modality-invariant features without any annotations. To address this problem, we propose a novel Augmented Dual-Contrastive Aggregation (ADCA) learning framework. Specifically, a dual-path contrastive learning framework with two modality-specific memories is proposed to learn the intra-modality person representation. To associate positive cross-modality identities, we design a cross-modality memory aggregation module with count priority to select highly associated positive samples, and aggregate their corresponding memory features at the cluster level, ensuring that the optimization is explicitly concentrated on the modality-irrelevant perspective. Extensive experiments demonstrate that our proposed ADCA significantly outperforms existing unsupervised methods under various settings, and even surpasses some supervised counterparts, facilitating VI-ReID to real-world deployment. Code is available at https://github.com/yangbincv/ADCA. Bin Yang 0026, Mang Ye, Jun Chen 0001, Zesen Wu |
ACM Multimedia | 3 |
| 2022 | Online multiple object tracking based on fusing global and partial features
Jun Chen 0001, Mithun Mukherjee 0001, Chao Liang 0001, Weijian Ruan |
Neurocomputing | 2 |
| 2022 | Real-Time Video Deraining via Global Motion Compensation and Hybrid Multi-Scale Temporal CorrelationsabstractThe current video deraining algorithms mainly use adjacent frames to optimize the target frame information. However, they only consider the inter-frame temporal correlations of a uniform scale between frames, ignoring the inter-frame temporal correlations of different scales. In addition, the high computational cost is another drawback of the current video deraining algorithms. To this end, we propose a novel aggregation network that explores the inter-frame multi-scale temporal correlations for video deraining with the small computational cost. First, we construct a hybrid multi-scale feature extraction structure in the network to increase the receptive field of multi-scale features. For similar rain streaks at adjacent frames with different scales, a hybrid multi-scale residual block (HMSRB) is proposed to explore the complementary and redundant information at the temporal dimension to characterize the target frame. At the same time, we introduce an improved global context module (GCM) to avoid the complex motion estimation and motion compensation (ME&MC) operation as in previous video deraining approaches, while reducing the calculation complexity. Finally, a fusion block is utilized to adaptively merge the extracted features. Experiments demonstrate that our proposed network is proved to be more efficient and effective than the existing algorithms. Jun Chen 0001, Zhen Han 0002, Qikui Zhu, Weijian Ruan |
IEEE Signal Process. Lett. | 2 |
| 2022 | TICNet: A Target-Insight Correlation Network for Object TrackingabstractRecently, the correlation filter (CF) and Siamese network have become the two most popular frameworks in object tracking. Existing CF trackers, however, are limited by feature learning and context usage, making them sensitive to boundary effects. In contrast, Siamese trackers can easily suffer from the interference of semantic distractors. To address the above problems, we propose an end-to-end target-insight correlation network (TICNet) for object tracking, which aims at breaking the above limitations on top of a unified network. TICNet is an asymmetric dual-branch network involving a target-background awareness model (TBAM), a spatial-channel attention network (SCAN), and a distractor-aware filter (DAF) for end-to-end learning. Specifically, TBAM aims to distinguish a target from the background in the pixel level, yielding a target likelihood map based on color statistics to mine distractors for DAF learning. SCAN consists of a basic convolutional network, a channel-attention network, and a spatial-attention network, aiming to generate attentive weights to enhance the representation learning of the tracker. Especially, we formulate a differentiable DAF and employ it as a learnable layer in the network, thus helping suppress distracting regions in the background. During testing, DAF, together with TBAM, yields a response map for the final target estimation. Extensive experiments on seven benchmarks demonstrate that TICNet outperforms the state-of-the-art methods while running at real-time speed. Weijian Ruan, Mang Ye, Yi Wu 0001, Wu Liu 0005, Jun Chen 0001, Chao Liang 0001, Ge Li 0002, Chia-Wen Lin |
IEEE Trans. Cybern. | 5 |
| 2021 | Long-short Term Prediction for Occluded Multiple Object TrackingabstractOnline multiple object tracking (MOT) is a challenging problem in complex scenes due to frequent occlusions. Most of the existing MOT methods tend to focus on addressing an individual type of occlusion, which cannot meet the requirements of real complex scenes. In this paper, we propose a unified MOT framework that combines long- and short-term prediction models for online multiple object tracking. Basically, The short-term prediction model consists of an appearance-based model and a motion-based model, aiming at exploiting the appearance and motion of objects to handle different types of occlusions jointly. Furthermore, we adopt a cubic spline interpolation as a long-term prediction model to estimate the trajectory of the target in occluded frames. To handle different lengths of occlusions, an adaptive weighted fusion model is proposed to combine the short-term prediction model, and the long-term prediction model. Experimental results on several challenging datasets demonstrate that the proposed method outperforms state-of-the-art methods. Jun Chen 0001, Mithun Mukherjee 0001, Weijian Ruan, Chao Liang 0001, Yi Yu 0001 |
GLOBECOM | 2 |
| 2021 | Illuminate Low-Light Image via Coarse-to-fine Multi-level Network
Yansheng Qiu, Jun Chen 0001, Xiao Wang 0029, Kui Jang |
MMM (1) | 2 |
| 2021 | An Improved Online Multiple Pedestrian Tracking Based on Head and Body DetectionabstractMultiple Object Tracking (MOT) is an important computer vision task which has gained increasing attention due to its academic and commercial potential. Although many researchers have proposed effective method, they failed in crowd scene. The reason is that body detection and tracking is used in existing MOT methods. In crowd scene, many detections are missed detect and many overlap body bounding boxes decrease the quality of data association. To handle this issue, this paper propsed an online novel multiple pedestrians tracking, which is based on head detection. We first fuse the head and body detection to improve the detection result. Then, we use the head detection bounding box to replace the body detection bounding box for tracking. Finally, the experimental results demonstrate the effectiveness of our proposed method and achieve best performance with the state-of-the-art MOT trackers. Jun Chen 0001, Mithun Mukherjee 0001, Haihui Wang, Dang Zhang |
MSN | 2 |
| 2021 | Remote Sensing Image Generation From AudioabstractGenerating image from other modal data has attracted much attention in cross-modal studies, since the generated image offers intuitive vision information. Unlike the previous works which generate an image from text, a novel task is introduced, generating an image from audio. However, semantic gap intrinsically exists in cross-modal data, which disturbs the generative results. In order to explore the relevance between the audio and image, a novel reranking audio-image translation method is proposed. The proposed method: 1) maps the audio and image into a uniform feature space; 2) designs an audio-audio matching network to match the related audio; and 3) adopts an audio-image matching network for every matched audio to generate a related image, and the most frequent image is voted as the final result. Extensive experiments on two remote sensing cross-modal data sets demonstrate that the proposed method can visualize the content of audio. Jun Chen 0001, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Occluded suspect search via channel-guided mechanism
Wenxin Huang, Ruimin Hu, Xiao Wang 0029, Chao Liang 0001, Jun Chen 0001 |
Neural Comput. Appl. | 5 |
| 2021 | Spatio-Spectral Feature Fusion for Low-Light Image EnhancementabstractLow-light image enhancement aims to improve an image's visual quality, which is essential for many downstream computer vision and multimedia tasks. Existing spatial-domain low-light enhancement methods barely focus on the regions containing object boundaries, which take the most informative characteristics. However, solely focusing on enhancing high-frequency details not only causes over-sharpening of an image but also leads to color distortion. In this paper, we propose a novel spatio-spectral feature fusion network (S2F2N), that involves a frequency-feature representation branch (FRB) and a spatial-feature representation branch (SRB) to learn the domain-specific representation individually. Moreover, a spatial-channel mixed attention block (MAB) is introduced to learn the joint representation of spatio-spectral features for final image relighting. Extensive experiments on several benchmark datasets demonstrate that our method can produce high fidelity results for low-light images. Yansheng Qiu, Jun Chen 0001, Zheng Wang 0007, Xiao Wang 0029, Chia-Wen Lin |
IEEE Signal Process. Lett. | 2 |
| 2021 | A Survey of Multiple Pedestrian Tracking Based on Tracking-by-Detection FrameworkabstractMultiple pedestrian tracking (MPT) has gained significant attention due to its huge potential in a commercial application. It aims to predict multiple pedestrian trajectories and maintain their identities, given a video sequence. In the past decade, due to the advancement in pedestrian detection algorithms, Tracking-by-Detection (TBD) based algorithms have achieved tremendous successes. TBD has become the most popular MPT framework, and it has been actively studied in the past decade. In this paper, we give a comprehensive survey of recent advances in TBD-based MPT algorithms. We systematically analyze the existing TBD-based algorithms and organize the survey into four major parts. At first, this survey draws a timeline to introduce the milestones of TBD-based works which briefly reviews the development of the existing TBD-based methods. Second, the main procedures of the TBD framework are summarized, and each stage in the procedure is described in detail. Afterward, this survey analyzes the performance of existing TBD-based algorithms on MOT challenge datasets and discusses the factors that affect tracking performance. Finally, open issues and future directions in the TBD framework are discussed. Jun Chen 0001, Chao Liang 0001, Weijian Ruan, Mithun Mukherjee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Expression-Aware Face Reconstruction via a Dual-Stream NetworkabstractRecently, 3D face reconstruction from a single image has achieved promising progress by adopting the 3D Morphable Model (3DMM). However, face images taken in-the-wild usually involve expressions with a large range of variety. This poses difficulty to use 3DMM to represent such various facial expressions owing to the limited expressive ability of its linear model, thereby resulting in distortion and ambiguity in local facial regions. To tackle this problem, we present a novel dual-stream network composed of a geometry stream and a texture stream to deal with expression variations. Specifically, in the geometry stream, we propose novel Attribute Spatial Maps (ASMs) to decompose a face into the identity and expression attributes and then separately record the essential spatial information of the two facial attributes in the 2D image space. This avoids the interaction between the two attributes, thus preserving the identity information and further improving the ability of coping with expression variations. In the texture stream, we propose to generate facial appearance with realistic texture and canonical layout by our Semantic Region Stylization Mechanism (SRSM), that transfers the style from an input face to a 3DMM albedo map in a region-adaptive manner. Moreover, we also propose a Shared Semantic Region Prediction Module (SSRPM) to explore the common correspondence of semantic regions between the above two face texture representations. Both quantitative and qualitative evaluations on public datasets demonstrate the effectiveness of our approach in face reconstruction under expression variations. Xiaoyu Chai, Jun Chen 0001, Chao Liang 0001, Dongshu Xu, Chia-Wen Lin |
IEEE Trans. Multim. | 2 |
| 2021 | Correlation Discrepancy Insight Network for Video Re-identificationabstractVideo-based person re-identification (ReID) aims at re-identifying a specified person sequence from videos that were captured by disjoint cameras. Most existing works on this task ignore the quality discrepancy across frames by using all video frames to develop a ReID method. Additionally, they adopt only the person self-characteristic as the representation, which cannot adapt to cross-camera variation effectively. To that end, we propose a novel correlation discrepancy insight network for video-based person ReID, which consists of an unsupervised correlation insight model (CIM) for video purification and a discrepancy description network (DDN) for person representation. Concretely, CIM is constructed by using kernelized correlation filters to encode person half-parts, which evaluates the frame quality by the cross correlation across frames for selecting discriminative video fragments. Furthermore, DDN exploits the selected video fragments to generate a discrepancy descriptor using a compression network, which aims at employing the discrepancies with other persons’ to facilitate the representation of the target person rather than only using the self-characteristic. Due to the advantage in handling cross-domain variation, the discrepancy descriptor is expected to provide a new pattern for the object representation in cross-camera tasks. Experimental results on three public benchmarks demonstrate that the proposed method outperforms several state-of-the-art methods. Weijian Ruan, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Wu Liu 0005, Jun Chen 0001, Jiayi Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2020 | Manifold Projection for Adversarial Defense on Face Recognition
Jianli Zhou, Chao Liang 0001, Jun Chen 0001 |
ECCV (30) | 3 |
| 2020 | Image Super-Resolution Using Residual Global Context NetworkabstractRecent studies have showed that convolutional neural networks (CNN) can effectively improve the performance of single image super-resolution (SR). However, previous methods rarely considered long-range dependencies between pixels and channel-wise interdependencies at the same time. They ignores the fact that natural images have strong internal data repetition which requires the network to capture long-range dependencies between pixels and considering the interdepen-dencies between channels can better exploit the input information of the network. In addition, although past studies have proved that deep convolutional neural network benefit the performance of image super-resolution, it also means that the network needs more memory consumption and higher computational complexity. To solve these problem,we introduce Global Context block (GCB) and design a comparative shallow network called Residual Global Context Networks (RGC-N). It achieves a better trade-off between the amount of parameter and the quality of image reconstruction. Extensive experiments demonstrate that the proposed method is superior to the state-of-the-art methods. Kuangye Liu, Zhen Han 0002, Junkui Chen, Chunlei Liu 0006, Jun Chen 0001, Zhongyuan Wang 0001 |
ICASSP | 5 |
| 2020 | Expression-Aware Face Reconstruction Via A Dual-Stream NetworkabstractRecently, 3D face reconstruction from a single image has achieved promising results by adopting the 3D Morphable Model (3DMM). However, as face images in-the-wild have various expressions, it is difficult for 3DMM to handle diverse facial expressions with a large range of variations, due to the limited expressive ability of its linear model, thereby resulting in distortion and ambiguity on facial local regions. To tackle this issue, we present a novel dual-stream network to deal with expression variations. Specifically, in the geometry stream, we propose novel Attribute Spatial Maps to record the spatial information of facial identity and expression attributes in the 2D image space separately. This avoids the interaction between the two attributes, thus keeping the identity information and further improving the ability to cope with expression changes. In the texture stream, we utilize the 3DMM albedo map to a style transfer based method for synthesizing facial appearance, which results in expression-irrelevant as well as realistic face textures. Both quantitative and qualitative evaluations on public datasets demonstrate the ability of our approach to achieve comparable results in face reconstructions under expression variations. Xiaoyu Chai, Jun Chen 0001, Chao Liang 0001, Dongshu Xu, Chia-Wen Lin |
ICME | 2 |
| 2020 | When Pedestrian Detection Meets Nighttime Surveillance: A New BenchmarkabstractPedestrian detection at nighttime is a crucial and frontier problem in surveillance, but has not been well explored by the computer vision and artificial intelligence communities. Most of existing methods detect pedestrians under favorable lighting conditions (e.g. daytime) and achieve promising performances. In contrast, they often fail under unstable lighting conditions (e.g. nighttime). Night is a critical time for criminal suspects to act in the field of security. The existing nighttime pedestrian detection dataset is captured by a car camera, specially designed for autonomous driving scenarios. The dataset for nighttime surveillance scenario is still vacant. There are vast differences between autonomous driving and surveillance, including viewpoint and illumination. In this paper, we build a novel pedestrian detection dataset from the nighttime surveillance aspect: NightSurveillance1. As a benchmark dataset for pedestrian detection at nighttime, we compare the performances of state-of-the-art pedestrian detectors and the results reveal that the methods cannot solve all the challenging problems of NightSurveillance. We believe that NightSurveillance can further advance the research of pedestrian detection, especially in the field of surveillance security at nighttime. Xiao Wang 0029, Jun Chen 0001, Zheng Wang 0007, Wu Liu 0005, Shin'ichi Satoh 0001, Chao Liang 0001, Chia-Wen Lin |
IJCAI | 2 |
| 2020 | One-Shot Face Recognition with Feature Rectification via Adversarial Learning
Jianli Zhou, Jun Chen 0001, Chao Liang 0001 |
MMM (1) | 2 |
| 2020 | Learning heterogeneous information network embeddings via relational triplet network
Xiyue Gao, Jun Chen 0001, Zexing Zhan |
Neurocomputing | 2 |
| 2020 | SIST: Online Scale-Adaptive Object tracking with Stepwise Insight
Weijian Ruan, Chao Liang 0001, Yi Yu 0001, Jun Chen 0001, Ruimin Hu |
Neurocomputing | 4 |
| 2020 | Single image de-raining via clique recursive feedback mechanism
Jun Chen 0001, Kui Jiang, Zhen Han 0002, Weijian Ruan, Zhongyuan Wang 0001, Chao Liang 0001 |
Neurocomputing | 2 |
| 2020 | Identity-Aware Face Super-Resolution for Low-Resolution Face RecognitionabstractAlthough deep learning-based face recognition techniques have achieved amazing performance in recent years, low-resolution (LR) face recognition remains challenging. In this letter, we address this problem by proposing an identity-aware face super-resolution network to recover identity information of LR faces. To learn identity-aware features effectively, the identity features are explicitly disentangled to two orthogonal components: the magnitude and angle of features that project identity features to a hypersphere space. We show that the magnitude of features is related to the quality of a face. The proposed approach shows its superiority on recovering identity-related textures which are beneficial to recover identity information for recognition. Extensive experiments demonstrate the effectiveness of the proposed algorithm in LR face recognition. Jun Chen 0001, Zheng Wang 0007, Chao Liang 0001, Chia-Wen Lin |
IEEE Signal Process. Lett. | 2 |
| 2020 | S3D: Scalable Pedestrian Detection via Score Scale Surface DiscriminationabstractPedestrian detection has remained an important research topic in both the computer vision and multimedia communities because of its importance in practical applications, such as driving assistance and video surveillance. Existing methods compare the response score with a fixed threshold to determine whether a candidate region contains pedestrians and produce dissatisfactory results that contain either missed detections or false detections, which are difficult to balance. This situation has a serious impact under the condition of variable scale. This paper investigates the functional relationship between the scores and scales of pedestrians. By designing experiments with multiple scales, we have found a discriminant surface in the score scale space. Pedestrians can be distinguished at various scale levels according to their locations on the discriminant surface. The proposed approach is evaluated using four challenging pedestrian detection datasets, including Caltech, INRIA, ETH, and KITTI, and the superior experimental results are achieved when compared with baseline methods. Xiao Wang 0029, Chao Liang 0001, Chen Chen 0001, Jun Chen 0001, Zheng Wang 0007, Zhen Han 0002, Chunxia Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Listen, Look, and Find the One: Robust Person Search with Multimodality IndexabstractPerson search with one portrait, which attempts to search the targets in arbitrary scenes using one portrait image at a time, is an essential yet unexplored problem in the multimedia field. Existing approaches, which predominantly depend on the visual information of persons, cannot solve problems when there are variations in the person’s appearance caused by complex environments and changes in pose, makeup, and clothing. In contrast to existing methods, in this article, we propose an associative multimodality index for person search with face, body, and voice information. In the offline stage, an associative network is proposed to learn the relationships among face, body, and voice information. It can adaptively estimate the weights of each embedding to construct an appropriate representation. The multimodality index can be built by using these representations, which exploit the face and voice as long-term keys and the body appearance as a short-term connection. In the online stage, through the multimodality association in the index, we can retrieve all targets depending only on the facial features of the query portrait. Furthermore, to evaluate our multimodality search framework and facilitate related research, we construct the Cast Search in Movies with Voice (CSM-V) dataset, a large-scale benchmark that contains 127K annotated voices corresponding to tracklets from 192 movies. According to extensive experiments on the CSM-V dataset, the proposed multimodality person search framework outperforms the state-of-the-art methods. Xiao Wang 0029, Wu Liu 0005, Jun Chen 0001, Xiaobo Wang 0001, Chenggang Yan 0001, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | Rain Streak Removal via Multi-scale Mixture Exponential Power ModelabstractRain streaks severely hamper the visible performance of the outdoor surveillance videos, which becomes an attractive issue in recent computer vision research. Existing methods usually encode rain streaks into Gaussian Mixture Model (GM-M). However, the limited number of Gaussian components in the GMM compromises the ability of the model in fitting real noise, such as sparse noise, which is exactly the characteristic of the rain streaks. In this paper, a novel model named Mixture Exponential Power Model (MEPM) is exploited. It sets multiple Laplace noise components and expands the representation capability for the sparse noise. Moreover, considering that the rain streaks in a video occur in different distances from the camera, we encode rain streaks into Multi-scale Mixture Exponential Power Model. The model is opti-mized by expectation-maximization (EM) algorithm and La-grange multiplier strategy. Experiments are implemented on synthetic and real rain videos and verify the superiority of the proposed method, compared with state-of-the-art methods. Jun Chen 0001, Zhen Han 0002, Mingfu Xiong, Chao Liang 0001, Zhongyuan Wang 0001 |
ICASSP | 2 |
| 2019 | Cross-view Identical Part Area Alignment for Person Re-identificationabstractPerson re-identification aims to associate images captured by non-overlapping cameras. It is a challenging task because images are often in different conditions such as background clutter, illumination variation, viewpoint changes and different camera settings. Viewpoint changes and pose variations often cause body part self-occlusion and misalignment. To deal with the problem, local features from human body parts are extracted. However, with viewpoint changes, the body parts also rotate horizontally. It is inappropriate to extract feature from entire area of body parts directly because the visible surface of body parts would turn away if viewpoint changes. Comparing identical areas provides a new way to pay attention to the details of person images. In this paper, we propose a Rotation Invariant Network to find the identical areas in cross-view images to extract robust local features. Extensive experiment show the effectiveness of our method on public datasets including CUHK03, Market1501 and DukeMTMC. Dongshu Xu, Jun Chen 0001, Chao Liang 0001, Zheng Wang 0007, Ruimin Hu |
ICASSP | 2 |
| 2019 | POINet: Pose-Guided Ovonic Insight Network for Multi-Person Pose TrackingabstractMulti-person pose tracking aims to jointly estimate and track multi-person keypoints in the unconstrained videos. The most popular solution to this task follows the tracking-by-detection strategy that relies on human detection and data association. While human detection has been boosted by deep learning, existing works mainly exploit several separated stages with hand-crafted metrics to realize data association, leading to great uncertainty and feeble adaption in complex scenes. To handle these problems, we propose an end-to-end pose-guided ovonic insight network (POINet) for the data association in multi-person pose tracking, which jointly learns feature extraction, similarity estimation, and identity assignment. Specifically, we design a pose-guided representation network to integrate pose information into hierarchical convolutional features, generating a pose-aligned person representation for person, which helps handle partial occlusions. Moreover, we propose an ovonic insight network to adaptively encode the cross-frame identity transformation, which can cope with the tough tracking cases of person leaving and entering the scene. In general, the proposed POINet provides a new insight to realize multi-person pose tracking in an end-to-end fashion. Extensive experiments conducted on the PoseTrack benchmark demonstrate that our POINet outperforms the state-of-the-art methods. Weijian Ruan, Wu Liu 0005, Qian Bao, Jun Chen 0001, Yuhao Cheng, Tao Mei 0001 |
ACM Multimedia | 4 |
| 2019 | Action Recognition Using Visual Attention with Reinforcement Learning
Jun Chen 0001, Ruimin Hu, Zengmin Xu |
MMM (2) | 2 |
| 2019 | Meta-Circuit machine: Inferencing human collaborative relationships in heterogeneous information networks
Xiyue Gao, Jun Chen 0001, Nian Huai |
Inf. Process. Manag. | 2 |
| 2019 | Action recognition for depth video using multi-view dynamic images
Yang Xiao 0007, Jun Chen 0001, Yancheng Wang 0002, Zhiguo Cao 0001, Joey Tianyi Zhou, Xiang Bai |
Inf. Sci. | 2 |
| 2019 | Person re-identification with multiple similarity probabilities using deep metric learning for efficient smart security applications
Mingfu Xiong, Dan Chen 0001, Jun Chen 0001, Jingying Chen 0001, Benyun Shi, Chao Liang 0001, Ruimin Hu |
J. Parallel Distributed Comput. | 3 |
| 2019 | Multi-Correlation Filters With Triangle-Structure Constraints for Object TrackingabstractCorrelation filters (CFs) have been extensively used in tracking tasks due to their high efficiency although most of them regard the tracked target as a whole and are minimally effective in handling partial occlusion. In this study, we incorporate a part-based strategy into the framework of CFs and propose a novel multipart correlation tracker with triangle-structure constraints. Specifically, we train multiple CFs for the global object and local parts, which are then jointly applied to obtain the correlation response of any candidate during tracking. The tracker is robust in handling partial occlusion because of the use of part-based representation. The remaining global representation can contribute reliable cues in cases wherein several local filters drift away in a specific scene. We further propose a triangle-structure model to measure the structural similarity of candidates. The model employs multiple triangles to determine the spatial relationship among parts and helps constrain the location of the target. Moreover, we introduce an effective part selection scheme based on energy and integrity, which is generally applicable to part-tracking models. Extensive experiments on two public benchmarks demonstrate the superiority of the proposed method over the state-of-the-art approaches. Weijian Ruan, Jun Chen 0001, Yi Wu 0001, Jinqiao Wang, Chao Liang 0001, Ruimin Hu, Junjun Jiang |
IEEE Trans. Multim. | 2 |
| 2019 | Semisupervised Discriminant Multimanifold Analysis for Action RecognitionabstractAlthough recent semisupervised approaches have proven their effectiveness when there are limited training data, they assume that the samples from different actions lie on a single data manifold in the feature space and try to uncover a common subspace for all samples. However, this assumption ignores the intraclass compactness and the interclass separability simultaneously. We believe that human actions should occupy multimanifold subspace and, therefore, model the samples of the same action as the same manifold and those of different actions as different manifolds. In order to obtain the optimum subspace projection matrix, the current approaches may be mathematically imprecise owe to the badly scaled matrix and improper convergence. To address these issues in unconstrained convex optimization, we introduce a nontrivial spectral projected gradient method and Karush-Kuhn-Tucker conditions without matrix inversion. Through maximizing the separability between different classes by using labeled data points and estimating the intrinsic geometric structure of the data distributions by exploring unlabeled data points, the proposed algorithm can learn global and local consistency and boost the recognition performance. Extensive experiments conducted on the realistic video data sets, including JHMDB, HMDB51, UCF50, and UCF101, have demonstrated that our algorithm outperforms the compared algorithms, including deep learning approach when there are only a few labeled samples. Zengmin Xu, Ruimin Hu, Jun Chen 0001, Chen Chen 0001, Junjun Jiang, Jiaofen Li |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Reinforcing Pedestrian Parsing on Small Scale Dataset
Jun Chen 0001, Junjun Jiang, Ruimin Hu |
MMM (1) | 2 |
| 2017 | Transferring clothing parsing from fashion dataset to surveillanceabstractIn this paper we address the problem of automatic clothing parsing in surveillance video with the information from user-generated tags such as “jeans” and “T-shirt”. Although clothing parsing has achieved great success in fashion clothing, it is quite challenging to parse clothing in practical surveillance conditions due to complicated environmental interferences, such as illumination change, scale zooming, viewpoint variation and etc. Our method is developed to capture the clothing information from the fashion field and apply it to surveillance domain by weakly-supervised transfer learning. Most of attribute labels in surveillance images convey strong location information, which can be considered as weak labels to deal with the transfer method. Both quantitative and qualitative experiments conducted on practical surveillance datasets have shown the effectiveness of the proposed method. Jun Chen 0001, Chao Liang 0001, Wenhua Fang, Xiaoyuan Jing, Ruimin Hu |
ICASSP | 2 |
| 2017 | Action recognition with gradient boundary convolutional networkabstractDeep learning features for video action recognition are usually learned from RGB/gray images, image gradients, and optical flows. The single modality of the input data can describe one characteristic of the human action such as appearance structure or motion information. In this paper, we propose a high efficient gradient boundary convolutional network (ConvNet) to simultaneously learn spatio-temporal feature from the single modality data of gradient boundaries. The gradient boundaries represent both local spacial structure and motion information of action video. The gradient boundaries also have less background noise compared to RGB/gray images and image gradients. Extensive experiments are conducted on two popular and challenging action benchmarks, the UCF101 and the HMDB51 action datasets. The proposed deep gradient boundary feature achieves competitive performances on both benchmarks. Jun Chen 0001, Chen Chen 0001, Ruimin Hu |
ICIP | 2 |
| 2017 | Non-rigid feature matching for image retrieval using global and local regularizationsabstractIn this paper, we propose a probabilistic method for feature matching of near-duplicate images undergoing non-rigid transformations. We start by creating a set of putative correspondences based on the feature similarity, and then focus on removing outliers from the putative set and estimating the transformation as well. This is formulated as a maximum likelihood estimation of a Bayesian model with latent variables indicating whether matches in the putative set are inliers or outliers. We impose the non-parametric global geometrical constraints on the correspondence using Tikhonov regularizers in a reproducing kernel Hilbert space. We also introduce a local geometrical constraint to preserve local structures among neighboring feature points. The problem is solved by using the Expectation Maximization algorithm, and the closed-form solution of the transformation is derived in the maximization step. Moreover, a fast implementation based on sparse approximation is given which reduces the method computation complexity to linearithmic without performance sacrifice. Extensive experiments on real near-duplicate images for both feature matching and image retrieval demonstrate accurate results of the proposed method which outperforms current state-of-the-art methods, especially in case of severe outliers. Yong Ma 0001, Huabing Zhou, Jun Chen 0001, Jingshu Shi, Zhongyuan Wang 0001 |
ICME | 3 |
| 2017 | Object tracking via online trajectory optimization with multi-feature fusionabstractThe goal of object tracking is to estimate the state and trajectory of an interested target in a video sequence, thus both spatial and temporal information are of critical importance for tracking. However, most existing trackers usually determine targets just by the judgement like confidence from a single frame, which tend to treat tracking as a static detecting problem while neglecting the spatial-temporal relationship. In this paper, we propose a novel tracking method of online trajectory optimization with multi-feature fusion (TOFF). Considering the trajectory continuity, the accurate targets are determined by estimating the optimal short trajectories over all the video fragments, which can be obtained by the observation models that are iteratively updated based on the selected most reliable proposals. By employing the structured samples instead of binary-labeled samples, we construct a structured output model with representing descriptor of multi-feature fusion as the basic tracker. Extensive experiments on various challenging image sequences demonstrate the superiority of our method to several state-of-the-art methods. Weijian Ruan, Jun Chen 0001, Chao Liang 0001, Yi Wu 0001, Ruimin Hu |
ICME | 2 |
| 2017 | Low-resolution pedestrian detection via a novel resolution-score discriminative surfaceabstractPedestrian detection, as an important task in video surveillance and forensics applications, has been widely studied. However, its performance is unsatisfactory especially in the low resolution conditions. In realistic scenarios, the size of pedestrians in the images is often small, and detection can be challenging. To solve this problem, this paper proposes a novel resolution-score discriminative surface method to investigate the variation behaviors of detection scores under different pedestrian and non-pedestrian image resolutions. The discriminative surface consists of a series of positive and negative resolution-score lines, and each of them is a connected line to depict the variation relationship between pedestrian's detection scores under various image resolutions. On this basis, the resolution-score discriminative surface can classify a resolution-score line as a pedestrian or not according to whether it lies in the positive or the negative region. Experimental results on two public datasets and one campus surveillance dataset demonstrate the effectiveness of the proposed method. Xiao Wang 0029, Jun Chen 0001, Chao Liang 0001, Chen Chen 0001, Zheng Wang 0007, Ruimin Hu |
ICME | 2 |
| 2017 | Efficient Pedestrian Detection in the Low Resolution via Sparse Representation with Sparse Support Regression
Wenhua Fang, Jun Chen 0001, Ruimin Hu |
PAKDD (2) | 2 |
| 2017 | Action recognition by saliency-based dense sampling
Zengmin Xu, Ruimin Hu, Jun Chen 0001, Chen Chen 0001, Qingquan Sun |
Neurocomputing | 3 |
| 2016 | Person Re-Identification via Multiple Coarse-to-Fine Deep MetricsabstractPerson re-identification, aiming to identify images of the same person from various cameras views in different places, has attracted a lot of research interests in the field of artificial intelligence and multimedia. As one of its popular research directions, the metric learning method plays an important role for seeking a proper metric space to generate accurate feature comparison. However, the existing metric learning methods mainly aim to learn an optimal distance metric function through a single metric, making them difficult to consider multiple similar relationships between the samples. To solve this problem, this paper proposes a coarse-to-fine deep metric learning method equipped with multiple different Stacked Auto-Encoder (SAE) networks and classification networks. In the perspective of the human's visual mechanism, the multiple different levels of deep neural networks simulate the information processing of the brain's visual system, which employs different patterns to recognize the character of objects. In addition, a weighted assignment mechanism is presented to handle the different measure manners for final recognition accuracy. The experimental results conducted on two public datasets, i.e., VIPeR and CUHK have shown the prospective performance of the proposed method. Mingfu Xiong, Jun Chen 0001, Zheng Wang 0007, Zhongyuan Wang 0001, Ruimin Hu, Chao Liang 0001, Daming Shi 0001 |
ECAI | 2 |
| 2016 | Multiple instance discriminative dictionary learning for action recognitionabstractAction recognition from video is a prominent research area in computer vision, with far-reaching applications. Current state-of-the-art action recognition methods is Fisher Vector (FV) coding model based on spatio-temporal local features. Though high dimensional local features have more representative, the high dimensions are challenge for the dictionary learning of FV model. This paper proposes a Multiple Instance Discriminative Dictionary Learning (MIDDL) method for action recognition. We introduce cross-validation method in multiple instance learning procedure, which prevents training from prematurely locking onto erroneous initial instances. In order to balance the positive instance number between positive bags, only the top ranked instances are labeled as positive in the step of iterative training classifiers. Taking these classifiers as discriminative visual words, we get the video global representation based on classifier response. The experimental results demonstrate the effectiveness of applying the learned discriminative classifiers as visual word on challenging action data sets, i.e. UCF50 and HMDB51. Jun Chen 0001, Zengmin Xu, Ruimin Hu |
ICASSP | 2 |
| 2016 | Depth image in-loop filter via graph cutabstractWith the ability of representing 3D scene geometry, depth maps are used to synthesize virtual views in free view video (FVV) or 3DTV. However, compression artefact of the depth images always lead to seriously geometry distortions in synthesized view, which severely affects the visual perception of 3D display. To solve the problem caused by depth artifact, the bilateral filter based method is presented to alleviate the noise by local weighted sum of neighboring pixels. However, to overcome the unstable feature of local weight of noisy pixels, we propose a novel graph cuts algorithm for depth filter with the constraint of corresponding structure information in color image. As a depth in-loop filter, the filter is incorporated into the framework of H.264/MVC. The proposed approach offers 0.5dB and 0.8dB average PSNR gains in terms of video rendering quality and depth coding efficiency comparing with the state-of-the-art. Liguo Zhou, Zhongyuan Wang 0001, Youming Fu, Jun Chen 0001, Rui Xiang, Shizheng Wang |
ICIP | 4 |
| 2016 | Boosted local classifiers for visual trackingabstractMost existing discriminative tracking methods model a target object as a whole and train a tracker based on holistic templates, which cannot effectively deal with partial occlusions. Instead, in this paper, by treating the target as a collection of local patches, we propose a novel tracking approach based on boosted local classifiers. Initially, a set of local patches are sampled to train a set of local classifiers, and the weight of each classifier is given based on the estimated error. In addition, the positive examples and negative examples are sampled for model update with two constraints during the tracking process, which helps obtain more negatives for updating the appearance model and improve the updating efficiency. With updating the weights of local classifiers based on the temporal stability, the tracker can effectively handle partial occlusions. Extensive experiments on various challenging image sequences demonstrate the superiority to several state-of-the-art methods. Weijian Ruan, Jun Chen 0001, Jinqiao Wang, Bo Luo, Ruimin Hu |
ICME | 2 |
| 2016 | Global Contrast Based Salient Region Boundary Sampling for Action Recognition
Zengmin Xu, Ruimin Hu, Jun Chen 0001 |
MMM (1) | 3 |
| 2016 | Group recursive discriminant subspace learning with image set decomposition
Fei Wu 0004, Xiaoyuan Jing, Yong-Fang Yao, Dong Yue 0001, Jun Chen 0001 |
Neural Comput. Appl. | 5 |
| 2016 | Zero-Shot Person Re-identification via Cross-View ConsistencyabstractPerson re-identification, aiming to identify images of the same person from various cameras configured in different places, has attracted much attention in the multimedia retrieval community. In this problem, choosing a proper distance metric is a crucial aspect, and many classic methods utilize a uniform learnt metric. However, their performance is limited due to ignoring the zero-shot and fine-grained characteristics presented in real person re-identification applications. In this paper, we investigate two consistencies across two cameras, which are cross-view support consistency and cross-view projection consistency. The philosophy behind it is that, in spite of visual changes in two images of the same person under two camera views, the support sets in their respective views are highly consistent, and after being projected to the same view, their context sets are also highly consistent. Based on the above phenomena, we propose a data-driven distance metric (DDDM) method, re-exploiting the training data to adjust the metric for each query-gallery pair. Experiments conducted on three public data sets have validated the effectiveness of the proposed method, with a significant improvement over three baseline metric learning methods. In particular, on the public VIPeR dataset, the proposed method achieves an accuracy rate of 42.09% at rank-1, which outperforms the state-of-the-art methods by 4.29%. Zheng Wang 0007, Ruimin Hu, Chao Liang 0001, Yi Yu 0001, Junjun Jiang, Mang Ye, Jun Chen 0001, Qingming Leng |
IEEE Trans. Multim. | 7 |
| 2016 | Person Reidentification via Ranking Aggregation of Similarity Pulling and Dissimilarity PushingabstractPerson reidentification is a key technique to match different persons observed in nonoverlapping camera views. Many researchers treat it as a special object-retrieval problem, where ranking optimization plays an important role. Existing ranking optimization methods mainly utilize the similarity relationship between the probe and gallery images to optimize the original ranking list, but seldom consider the important dissimilarity relationship. In this paper, we propose to use both similarity and dissimilarity cues in a ranking optimization framework for person reidentification. Its core idea is that the true match should not only be similar to those strongly similar galleries of the probe, but also be dissimilar to those strongly dissimilar galleries of the probe. Furthermore, motivated by the philosophy of multiview verification, a ranking aggregation algorithm is proposed to enhance the detection of similarity and dissimilarity based on the following assumption: the true match should be similar to the probe in different baseline methods. In other words, if a gallery blue image is strongly similar to the probe in one method, while simultaneously strongly dissimilar to the probe in another method, it will probably be a wrong match of the probe. Extensive experiments conducted on public benchmark datasets and comparisons with different baseline methods have shown the great superiority of the proposed ranking optimization method. Mang Ye, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Qingming Leng, Chunxia Xiao, Jun Chen 0001, Ruimin Hu |
IEEE Trans. Multim. | 7 |
| 2015 | Exploiting effects of parts in fine-grained categorization of vehiclesabstractFine-grained categorization has become a hot topic in computer vision. Based on the theory that part information is crucial for fine-grained categorization, we proposed a part-based categorization method for vehicles, consisting vehicle parts localization, part-based vehicle representation and classification. There were three contributions we made in this work: 1) we analyzed discriminative powers of parts for fine-grained categorization; 2) we proposed a frame of how to integrate discriminative powers of parts into categorization, and proved that it can achieve better performance than treating every part equally; 3) we provided an annotated dataset with parts for vehicle categorization. Ruimin Hu, Jun Xiao 0004, Jing Xiao 0004, Jun Chen 0001 |
ICIP | 6 |
| 2015 | How much bandwidth does surveillance system require?abstractOne of the main challenges in surveillance systems lies in the massive amount of video involved in providing potential key content with sufficient resolution. This paper shows that there exists a sweet spot, which we term critical video quality that can be used to reduce bitrate of video transmission without significantly affecting the accuracy of the surveillance tasks. We present a new city surveillance dataset which was divided into three types of scenarios, and we analyze subjective data collected via human subjective testing for object identification. These data are then used to create objective measurements (models) to drive video compression ratio based on the detection probability. The main idea is to find out the lowest bitrate of video transmission while maximizes the probability of detecting objects which are carried or abandoned. Experiment results shown that our generalized models can predict acceptable video quality for object identification in rational ways. Zengmin Xu, Ruimin Hu, Jun Chen 0001 |
ICIP | 3 |
| 2015 | Specific Person Retrieval via Incomplete Text DescriptionabstractSearching for specific persons from surveillance videos captured by different cameras, is a key yet under-addressed challenge in multimedia system. Related person retrieval works mainly focus on searching person by visual appearance, known as person re-identification. However, the initial visual image may not be available in some practical applications. For example, the criminal is described by a text description indirectly, "A young woman wearing a red casual with a backpack", the traditional methods can not conquer this issue. Based on a set of pre-defined attributes that the text description query can be transformed to an attribute vector, thus can be used to retrieval in the gallery set. And yet, the user-provided attributes are sometimes incomplete. This new issue is defined as Specific Person Retrieval via Incomplete Text Description. In this paper, we conduct a specific attribute completion to enrich the original text query and generate a more expressive attribute vector. Then, a pairwise-based metric learning is introduced for completed attribute vectors. Extensive experiments conducted on two benchmark datasets have shown our superior performance. Mang Ye, Chao Liang 0001, Zheng Wang 0007, Qingming Leng, Jun Chen 0001, Jun Liu 0036 |
ICMR | 5 |
| 2015 | Ranking Optimization for Person Re-identification via Similarity and DissimilarityabstractPerson re-identification is a key technique to match different persons observed in non-overlapping camera views.Many researchers treat it as a special object retrieval problem, where ranking optimization plays an important role. Existing ranking optimization methods utilize the similarity relationship between the probe and gallery images to optimize the original ranking list in which dissimilarity relationship is seldomly investigated. In this paper, we propose to use both similarity and dissimilarity cues in a ranking optimization framework for person re-identification. Its core idea is based on the phenomenon that the true match should not only be similar to the strong similar samples of the probe but also dissimilar to the strong dissimilar samples. Extensive experiments have shown the great superiority of the proposed ranking optimization method. Mang Ye, Chao Liang 0001, Zheng Wang 0007, Qingming Leng, Jun Chen 0001 |
ACM Multimedia | 5 |
| 2015 | Object Detection in Low-Resolution Image via Sparse Representation
Wenhua Fang, Jun Chen 0001, Chao Liang 0001, Xiao Wang 0029, Yuanyuan Nan, Ruimin Hu |
MMM (1) | 2 |
| 2015 | Sparsity-Based Occlusion Handling Method for Person Re-identification
Bingyue Huang, Jun Chen 0001, Chao Liang 0001, Zheng Wang 0007, Kaimin Sun |
MMM (2) | 2 |
| 2015 | Coupled Discriminant Multi-Manifold Analysis with Application to Low-Resolution Face Recognition
Junjun Jiang, Ruimin Hu, Zhen Han 0002, Liang Chen 0026, Jun Chen 0001 |
MMM (1) | 5 |
| 2015 | Coupled-View Based Ranking Optimization for Person Re-identification
Mang Ye, Jun Chen 0001, Qingming Leng, Chao Liang 0001, Zheng Wang 0007, Kaimin Sun |
MMM (1) | 2 |
| 2015 | Car re-identification from large scale images using semantic attributesabstractCar re-identification, searching a specific car object from a large-scale car image database, is investigated in this paper. Previous work mainly focuses on fixed pose and overlooks the special appearance. However, avoiding matching other poses would lead to coarse results of the car retrieval. And some special attributes like individual paintings which are greatly helpful for car retrieval have not drawn enough attention. This paper addresses these problems through multi-poses matching and re-ranking based on special attributes. Our core idea lies in query expansion method that can capture weighted attributes to build the retrieval model, which allows us to estimate invisible attributes by the visible ones to construct complete attributes vectors to car retrieval in any poses. Furthermore, we divide all attributes into two groups, special attributes and common attributes. Here special attributes represent the abnormal appearance like individual paintings or car damage while common attributes denote the intrinsic appearance of car. Using special attributes to re-rank results turns out to be beneficial to improve the retrieval performance. In the end, the experiments demonstrate the effectiveness of our approach on the car datasets. Chao Liang 0001, Wenhua Fang, Da Xiang, Chengping Ren, Jun Chen 0001 |
MMSP | 7 |
| 2015 | An empirical study on software defect prediction with a simplified metric set
Bing Li 0010, Jun Chen 0001, Yutao Ma |
Inf. Softw. Technol. | 4 |
| 2015 | Person re-identification with content and context re-ranking
Qingming Leng, Ruimin Hu, Chao Liang 0001, Jun Chen 0001 |
Multim. Tools Appl. | 5 |
| 2014 | Pedestrian detection from salient regionsabstractClassic algorithms of pedestrian detection usually locate the latent position via sliding window techniques, which resize the matching window and/or original images at different scales and scan the image. However, this method has two main drawbacks. First, resizing at a fix rate cannot search through the whole scale space, resulting in the failure of accurate object location. Second, resizing and scanning at various scales is usually time-consuming, which is improper for practical applications. To conquer the above difficulties, a novel pedestrian detection method with salient information is proposed. In this paper, the salient detection model and the traditional covariance matrix descriptor are combined in a Bayesian framework to detect pedestrians in the still image. Finally, the efficiency of our approach compared with state-of-the-art results is demonstrated on the public INRIA dataset. Xiao Wang 0029, Jun Chen 0001, Wenhua Fang, Chao Liang 0001, Chunjie Zhang 0001, Ruimin Hu |
ICIP | 2 |
| 2014 | Face hallucination via re-identified K-nearest neighbors embeddingabstractBased on locally linear embedding (LLE) manifold learning theory, which assumes that the low-resolution (LR) manifold and high-resolution (HR) manifold spaces share the same local geometry structure, neighbor embedding based super-resolution(SR) methods search K-nearest neighbors(K-NN) of LR patch, then use the counterpart HR patches to estimate HR patch. The primary issue of these methods is how to search the optimal K-NN. However, due to the “one-to-many” mapping between the LR image and HR ones in practice, the neighborhood relationship of the LR patch in LR space is very different with its HR counterpart's. In this paper, we explore a novel and effective re-identified K-NN(RIKNN) method to search neighbors of LR patch by taking into consideration the neighbor information in the HR space. It searches K-NN of LR patch in the LR space and then refines the searching results by re-identifying in the HR space, thus giving rise to accurate K-NN and improvement performance. Experimental results with application to face hallucination demonstrate that our method outperforms state of the art in terms of subjective and objective results and computational complexity. Shenming Qu, Ruimin Hu, Junjun Jiang, Zhongyuan Wang 0001, Jun Chen 0001 |
ICME | 6 |
| 2014 | Collaborative Web Service QoS Prediction on Unbalanced Data DistributionabstractQoS prediction is critical to Web service selection and recommendation. This paper proposes a collaborative approach to quality-of-service (QoS) prediction of web services on unbalanced data distribution by utilizing the past usage history of service users. It avoids expensive and time-consuming web service invocations. There existed several methods which search top-k similar users or services in predicting QoS values of Web services, but they did not consider unbalanced data distribution. Then, we improve existed methods in similar neighbors' selection by sampling importance resampling. To validate our approach, large-scale experiments are conducted based on a real-world Web service dataset, WSDream. The results show that our proposed approach achieves higher prediction accuracy than other approaches. Bing Li 0010, Lulu He, Jun Chen 0001 |
ICWS | 5 |
| 2013 | Coupled-layer neighbor embedding for surveillance face hallucinationabstractAs the face image captured by a surveillance camera is typically very low-resolution (LR), blurred and noisy, traditional neighbor embedding method considers only one manifold (the LR image manifold) and fails very often to reliably estimate the intention geometrical structure. In this paper, we introduce the notion of neighbor embedding from the LR image manifold and the high-resolution (HR) one simultaneously and propose a novel neighbor embedding model, termed the coupled-layer neighbor embedding (CLNE), for surveillance face hallucination. CLNE differs substantially from other neighbor embedding models in that the former has two layers: the LR layer and the the HR layer. The LR layer in this model is the local geometrical structure of the LR patch manifold, which is characterized by the reconstruction weights; the HR layer in this model is a set of HR training patches that guide the K-nearest neighbor (K-NN) searching and geometrically constrain the reconstruction weights. By this coupled constraint paradigm between the adaptation of the LR layer and the HR one, CLNE can achieve a more robust neighbor embedding through the significant degradation process. Indeed, the experimental results confirm that our method outperforms the related state-of-the-art methods by having better objective values as well as better visual results. Junjun Jiang, Ruimin Hu, Liang Chen 0026, Zhen Han 0002, Tao Lu 0001, Jun Chen 0001 |
ICIP | 6 |
| 2013 | Locality-constraint iterative neighbor embedding for face hallucinationabstractBased on the assumption that low-resolution (LR) and high-resolution (HR) patch manifolds are locally isometric, the neighbor embedding based super-resolution algorithms try to preserve the local geometry of the patch manifold for the reconstructed HR patch manifold. However, due to “one-to-many” mappings between LR and HR images, the neighborhood relationship of the LR patch manifold can't reflect the inherent data structure. In this paper, we explore the data structure by both considering the LR patch and HR patch manifolds instead of only considering one manifold (LR patch manifold). By incorporating the position prior of face and local geometry of HR patch manifold, we propose an improved neighbor embedding method to face hallucination, namely locality-constraint iterative neighbor embedding (LINE), in which we iteratively update the K-nearest neighbors (K-NN) and reconstruction weights based on the result (the hallucinated HR patch) from previous iteration, giving rise to improved performance compared with traditional neighbor embedding algorithms. Experimental results with application to face hallucination on simulated LR face images and real world ones demonstrate the effectiveness of the proposed method. Junjun Jiang, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001, Tao Lu 0001, Jun Chen 0001 |
ICME | 6 |
| 2013 | Bidirectional ranking for person re-identificationabstractThis paper proposes a simple but efficient bidirectional ranking method to improve person re-identification results across non-overlapping cameras. Previous methods treat person reidentification as a special object retrieval problem, and compute the final rank result purely based on a unidirectional matching between the probe and all gallery images. However, the expected person image may be excluded from the probe's ??-nearest neighbor due to appearance changes caused by variations in illuminations, poses, viewpoints and occlusion. To solve the above problem, our method queries every gallery image in a new gallery composed of the original probe image and other gallery images, and revises the initial query result in accordance with both content and context similarities between bidirectional ranking lists. A latent assumption of our method is that images of the same person should not only have similar visual content, known as content similarity, but also possess similar k-nearest neighbors, known as context similarity. Extensive experiments conducted on a series of standard data sets have validated the effectiveness of our proposed method with an average improvement of 5-10% over original baseline methods. Qingming Leng, Ruimin Hu, Chao Liang 0001, Jun Chen 0001 |
ICME | 5 |