EDBT 2026 Demo / reviewers in the wild / expert
Weijian Ruan
dblp:184/9444
· DBLP profile ↗
31ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0003-3710-8739ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 9 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SFIR: Optimizing spatial and frequency domains for image restoration
Yubin Gu, Siting Chen, Jiayi Ji, Xiaoshuai Sun, Weijian Ruan, Rongrong Ji |
Pattern Recognit. | 6 |
| 2026 | Wavelet-based learning and optimized sampling for image deraining
Yubin Gu, Xiaoshuai Sun, Jiayi Ji, Weijian Ruan, Rongrong Ji |
Pattern Recognit. | 5 |
| 2026 | RWKV4Rec: RWKV-Based Personalized Sequential Recommendation ModelabstractSequential recommender systems, as a core technology for personalized services, have been widely applied across various domains including e-commerce and social networks. Transformer-based sequential recommendation models face challenges such as high resource consumption and inefficient sequential processing, limiting their deployment and real-time recommendation performance. Inspired by the efficiency of the RWKV architecture in handling long sequences in natural language processing, this article proposes RWKV4Rec (RWKV for Sequential Recommendation), the first application of RWKV to sequential recommendation tasks. Our RWKV4Rec model leverages RWKV’s efficient long-sequence processing capability and linear computational complexity, reducing resource demands and eliminating additional sequential processing operations. We design an item-RWKV block module and propose a Low-Rank Time Mix based on LoRA technology, which adaptively assigns weights to historical items at each timestep to generate expressive behavior sequence features. Combined with explicit user embeddings and personalized local item matrices, it enhances sequence modeling for efficient and accurate interest prediction. Experiments on four benchmark datasets demonstrate that RWKV4Rec outperforms state-of-the-art methods, achieving 1.80–3.76% improvements in NDCG@10. Mengwei Yuan, Linkai Wan, Zengmin Xu, Ziyuan Xu, Weijian Ruan |
ACM Trans. Knowl. Discov. Data | 5 |
| 2025 | NightReID: A Large-Scale Nighttime Person Re-Identification BenchmarkabstractPerson re-identification (Re-ID) is crucial for intelligent surveillance systems, facilitating the identification of individuals across multiple camera views. While significant advancements have been made for daytime scenarios, ensuring reliable Re-ID performance during nighttime remains a significant challenge. Given the cost and limited accessibility of infrared cameras, we investigate a critical question: Can RGB cameras be effectively utilized for accurate Re-ID during nighttime? To address this, we introduce NightReID, a large-scale RGB Re-ID dataset collected from a real-world nighttime surveillance system. NightReID includes 1,500 identities and over 53,000 images, capturing diverse scenes with complex lighting and adverse weather conditions. This rich dataset provides a valuable benchmark for advancing nighttime Re-ID research. Moreover, we propose the Enhancement, Denoising, and Alignment (EDA) framework with two novel modules to enhance nighttime Re-ID performance. First, an unsupervised Image Enhancement and Denoising (IED) method is designed to improve the quality of nighttime images, preserving critical details while removing noise without requiring paired ground truth. Second, we introduce Data Distribution Alignment (DDA) through statistical priors, aligning the distributions between pre-training data and nighttime data to mitigate domain shift. Extensive experiments on multiple nighttime Re-ID datasets demonstrate the significance of NightReID and validate the efficacy, flexibility, and applicability of the EDA framework. Weijian Ruan, He Li 0054, Mang Ye |
AAAI | 2 |
| 2025 | ProjAttacker: A Configurable Physical Adversarial Attack for Face Recognition via ProjectorabstractPrevious physical adversarial attacks have shown that carefully crafted perturbations can deceive face recognition systems, revealing critical security vulnerabilities. However, these attacks often struggle to impersonate multiple targets and frequently fail to bypass liveness detection. For example, attacks using human-skin masks [28] are challenging to fabricate, inconvenient to swap between users, and often fail liveness detection due to facial occlusions. A projector, however, can generate content-rich light without obstructing the face, making it ideal for non-intrusive attacks. Thus, we propose a novel physical adversarial attack using a projector and explore the superposition of projected and natural light to create adversarial facial images. This approach eliminates the need for physical artifacts on the face, effectively overcoming these limitations. Specifically, our proposed ProjAttacker generates adversarial 3D textures that are projected onto human faces. To ensure physical realizability, we introduce a light reflection function that models complex optical interactions between projected light and human skin, accounting for reflection and diffraction effects. Furthermore, we incorporate camera Image Signal Processing (ISP) simulation to maintain the robustness of adversarial perturbations across real-world diverse imaging conditions. Comprehensive evaluations conducted in both digital and physical scenarios validate the effectiveness of our method. Yuanwei Liu, Hui Wei 0004, Ruqi Xiao, Weijian Ruan, Xingxing Wei 0001, Joey Tianyi Zhou, Zheng Wang 0007 |
CVPR | 5 |
| 2025 | Precise occlusion-aware and feature-level reconstruction for occluded person re-identification
Xiujun Shu, Hanjun Li 0002, Ruizhi Qiao, Weijian Ruan, Hanjing Su, Bo Wang 0162, Shouzhi Chen |
Neurocomputing | 6 |
| 2024 | FSNA: Few-Shot Object Detection via Neighborhood Information Adaption and All AttentionabstractFew-shot object detection (FSOD), a formidable task centered around developing inclusive models with annotated constrained samples, has attracted increasing interest in recent years. This discipline addresses unbalanced data distributions, which are particularly relevant to authentic scenarios. Although recent FSOD efforts have achieved considerable success in terms of localization, recognition remains a formidable obstacle. This stems from the fact that typical FSOD models evolve from general object detection frameworks predicated on extensive training data, and they underutilize and mine data information in scenarios with restricted samples, resulting in subpar performance. To address this deficiency, we introduce a groundbreaking methodology that is specifically tailored to overcome the inadequate sample challenge in FSOD tasks. Our approach incorporates a neighborhood information adaption (NIA) module that is designed to dynamically utilize information near the target, assisting in robustly performing object identification within the target domain. In addition, we propose an innovative attention mechanism called all attention, which not only encapsulates the dependencies of each position within a single feature map but also leverages correlations with other feature maps. This methodology culminates in more refined feature representations, which are particularly advantageous in situations with limited data. Comprehensive experiments conducted on the PASCAL VOC and COCO datasets illustrate that our technique achieves a substantial improvement with regard to addressing the FSOD task. Jinxiang Zhu, Qi Wang 0079, Weijian Ruan, Liang Lei, Gefei Hao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Label-Aware Calibration and Relation-Preserving in Visual Intention UnderstandingabstractVisual intention understanding is a challenging task that explores the hidden intention behind the images of publishers in social media. Visual intention represents implicit semantics, whose ambiguous definition inevitably leads to label shifting and label blemish. The former indicates that the same image delivers intention discrepancies under different data augmentations, while the latter represents that the label of intention data is susceptible to errors or omissions during the annotation process. This paper proposes a novel method, called Label-aware Calibration and Relation-preserving (LabCR) to alleviate the above two problems from both intra-sample and inter-sample views. First, we disentangle the multiple intentions into a single intention for explicit distribution calibration in terms of the overall and the individual. Calibrating the class probability distributions in augmented instance pairs provides consistent inferred intention to address label shifting. Second, we utilize the intention similarity to establish correlations among samples, which offers additional supervision signals to form correlation alignments in instance pairs. This strategy alleviates the effect of label blemish. Extensive experiments have validated the superiority of the proposed method LabCR in visual intention understanding and pedestrian attribute recognition. Code is available at https://github.com/ShiQingHongYa/LabCR. Qinghongya Shi, Mang Ye, Wenke Huang 0003, Weijian Ruan, Bo Du 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | CDKM: Common and Distinct Knowledge Mining Network With Content Interaction for Dense CaptioningabstractThe dense captioning task aims at detecting multiple salient regions of an image and describing them separately in natural language. Although significant advancements in the field of dense captioning have been made, there are still some limitations to existing methods in recent years. On the one hand, most dense captioning methods lack strong target detection capabilities and struggle to cover all relevant content when dealing with target-intensive images. On the other hand, current transformer-based methods are powerful but neglect the acquisition and utilization of contextual information, hindering the visual understanding of local areas. To address these issues, we propose a common and distinct knowledge-mining network with content interaction for the task of dense captioning. Our network has a knowledge mining mechanism that improves the detection of salient targets by capturing common and distinct knowledge from multi-scale features. We further propose a content interaction module that combines region features into a unique context based on their correlation. Our experiments on various benchmarks have shown that the proposed method outperforms the current state-of-the-art methods. Hongyu Deng, Yushan Xie, Qi Wang 0079, Weijian Ruan, Wu Liu 0005, Yong-Jin Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | ALFPN: Adaptive Learning Feature Pyramid Network for Small Object DetectionabstractObject detection has become a crucial technology in intelligent vision systems, enabling automatic detection of target objects. While most detectors perform well on open datasets, they often struggle with small‐scale objects. This is due to the traditional top‐down feature fusion methods that weaken the semantic and location information of small objects, leading to poor classification performance. To address this issue, we propose a novel feature pyramid network, the adaptive learnable feature pyramid network (ALFPN). Our approach features an adaptive feature inspection that incorporates learnable fusion coefficients in the fusion of different levels of feature layers, aiding the network in learning features with less noise. In addition, we construct a context‐aligned supervisor that adjusts the feature maps fused at different levels to avoid scaling‐related offset effects. Our experiments demonstrate that our method achieves state‐of‐the‐art results and is highly robust for the small object detection on the TT‐100K, PASCAL VOC, and COCO datasets. These findings indicate that a model’s ability to extract discriminant features is positively correlated with its performance in detecting small objects. Qi Wang 0079, Weijian Ruan, Jingxiang Zhu, Liang Lei, Gefei Hao |
Int. J. Intell. Syst. | 3 |
| 2022 | Fine-grained code-comment semantic interaction analysisabstractCode comment, i.e., the natural language text to describe code, is considered as a killer for program comprehension. Current literature approaches mainly focus on comment generation or comment update, and thus fall short on explaining which part of the code leads to a specific content in the comment. In this paper, we propose that addressing such a challenge can better facilitate code understanding. We propose Fosterer, which can build fine-grained semantic interactions between code statements and comment tokens. It not only leverages the advanced deep learning techniques like cross-modal learning and contrastive learning, but also borrows the weapon of pre-trained vision models. Specifically, it mimics the comprehension practice of developers, treating code statements as image patches and comments as texts, and uses contrastive learning to match the semantically-related part between the visual and textual information. Experiments on a large-scale manually-labelled dataset show that our approach can achieve an F1-score around 80%, and such a performance exceeds a heuristic-based baseline to a large extent. We also find that Fosterer can work with a high efficiency, i.e., it only needs 1.5 seconds for inferring the results for a code-comment pair. Furthermore, a user study demonstrates its usability: for 65% cases, its prediction results are considered as useful for improving code understanding. Therefore, our research sheds light on a promising direction for program comprehension. Mingyang Geng, Shangwen Wang, Dezun Dong, Shanzhi Gu, Weijian Ruan, Xiangke Liao |
ICPC | 6 |
| 2022 | Online multiple object tracking based on fusing global and partial features
Jun Chen 0001, Mithun Mukherjee 0001, Chao Liang 0001, Weijian Ruan |
Neurocomputing | 5 |
| 2022 | Temporal Weighting Appearance-Aligned Network for Nighttime Video RetrievalabstractVideo-based person re-identification (ReID) aims at re-identifying video sequences of a specified person from videos captured by disjoint cameras. Existing datasets and works on this task all focus on daytime scenarios and cannot adapt well to the nighttime scenarios, which is also of significant importance for practical applications. In this paper, we contribute a new dataset for nighttime video-based ReID, termed NIVIR, which contains 800 identities with over 228,000 images. NIVIR contains video shots under various lighting conditions, different weathers, and complex scenarios, which is consistent with the real nighttime outdoor surveillance. Furthermore, we propose a temporal weighting appearance-aligned network (TWAN) for nighttime video-based ReID, which is composed of a correlation-based appearance-aligned module (CAM) and a temporal weighting module (TWM). Specifically, CAM is proposed to reconstruct the adjacent feature maps to guarantee the appearance alignment between the central frame and its adjacent frames. TWM is designed to evaluate the frame quality of a tracklet and generate temporal weights to enhance the video representation. Extensive experiments conducted on our new NIVIR dataset demonstrate that the proposed TWAN outperforms the state-of-the-art methods. We believe that our NIVIR dataset and the comprehensive attempts for solving the nighttime ReID problem will push forward the development of the ReID research community. Weijian Ruan, Yiran Tao, Linjun Ruan, Xiujun Shu, Yu Qiao 0001 |
IEEE Signal Process. Lett. | 1 |
| 2022 | Real-Time Video Deraining via Global Motion Compensation and Hybrid Multi-Scale Temporal CorrelationsabstractThe current video deraining algorithms mainly use adjacent frames to optimize the target frame information. However, they only consider the inter-frame temporal correlations of a uniform scale between frames, ignoring the inter-frame temporal correlations of different scales. In addition, the high computational cost is another drawback of the current video deraining algorithms. To this end, we propose a novel aggregation network that explores the inter-frame multi-scale temporal correlations for video deraining with the small computational cost. First, we construct a hybrid multi-scale feature extraction structure in the network to increase the receptive field of multi-scale features. For similar rain streaks at adjacent frames with different scales, a hybrid multi-scale residual block (HMSRB) is proposed to explore the complementary and redundant information at the temporal dimension to characterize the target frame. At the same time, we introduce an improved global context module (GCM) to avoid the complex motion estimation and motion compensation (ME&MC) operation as in previous video deraining approaches, while reducing the calculation complexity. Finally, a fusion block is utilized to adaptively merge the extracted features. Experiments demonstrate that our proposed network is proved to be more efficient and effective than the existing algorithms. Jun Chen 0001, Zhen Han 0002, Qikui Zhu, Weijian Ruan |
IEEE Signal Process. Lett. | 5 |
| 2022 | TICNet: A Target-Insight Correlation Network for Object TrackingabstractRecently, the correlation filter (CF) and Siamese network have become the two most popular frameworks in object tracking. Existing CF trackers, however, are limited by feature learning and context usage, making them sensitive to boundary effects. In contrast, Siamese trackers can easily suffer from the interference of semantic distractors. To address the above problems, we propose an end-to-end target-insight correlation network (TICNet) for object tracking, which aims at breaking the above limitations on top of a unified network. TICNet is an asymmetric dual-branch network involving a target-background awareness model (TBAM), a spatial-channel attention network (SCAN), and a distractor-aware filter (DAF) for end-to-end learning. Specifically, TBAM aims to distinguish a target from the background in the pixel level, yielding a target likelihood map based on color statistics to mine distractors for DAF learning. SCAN consists of a basic convolutional network, a channel-attention network, and a spatial-attention network, aiming to generate attentive weights to enhance the representation learning of the tracker. Especially, we formulate a differentiable DAF and employ it as a learnable layer in the network, thus helping suppress distracting regions in the background. During testing, DAF, together with TBAM, yields a response map for the final target estimation. Extensive experiments on seven benchmarks demonstrate that TICNet outperforms the state-of-the-art methods while running at real-time speed. Weijian Ruan, Mang Ye, Yi Wu 0001, Wu Liu 0005, Jun Chen 0001, Chao Liang 0001, Ge Li 0002, Chia-Wen Lin |
IEEE Trans. Cybern. | 1 |
| 2021 | Long-short Term Prediction for Occluded Multiple Object TrackingabstractOnline multiple object tracking (MOT) is a challenging problem in complex scenes due to frequent occlusions. Most of the existing MOT methods tend to focus on addressing an individual type of occlusion, which cannot meet the requirements of real complex scenes. In this paper, we propose a unified MOT framework that combines long- and short-term prediction models for online multiple object tracking. Basically, The short-term prediction model consists of an appearance-based model and a motion-based model, aiming at exploiting the appearance and motion of objects to handle different types of occlusions jointly. Furthermore, we adopt a cubic spline interpolation as a long-term prediction model to estimate the trajectory of the target in occluded frames. To handle different lengths of occlusions, an adaptive weighted fusion model is proposed to combine the short-term prediction model, and the long-term prediction model. Experimental results on several challenging datasets demonstrate that the proposed method outperforms state-of-the-art methods. Jun Chen 0001, Mithun Mukherjee 0001, Weijian Ruan, Chao Liang 0001, Yi Yu 0001 |
GLOBECOM | 4 |
| 2021 | Channel Augmented Joint Learning for Visible-Infrared RecognitionabstractThis paper introduces a powerful channel augmented joint learning strategy for the visible-infrared recognition problem. For data augmentation, most existing methods directly adopt the standard operations designed for single-modality visible images, and thus do not fully consider the imagery properties in visible to infrared matching. Our basic idea is to homogenously generate color-irrelevant images by randomly exchanging the color channels. It can be seamlessly integrated into existing augmentation operations without modifying the network, consistently improving the robustness against color variations. Incorporated with a random erasing strategy, it further greatly enriches the diversity by simulating random occlusions. For cross-modality metric learning, we design an enhanced channel-mixed learning strategy to simultaneously handle the intra-and cross-modality variations with squared difference for stronger discriminability. Besides, a channel-augmented joint learning strategy is further developed to explicitly optimize the outputs of augmented images. Extensive experiments with insightful analysis on two visible-infrared recognition tasks show that the proposed strategies consistently improve the accuracy. Without auxiliary information, it improves the state-of-the-art Rank-1/mAP by 14.59%/13.00% on the large-scale SYSU-MM01 dataset. Mang Ye, Weijian Ruan, Bo Du 0001, Zheng Shou 0001 |
ICCV | 2 |
| 2021 | Semantic-Guided Pixel Sampling for Cloth-Changing Person Re-IdentificationabstractCloth-changing person re-identification (re-ID) is a new rising research topic that aims at retrieving pedestrians whose clothes are changed. This task is quite challenging and has not been fully studied to date. Current works mainly focus on body shape or contour sketch, but they are not robust enough due to view and posture variations. The key to this task is to exploit cloth-irrelevant cues. This paper proposes a semantic-guided pixel sampling approach for the cloth-changing person re-ID task. We do not explicitly define which feature to extract but force the model to automatically learn cloth-irrelevant cues. Specifically, we firstly recognize the pedestrian's upper clothes and pants, then randomly change them by sampling pixels from other pedestrians. The changed samples retain the identity labels but exchange the pixels of clothes or pants among different pedestrians. Besides, we adopt a loss function to constrain the learned features to keep consistent before and after changes. In this way, the model is forced to learn cues that are irrelevant to upper clothes and pants. We conduct extensive experiments on the latest released PRCC dataset. Our method achieved 65.8% on Rank1 accuracy, which outperforms previous methods with a large margin. The code is available athttps://github.com/shuxjweb/pixel_sampling.git. Xiujun Shu, Ge Li 0002, Xiao Wang 0014, Weijian Ruan, Qi Tian 0001 |
IEEE Signal Process. Lett. | 4 |
| 2021 | A Survey of Multiple Pedestrian Tracking Based on Tracking-by-Detection FrameworkabstractMultiple pedestrian tracking (MPT) has gained significant attention due to its huge potential in a commercial application. It aims to predict multiple pedestrian trajectories and maintain their identities, given a video sequence. In the past decade, due to the advancement in pedestrian detection algorithms, Tracking-by-Detection (TBD) based algorithms have achieved tremendous successes. TBD has become the most popular MPT framework, and it has been actively studied in the past decade. In this paper, we give a comprehensive survey of recent advances in TBD-based MPT algorithms. We systematically analyze the existing TBD-based algorithms and organize the survey into four major parts. At first, this survey draws a timeline to introduce the milestones of TBD-based works which briefly reviews the development of the existing TBD-based methods. Second, the main procedures of the TBD framework are summarized, and each stage in the procedure is described in detail. Afterward, this survey analyzes the performance of existing TBD-based algorithms on MOT challenge datasets and discusses the factors that affect tracking performance. Finally, open issues and future directions in the TBD framework are discussed. Jun Chen 0001, Chao Liang 0001, Weijian Ruan, Mithun Mukherjee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Correlation Discrepancy Insight Network for Video Re-identificationabstractVideo-based person re-identification (ReID) aims at re-identifying a specified person sequence from videos that were captured by disjoint cameras. Most existing works on this task ignore the quality discrepancy across frames by using all video frames to develop a ReID method. Additionally, they adopt only the person self-characteristic as the representation, which cannot adapt to cross-camera variation effectively. To that end, we propose a novel correlation discrepancy insight network for video-based person ReID, which consists of an unsupervised correlation insight model (CIM) for video purification and a discrepancy description network (DDN) for person representation. Concretely, CIM is constructed by using kernelized correlation filters to encode person half-parts, which evaluates the frame quality by the cross correlation across frames for selecting discriminative video fragments. Furthermore, DDN exploits the selected video fragments to generate a discrepancy descriptor using a compression network, which aims at employing the discrepancies with other persons’ to facilitate the representation of the target person rather than only using the self-characteristic. Due to the advantage in handling cross-domain variation, the discrepancy descriptor is expected to provide a new pattern for the object representation in cross-camera tasks. Experimental results on three public benchmarks demonstrate that the proposed method outperforms several state-of-the-art methods. Weijian Ruan, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Wu Liu 0005, Jun Chen 0001, Jiayi Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Crowdsourcing-Based Ranking Aggregation for Person Re-IdentificationabstractPerson re-identification (re-ID) is widely applied in surveillance and criminal detection applications. The existing research focus on devising the stand-alone re-ID methods, ignoring their practical application in the multi-person collaboration scenario. To improve the search efficiency, a group of investigators are usually assigned the same task to re-identify a suspect from a shared gallery set. Due to their personalized viewpoints and search feedback operations, different investigators may obtain diverse search results of the same query target. In this case, merging different rankings and generating an improved result is of great importance. To this end, this paper proposes a crowdsourcing-based ranking aggregation to adaptively fuse multiple ranking lists for re-ID problem. The method estimates the reliability of individual investigators, with a specifically designed long tail distribution to fit the top ranking demand, and is feasible for human-machine interaction. Extensive experiments conducted on four dataset-s demonstrate the superiority of the proposed method. Yinxue Yu, Chao Liang 0001, Weijian Ruan, Longxiang Jiang |
ICASSP | 3 |
| 2020 | Dual-Direction Perception and Collaboration Network for Near-Online Multi-Object TrackingabstractBackward tracks have been exploited to improve performance of multi-object tracking (MOT). The existing method brings a stable similarity measurement but neglects unreliable detection. Exploiting predictions of forward tracks has emerged as a popular approach to tackle the task of tracking-by-detection. However, it's observed that missing detection has not been solved well enough which would significantly influence the tracking accuracy. Thus, obtaining more proposals from dual-direction tracking and predictions of tracks is concerned to address the problem of missing detection. In this paper, we propose a dual-direction perception and collaboration network (DPCNet) for MOT that exploits forward and backward tracking to collaboratively track objects. It collects candidates from the dual directions so that they can complement each other in different scenarios. Moreover, we propose a near-online tracking model based on DPCNet to improve the efficiency, which batches the tracking and makes forward and backward tracking in parallel. Experiments conducted on MOT challenge benchmarks demonstrate that the proposed method outperforms the state-of-the-arts. Xian Zhong, Weijian Ruan, Wenxin Huang, Jingling Yuan |
ICIP | 3 |
| 2020 | SIST: Online Scale-Adaptive Object tracking with Stepwise Insight
Weijian Ruan, Chao Liang 0001, Yi Yu 0001, Jun Chen 0001, Ruimin Hu |
Neurocomputing | 1 |
| 2020 | Single image de-raining via clique recursive feedback mechanism
Jun Chen 0001, Kui Jiang, Zhen Han 0002, Weijian Ruan, Zhongyuan Wang 0001, Chao Liang 0001 |
Neurocomputing | 5 |
| 2019 | POINet: Pose-Guided Ovonic Insight Network for Multi-Person Pose TrackingabstractMulti-person pose tracking aims to jointly estimate and track multi-person keypoints in the unconstrained videos. The most popular solution to this task follows the tracking-by-detection strategy that relies on human detection and data association. While human detection has been boosted by deep learning, existing works mainly exploit several separated stages with hand-crafted metrics to realize data association, leading to great uncertainty and feeble adaption in complex scenes. To handle these problems, we propose an end-to-end pose-guided ovonic insight network (POINet) for the data association in multi-person pose tracking, which jointly learns feature extraction, similarity estimation, and identity assignment. Specifically, we design a pose-guided representation network to integrate pose information into hierarchical convolutional features, generating a pose-aligned person representation for person, which helps handle partial occlusions. Moreover, we propose an ovonic insight network to adaptively encode the cross-frame identity transformation, which can cope with the tough tracking cases of person leaving and entering the scene. In general, the proposed POINet provides a new insight to realize multi-person pose tracking in an end-to-end fashion. Extensive experiments conducted on the PoseTrack benchmark demonstrate that our POINet outperforms the state-of-the-art methods. Weijian Ruan, Wu Liu 0005, Qian Bao, Jun Chen 0001, Yuhao Cheng, Tao Mei 0001 |
ACM Multimedia | 1 |
| 2019 | Multi-Correlation Filters With Triangle-Structure Constraints for Object TrackingabstractCorrelation filters (CFs) have been extensively used in tracking tasks due to their high efficiency although most of them regard the tracked target as a whole and are minimally effective in handling partial occlusion. In this study, we incorporate a part-based strategy into the framework of CFs and propose a novel multipart correlation tracker with triangle-structure constraints. Specifically, we train multiple CFs for the global object and local parts, which are then jointly applied to obtain the correlation response of any candidate during tracking. The tracker is robust in handling partial occlusion because of the use of part-based representation. The remaining global representation can contribute reliable cues in cases wherein several local filters drift away in a specific scene. We further propose a triangle-structure model to measure the structural similarity of candidates. The model employs multiple triangles to determine the spatial relationship among parts and helps constrain the location of the target. Moreover, we introduce an effective part selection scheme based on energy and integrity, which is generally applicable to part-tracking models. Extensive experiments on two public benchmarks demonstrate the superiority of the proposed method over the state-of-the-art approaches. Weijian Ruan, Jun Chen 0001, Yi Wu 0001, Jinqiao Wang, Chao Liang 0001, Ruimin Hu, Junjun Jiang |
IEEE Trans. Multim. | 1 |
| 2018 | Video-Based Person Re-Identification via Self Paced WeightingabstractPerson re-identification (re-id) is a fundamental technique to associate various person images, captured by differentsurveillance cameras, to the same person. Compared to the single image based person re-id methods, video-based personre-id has attracted widespread attentions because extra space-time information and more appearance cues that can beused to greatly improve the matching performance. However, most existing video-based person re-id methods equally treatall video frames, ignoring their quality discrepancy caused by object occlusion and motions, which is a common phenomenonin real surveillance scenario. Based on this finding, we propose a novel video-based person re-id method via self paced weighting (SPW). Firstly, we propose a self paced outlier detection method to evaluate the noise degree of video sub sequences. Thereafter, a weighted multi-pair distance metric learning approach is adopted to measure the distance of two person image sequences. Experimental results on two public datasets demonstrate the superiority of the proposed method over current state-of-the-art work. Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Weijian Ruan, Ruimin Hu |
AAAI | 5 |
| 2018 | Faster Seam Carving for Video RetargetingabstractVideo retargeting is to resize a video to a desired resolution or aspect ratio while preserving its salient content without visual distortion. The key to video retargeting is to reconcile spatio-temporal coherence of video frames, and most existing works use seam carving to achieve that by employing the dynamic programming to find optimal seams. However, these methods are too time-consuming due to high computational complexity of the dynamic programming. To this end, we propose a novel method which uses discontinuous and suboptimal seams for seam carving. Concretely, we obtain the discontinuous seams by allowing seams to move freely in homogeneous regions of the frame, which helps preserve the spatio-temporal coherence effectively. Then, the genetic algorithm is employed to find suboptimal seams, so as to reduce computational complexity. Finally, each frame can be retargeted to a new aspect ratio or size by repeatedly carving out seams. Compared to the-state-of-the-art methods, the proposed algorithm achieves comparable results at an average expense of only one third of their running time. Ruimin Hu, Chao Liang 0001, Chunxia Xiao, Weijian Ruan |
ICIP | 5 |
| 2017 | Object tracking via online trajectory optimization with multi-feature fusionabstractThe goal of object tracking is to estimate the state and trajectory of an interested target in a video sequence, thus both spatial and temporal information are of critical importance for tracking. However, most existing trackers usually determine targets just by the judgement like confidence from a single frame, which tend to treat tracking as a static detecting problem while neglecting the spatial-temporal relationship. In this paper, we propose a novel tracking method of online trajectory optimization with multi-feature fusion (TOFF). Considering the trajectory continuity, the accurate targets are determined by estimating the optimal short trajectories over all the video fragments, which can be obtained by the observation models that are iteratively updated based on the selected most reliable proposals. By employing the structured samples instead of binary-labeled samples, we construct a structured output model with representing descriptor of multi-feature fusion as the basic tracker. Extensive experiments on various challenging image sequences demonstrate the superiority of our method to several state-of-the-art methods. Weijian Ruan, Jun Chen 0001, Chao Liang 0001, Yi Wu 0001, Ruimin Hu |
ICME | 1 |
| 2017 | Structural superpixel descriptor for visual trackingabstractObject representation is a major component in object tracking, however, most conventional patch-based methods just simply decompose the object into patches with grid or stochastic rectangles. This kind of decomposition ignores the intrinsic structure of object, leading to low discriminative power and weak representation effectiveness when similar objects appear or under background clutters. In this paper, we propose an effective object descriptor based on a hierarchical representation with superpixels for visual tracking, called Structural Superpixel Descriptor (SSD). The proposed SSD not only exploits the superpixels to capture the structural information of object, but also preserves the spatial layout structure among the superpixels inside each target candidate. Moreover, we propose an adaptive patch weighting method based on spatial constraint to alleviate various adverse impacts of background information, making the tracker more robust against background noises. We show that the proposed SSD makes full use of the intrinsic structure inside target candidates. Extensive experiments conducted on various challenging sequences demonstrate that the proposed tracker performs well against state-of-the-art algorithms. Ruimin Hu, Chao Liang 0001, Weijian Ruan, Bo Luo |
IJCNN | 4 |
| 2016 | Boosted local classifiers for visual trackingabstractMost existing discriminative tracking methods model a target object as a whole and train a tracker based on holistic templates, which cannot effectively deal with partial occlusions. Instead, in this paper, by treating the target as a collection of local patches, we propose a novel tracking approach based on boosted local classifiers. Initially, a set of local patches are sampled to train a set of local classifiers, and the weight of each classifier is given based on the estimated error. In addition, the positive examples and negative examples are sampled for model update with two constraints during the tracking process, which helps obtain more negatives for updating the appearance model and improve the updating efficiency. With updating the weights of local classifiers based on the temporal stability, the tracker can effectively handle partial occlusions. Extensive experiments on various challenging image sequences demonstrate the superiority to several state-of-the-art methods. Weijian Ruan, Jun Chen 0001, Jinqiao Wang, Bo Luo, Ruimin Hu |
ICME | 1 |