Longyin Wen

dblp:119/1468 · DBLP profile ↗
← Back
60ranked-venue papers
11as first author
20since 2021 · last 2026
0000-0001-5525-492XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 40 · 8 first-author · 14 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Structured Context Learning for Generic Event Boundary Detection
abstract
Generic Event Boundary Detection (GEBD) aims to identify moments in videos that humans perceive as event boundaries. This paper proposes a novel method for addressing this task, called Structured Context Learning, which introduces the Structured Partition of Sequence (SPoS) to provide a structured context for learning temporal information. Our approach is end-to-end trainable and flexible, not restricted to specific temporal models like GRU, LSTM, and Transformers. This flexibility enables our method to achieve a better speed-accuracy trade-off. Specifically, we apply SPoS to partition the input frame sequence and provide a structured context for the subsequent temporal model. Notably, SPoS’s overall computational complexity is linear with respect to the video length. We next calculate group similarities to capture differences between frames, and a lightweight fully convolutional network is utilized to determine the event boundaries based on the grouped similarity maps. To remedy the ambiguities of boundary annotations, we adapt the Gaussian kernel to preprocess the ground-truth event boundaries. Our proposed method has been extensively evaluated on the challenging Kinetics-GEBD, TAPOS, and shot transition detection datasets, demonstrating its superiority over existing state-of-the-art methods.
Xin Gu 0003, Dexiang Hong, Libo Zhang 0001, Tiejian Luo, Longyin Wen, Heng Fan 0001
WACV7
2025 $\mathcal{D}$-Attn: Decomposed Attention for Large Vision-and-Language Models
Chia-Wen Kuo, Sijie Zhu, Xiaohui Shen, Longyin Wen
ICCV5
2025 SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing
abstract
Due to the challenges of manually collecting accurate editing data, existing datasets are typically constructed using various automated methods, leading to noisy supervision signals caused by the mismatch between editing instructions and original-edited image pairs. Recent efforts attempt to improve editing models through generating higher-quality edited images, pre-training on recognition tasks, or introducing vision-language models (VLMs) but fail to resolve this fundamental issue. In this paper, we offer a novel solution by constructing more effective editing instructions for given image pairs. This includes rectifying the editing instructions to better align with the original-edited image pairs and using contrastive editing instructions to further enhance their effectiveness. Specifically, we find that editing models exhibit specific generation attributes at different inference steps, independent of the text. Based on these prior attributes, we define a unified guide for VLMs to rectify editing instructions. However, there are some challenging editing scenarios that cannot be resolved solely with rectified instructions. To this end, we further construct contrastive supervision signals with positive and negative instructions and introduce them into the model training using triplet loss, thereby further facilitating supervision effectiveness. Our method does not require the VLM modules or pre-training tasks used in previous work, offering a more direct and efficient way to provide better supervision signals, and providing a novel, simple, and effective solution for instruction-based image editing. Results on multiple benchmarks demonstrate that our method significantly outperforms existing approaches. Compared with previous SOTA SmartEdit, we achieve 9.19% improvements on the Real-Edit benchmark with 30x less training data and 13x smaller model size.
Ming Li 0010, Xiaoying Xing, Longyin Wen, Chen Chen 0001, Sijie Zhu
ICCV5
2025 Multi-Reward as Condition for Instruction-based Image Editing
abstract
High-quality training triplets (instruction, original image, edited image) are essential for instruction-based image editing. Predominant training datasets (e.g., InsPix2Pix) are created using text-to-image generative models (e.g., Stable Diffusion, DALL-E) which are not trained for image editing. Accordingly, these datasets suffer from inaccurate instruction following, poor detail preserving, and generation artifacts. In this paper, we propose to address the training data quality issue with multi-perspective reward data instead of refining the ground-truth image quality. 1) we first design a quantitative metric system based on best-in-class LVLM (Large Vision Language Model), i.e., GPT-4o in our case, to evaluate the generation quality from 3 perspectives, namely, instruction following, detail preserving, and generation quality. For each perspective, we collected quantitative score in $0\sim 5$ and text descriptive feedback on the specific failure points in ground-truth edited images, resulting in a high-quality editing reward dataset, i.e., RewardEdit20K. 2) We further proposed a novel training framework to seamlessly integrate the metric output, regarded as multi-reward, into editing models to learn from the imperfect training triplets. During training, the reward scores and text descriptions are encoded as embeddings and fed into both the latent space and the U-Net of the editing models as auxiliary conditions. During inference, we set these additional conditions to the highest score with no text description for failure points, to aim at the best generation outcome. 3) We also build a challenging evaluation benchmark with real-world images/photos and diverse editing instructions, named as Real-Edit. Experiments indicate that our multi-reward conditioned model outperforms its no-reward counterpart on two popular editing pipelines, i.e., InsPix2Pix and SmartEdit. Code is released at https://github.com/bytedance/Multi-Reward-Editing.
Xin Gu 0003, Libo Zhang 0001, Longyin Wen, Tiejian Luo, Sijie Zhu
ICLR5
2024 CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
abstract
Recent advancements in Multimodal Large Language Models (LLMs) have focused primarily on scaling by increasing text-image pair data and enhancing LLMs to improve performance on multimodal tasks. However, these scaling approaches are computationally expensive and overlook the significance of efficiently improving model capabilities from the vision side. Inspired by the successful applications of Mixture-of-Experts (MoE) in LLMs, which improves model scalability during training while keeping inference costs similar to those of smaller models, we propose CuMo, which incorporates Co-upcycled Top-K sparsely-gated Mixture-of-experts blocks into both the vision encoder and the MLP connector, thereby enhancing the multimodal LLMs with neglectable additional activated parameters during inference. CuMo first pre-trains the MLP blocks and then initializes each expert in the MoE block from the pre-trained MLP block during the visual instruction tuning stage, with auxiliary losses to ensure a balanced loading of experts. CuMo outperforms state-of-the-art multimodal LLMs across various VQA and visual-instruction-following benchmarks within each model size group, all while training exclusively on open-sourced datasets.
Jiachen Li 0003, Sijie Zhu, Chia-Wen Kuo, Jitesh Jain, Humphrey Shi, Longyin Wen
NeurIPS9
2023 Text with Knowledge Graph Augmented Transformer for Video Captioning
abstract
Video captioning aims to describe the content of videos using natural language. Although significant progress has been made, there is still much room to improve the performance for real-world applications, mainly due to the long-tail words challenge. In this paper, we propose a text with knowledge graph augmented transformer (TextKG)for video captioning. Notably, TextKG is a two-stream transformer, formed by the external stream and internal stream. The external stream is designed to absorb additional knowledge, which models the interactions between the additional knowledge, e.g., pre-built knowledge graph, and the built-in information of videos, e.g., the salient object regions, speech transcripts, and video captions, to mitigate the long-tail words challenge. Meanwhile, the internal stream is designed to exploit the multi-modality information in videos (e.g., the appearance of video frames, speech transcripts, and video captions) to ensure the quality of caption results. In addition, the cross attention mechanism is also used in between the two streams for sharing information. In this way, the two streams can help each other for more accurate results. Extensive experiments conducted on four challenging video captioning datasets, i.e., YouCookII, ActivityNet Captions, MSR-VTT, and MSVD, demonstrate that the proposed method performs favorably against the state-of-the-art methods. Specifically, the proposed TextKG method out-performs the best published results by improving 18.7% absolute CIDEr scores on the YouCookII dataset.
Xin Gu 0003, Libo Zhang 0001, Tiejian Luo, Longyin Wen
CVPR6
2023 Accurate and Fast Compressed Video Captioning
abstract
Existing video captioning approaches typically require to first sample video frames from a decoded video and then conduct a subsequent process (e.g., feature extraction and/or captioning model learning). In this pipeline, manual frame sampling may ignore key information in videos and thus degrade performance. Additionally, redundant information in the sampled frames may result in low efficiency in the inference of video captioning. Addressing this, we study video captioning from a different perspective in compressed domain, which brings multi-fold advantages over the existing pipeline: 1) Compared to raw images from the decoded video, the compressed video, consisting of I-frames, motion vectors and residuals, is highly distinguishable, which allows us to leverage the entire video for learning without manual sampling through a specialized model design; 2) The captioning model is more efficient in inference as smaller and less redundant information is processed. We propose a simple yet effective end-to-end transformer in the compressed domain for video captioning that enables learning from the compressed video for captioning. We show that even with a simple design, our method can achieve state-of-the-art performance on different benchmarks while running almost 2× faster than existing approaches. Code is available at https://github.com/acherstyx/CoCap.
Yaojie Shen, Xin Gu 0003, Heng Fan 0001, Longyin Wen, Libo Zhang 0001
ICCV5
2023 DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training
Linchao Zhu, Longyin Wen, Yi Yang 0001
ICLR3
2022 End-to-End Compressed Video Representation Learning for Generic Event Boundary Detection
abstract
Generic event boundary detection aims to localize the generic, taxonomy-free event boundaries that segment videos into chunks. Existing methods typically require video frames to be decoded before feeding into the network, which demands considerable computational power and storage space. To that end, we propose a new end-to-end compressed video representation learning for event boundary detection that leverages the rich information in the compressed domain, i.e., RGB, motion vectors, residuals, and the internal group of pictures (GOP) structure, withoutfully decoding the video. Specifically, we first use the Con-vNets to extract features of the I-frames in the Gaps. After that, a light-weight spatial-channel compressed encoder is designed to compute the feature representations of the P-frames based on the motion vectors, residuals and representations of their dependent I-frames. A temporal contrastive module is proposed to determine the event boundaries of video sequences. To remedy the ambiguities of annotations and speed up the training process, we use the Gaussian kernel to preprocess the ground-truth event boundaries. Extensive experiments conducted on the Kinetics-GEBD dataset demonstrate that the proposed method achieves comparable results to the state-of-the-art methods with 4.5 x faster running speed.
Longyin Wen, Dexiang Hong, Tiejian Luo, Libo Zhang 0001
CVPR3
2022 Detection and Tracking Meet Drones Challenge
abstract
Drones, or general UAVs, equipped with cameras have been fast deployed with a wide range of applications, including agriculture, aerial photography, and surveillance. Consequently, automatic understanding of visual data collected from drones becomes highly demanding, bringing computer vision and drones more and more closely. To promote and track the developments of object detection and tracking algorithms, we have organized three challenge workshops in conjunction with ECCV 2018, ICCV 2019 and ECCV 2020, attracting more than 100 teams around the world. We provide a large-scale drone captured dataset, VisDrone, which includes four tracks, i.e., (1) image object detection, (2) video object detection, (3) single object tracking, and (4) multi-object tracking. In this paper, we first present a thorough review of object detection and tracking datasets and benchmarks, and discuss the challenges of collecting large-scale drone-based object detection and tracking datasets with fully manual annotations. After that, we describe our VisDrone dataset, which is captured over various urban/suburban areas of 14 different cities across China from North to South. Being the largest such dataset ever published, VisDrone enables extensive evaluation and investigation of visual analysis algorithms for the drone platform. We provide a detailed analysis of the current state of the field of large-scale object detection and tracking on drones, and conclude the challenge as well as propose future directions. We expect the benchmark largely boost the research and development in video analysis on drone platforms. All the datasets and experimental results can be downloaded from https://github.com/VisDrone/VisDrone-Dataset.
Pengfei Zhu 0001, Longyin Wen, Dawei Du, Xiao Bian, Heng Fan 0001, Qinghua Hu, Haibin Ling
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 ChaLearn Looking at People: IsoGD and ConGD Large-Scale RGB-D Gesture Recognition
abstract
The ChaLearn large-scale gesture recognition challenge has run twice in two workshops in conjunction with the International Conference on Pattern Recognition (ICPR) 2016 and International Conference on Computer Vision (ICCV) 2017, attracting more than 200 teams around the world. This challenge has two tracks, focusing on isolated and continuous gesture recognition, respectively. It describes the creation of both benchmark datasets and analyzes the advances in large-scale gesture recognition based on these two datasets. In this article, we discuss the challenges of collecting large-scale ground-truth annotations of gesture recognition and provide a detailed analysis of the current methods for large-scale isolated and continuous gesture recognition. In addition to the recognition rate and mean Jaccard index (MJI) as evaluation metrics used in previous challenges, we introduce the corrected segmentation rate (CSR) metric to evaluate the performance of temporal segmentation for continuous gesture recognition. Furthermore, we propose a bidirectional long short-term memory (Bi-LSTM) method, determining video division points based on skeleton points. Experiments show that the proposed Bi-LSTM outperforms state-of-the-art methods with an absolute improvement of 8.1% (from 0.8917 to 0.9639) of CSR.
Jun Wan 0001, Chi Lin 0002, Longyin Wen, Yunan Li 0001, Qiguang Miao, Sergio Escalera, Gholamreza Anbarjafari, Isabelle Guyon, Guodong Guo, Stan Z. Li
IEEE Trans. Cybern.3
2021 Rethinking Object Detection in Retail Stores
abstract
The conventional standard for object detection uses a bounding box to represent each individual object instance. However, it is not practical in the industry-relevant applications in the context of warehouses due to severe occlusions among groups of instances of the same categories. In this paper, we propose a new task, i.e., simultaneously object localization and counting, abbreviated as Locount, which requires algorithms to localize groups of objects of interest with the number of instances. However, there does not exist a dataset or benchmark designed for such a task. To this end, we collect a large-scale object localization and counting dataset with rich annotations in retail stores, which consists of 50,394 images with more than 1.9 million object instances in 140 categories. Together with this dataset, we provide a new evaluation protocol and divide the training and testing subsets to fairly evaluate the performance of algorithms for Locount, developing a new benchmark for the Locount task. Moreover, we present a cascaded localization and counting network as a strong baseline, which gradually classifies and regresses the bounding boxes of objects with the predicted numbers of instances enclosed in the bounding boxes, trained in an end-to-end manner. Extensive experiments are conducted on the proposed dataset to demonstrate its significance and the analysis is provided to indicate future directions. Dataset is available at https://isrc.iscas.ac.cn/gitlab/research/locount-dataset.
Yuanqiang Cai, Longyin Wen, Libo Zhang 0001, Dawei Du, Weiqiang Wang 0001
AAAI2
2021 Detection, Tracking, and Counting Meets Drones in Crowds: A Benchmark
abstract
To promote the developments of object detection, tracking and counting algorithms in drone-captured videos, we construct a benchmark with a new drone-captured large-scale dataset, named as DroneCrowd, formed by 112 video clips with 33, 600 HD frames in various scenarios. Notably, we annotate 20, 800 people trajectories with 4.8 million heads and several video-level attributes. Meanwhile, we design the Space-Time Neighbor-Aware Network (STNNet) as a strong baseline to solve object detection, tracking and counting jointly in dense crowds. STNNet is formed by the feature extraction module, followed by the density map estimation heads, and localization and association subnets. To exploit the context information of neighboring objects, we design the neighboring context loss to guide the association subnet training, which enforces consistent relative position of nearby objects in temporal domain. Extensive experiments on our DroneCrowd dataset demonstrate that STNNet performs favorably against the state-of-the-arts.
Longyin Wen, Dawei Du, Pengfei Zhu 0001, Qinghua Hu, Qilong Wang 0001, Liefeng Bo, Siwei Lyu
CVPR1
2021 Towards Real-World Prohibited Item Detection: A Large-Scale X-ray Benchmark
abstract
Automatic security inspection using computer vision technology is a challenging task in real-world scenarios due to various factors, including intra-class variance, class imbalance, and occlusion. Most of the previous methods rarely solve the cases that the prohibited items are deliberately hidden in messy objects due to the lack of large-scale datasets, restricted their applications in real-world scenarios. Towards real-world prohibited item detection, we collect a large-scale dataset, named as PIDray, which covers various cases in real-world scenarios for prohibited item detection, especially for deliberately hidden items. With an intensive amount of effort, our dataset contains 12 categories of prohibited items in 47, 677 X-ray images with high-quality annotated segmentation masks and bounding boxes. To the best of our knowledge, it is the largest prohibited items detection dataset to date. Meanwhile, we design the selective dense attention network (SDANet) to construct a strong baseline, which consists of the dense attention module and the dependency refinement module. The dense attention module formed by the spatial and channel-wise dense attentions, is designed to learn the discriminative features to boost the performance. The dependency refinement module is used to exploit the dependencies of multi-scale features. Extensive experiments conducted on the collected PIDray dataset demonstrate that the proposed method performs favorably against the state-of-the-art methods, especially for detecting the deliberately hidden items.
Boying Wang, Libo Zhang 0001, Longyin Wen, Xianglong Liu 0001
ICCV3
2021 RefineDet++: Single-Shot Refinement Neural Network for Object Detection
abstract
Convolutional neural network based methods have dominated object detection in recent years, which can be divided into the one-stage approach and the two-stage approach. In general, the two-stage approach (e.g., Faster R-CNN) achieves high accuracy, while the one-stage approach (e.g., SSD) has the advantage of high efficiency. To inherit the merits of both while overcoming their disadvantages, we propose a novel single-shot based detector, namely RefineDet++, which achieves better accuracy than two-stage methods and maintains comparable efficiency of one-stage methods. The proposed RefineDet++ consists of two inter-connected modules: the anchor refinement module and the alignment detection module. Specifically, the former module aims to (1) filter out negative anchors to reduce search space for the subsequent classifier, and (2) coarsely adjust the locations and sizes of anchors to provide better initialization for the subsequent regressor. The latter module takes (1) the refined anchors as the input from the former module with (2) a newly designed alignment convolution operation to further improve the regression accuracy and predict multi-class label. Meanwhile, we design a transfer connection block to transfer the features in the anchor refinement module to predict locations, sizes and class labels of objects in the object detection module. The multi-task loss function enables us to train the whole network in an end-to-end way. Extensive experiments on PASCAL VOC and MS COCO demonstrate that RefineDet++ achieves state-of-the-art detection accuracy with high efficiency.
Longyin Wen, Zhen Lei 0001, Stan Z. Li
IEEE Trans. Circuits Syst. Video Technol.2
2021 Multi-Drone-Based Single Object Tracking With Agent Sharing Network
abstract
Drones equipped with cameras (UAVs) can dynamically track the target in the air from a broader view compared with static cameras or moving sensors over the ground. However, it is still challenging to accurately track the target using a single drone due to several factors such as appearance variations and severe occlusions. To this end, we collect a newMulti-Drone singleObjectTracking (MDOT) dataset that consists of 92 groups of video clips with 113, 918 high resolution frames taken by two drones and 63 groups of video clips with 145, 875 high resolution frames taken by three drones. Besides, two evaluation metrics are specially designed for multi-drone single object tracking,i.e., automatic fusion score (AFS) and ideal fusion score (IFS). Moreover, the agent sharing network (ASNet) is proposed by integrating self-supervised template sharing, target re-detection, and view-aware fusion of the target from multiple drones into a unified framework, which can improve the tracking accuracy significantly compared with single drone tracking. Extensive experiments on MDOT show that our ASNet significantly outperforms recent state-of-the-art trackers. The dataset can be found inhttps://github.com/VisDrone/MultiDrone.
Pengfei Zhu 0001, Jiayu Zheng, Dawei Du, Longyin Wen, Yiming Sun 0003, Qinghua Hu
IEEE Trans. Circuits Syst. Video Technol.4
2021 Fast Online Video Pose Estimation by Dynamic Bayesian Modeling of Mode Transitions
abstract
We propose a fast online video pose estimation method to detect and track human upper-body poses based on a conditional dynamic Bayesian modeling of pose modes without referring to future frames. The estimation of human body poses from videos is an important task with many applications. Our method extends fast image-based pose estimation to live video streams by leveraging the temporal correlation of articulated poses between frames. Video pose estimation is inferred over a time window using a conditional dynamic Bayesian network (CDBN), which we term time-windowed CDBN. Specifically, latent pose modes and their transitions are modeled and co-determined from the combination of three modules: 1) inference based on current observations; 2) the modeling of mode-to-mode transitions as a probabilistic prior; and 3) the modeling of state-to-mode transitions using a multimode softmax regression. Given the predicted pose modes, the body poses in terms of arm joint locations can then be determined more accurately and robustly. Our method is suitable to investigate high frame rate (HFR) scenarios, where pose mode transitions can effectively capture action-related temporal information to boost performance. We evaluate our method on a newly collected HFR-Pose dataset and four major video pose datasets (VideoPose2, TUM Kitchen, FLIC, and Penn_Action). Our method achieves improvements in both accuracy and efficiency over existing online video pose estimation methods.
Ming-Ching Chang, Lipeng Ke, Honggang Qi, Longyin Wen, Siwei Lyu
IEEE Trans. Cybern.4
2021 Learning Self-Supervised Space-Time CNN for Fast Video Style Transfer
abstract
Style transfer on images has achieved significant advances in recent years, with the deep convolutional neural network (CNN). Directly applying image style transfer algorithms to each frame of a video independently often leads to flickering and unstable results. In this work, we present a self-supervised space-time convolutional neural network (CNN) based method for online video style transfer, named as VTNet, which is end-to-end trained from nearly unlimited unlabeled video data to produce temporally coherent stylized videos in real-time. Specifically, our VTNet transfer the style of a reference image to the source video frames, which is formed by the temporal prediction branch and the stylizing branch. The temporal prediction branch is used to capture discriminative spatiotemporal features for temporal consistency, pretrained in an adversarial manner from unlabeled video data. The stylizing branch is used to transfer the style image to a video frame with the guidance from the temporal prediction branch to ensure temporal consistency. To guide the training of VTNet, we introduce the style-coherence loss net (SCNet), which assembles the content loss, the style loss, and the new designed coherence loss. These losses are computed based on high-level features extracted from a pretrained VGG-16 network. The content loss is used to preserve high-level abstract contents of the input frames, and the style loss introduces new colors and patterns from the style image. Instead of using optical flow to explicitly redress the stylized video frames, we design the coherence loss to make the stylized video inherit the dynamics and motion patterns from the source video to remove temporal flickering. Extensive subjective and objective evaluations on various styles demonstrate that the proposed method achieves favorable results against the state-of-the-arts with high efficiency.
Kai Xu 0013, Longyin Wen, Guorong Li, Honggang Qi, Liefeng Bo, Qingming Huang
IEEE Trans. Image Process.2
2021 SiamCAN: Real-Time Visual Tracking Based on Siamese Center-Aware Network
abstract
In this article, we present a novel Siamese center-aware network (SiamCAN) for visual tracking, which consists of the Siamese feature extraction subnetwork, followed by the classification, regression, and localization branches in parallel. The classification branch is used to distinguish the target from background, and the regression branch is introduced to regress the bounding box of the target. To reduce the impact of manually designed anchor boxes to adapt to different target motion patterns, we design the localization branch to localize the target center directly to assist the regression branch generating accurate results. Meanwhile, we introduce the global context module into the localization branch to capture long-range dependencies for more robustness to large displacements of the target. A multi-scale learnable attention module is used to guide these three branches to exploit discriminative features for better performance. Extensive experiments on 9 challenging benchmarks, namely VOT2016, VOT2018, VOT2019, OTB100, LTB35, LaSOT, TC128, UAV123 and VisDrone-SOT2019 demonstrate that SiamCAN achieves leading accuracy with high efficiency. Our source code is available at https://isrc.iscas.ac.cn/gitlab/research/siamcan.
Wenzhang Zhou, Longyin Wen, Libo Zhang 0001, Dawei Du, Tiejian Luo
IEEE Trans. Image Process.2
2021 Self-Supervised Deep TripleNet for Video Object Segmentation
abstract
Most of previous video object segmentation methods require a large amount of pixel-level annotated video data to construct a robust model. It is quite expensive to label segmentation mask in video. In this paper, we propose a self-supervised triplenet for video object segmentation, which only leverages nearly unlimited unlabeled video data in training phase. Our method consists of two modules, i.e., the temporal motion module and appearance matching module. The temporal motion module is trained based on the pixel correspondence between two video frames in a self-supervised manner, which models the motion patterns between two video frames and propagates the labels from one frame to another. Meanwhile, the appearance matching module encodes the reference frame and its corresponding mask, and generates the segmentation mask of the same object in target frame. The appearance matching module can adjust and refine the outout of temporal motion module, and avoid error accumulation by matching the reference appearance. In order to train the appearance matching module in self-supervised manner, we propose two mask generation strategies: foreground region mask generation and random color region mask generation. Extensive experiments conducted on four challenging video object segmentation datasets, i.e., DAVIS-2017, Youtube-VOS, DAVIS- 2016 and SegTrack v2, demonstrate that the proposed method performs favorable against the state-of-the-art self-supervised methods, and performs even competitively with fully-supervised methods. We also show our self-supervised approach has actually superior generalizability to the majority of supervised methods.
Kai Xu 0013, Longyin Wen, Guorong Li, Qingming Huang
IEEE Trans. Multim.2
2020 Attention Convolutional Binary Neural Tree for Fine-Grained Visual Categorization
abstract
Fine-grained visual categorization (FGVC) is an important but challenging task due to high intra-class variances and low inter-class variances caused by deformation, occlusion, illumination, etc. An attention convolutional binary neural tree architecture is presented to address those problems for weakly supervised FGVC. Specifically, we incorporate convolutional operations along edges of the tree structure, and use the routing functions in each node to determine the root-to-leaf computational paths within the tree. The final decision is computed as the summation of the predictions from leaf nodes. The deep convolutional operations learn to capture the representations of objects, and the tree structure characterizes the coarse-to-fine hierarchical feature learning process. In addition, we use the attention transformer module to enforce the network to capture discriminative features. The negative log-likelihood loss is used to train the entire network in an end-to-end fashion by SGD with back-propagation. Several experiments on the CUB-200-2011, Stanford Cars and Aircraft datasets demonstrate that the proposed method performs favorably against the state-of-the-arts.
Ruyi Ji, Longyin Wen, Libo Zhang 0001, Dawei Du, Chen Zhao 0024, Xianglong Liu 0001, Feiyue Huang
CVPR2
2020 Learning Semantic Neural Tree for Human Parsing
Ruyi Ji, Dawei Du, Libo Zhang 0001, Longyin Wen, Chen Zhao 0024, Feiyue Huang, Siwei Lyu
ECCV (13)4
2020 Spatial Attention Pyramid Network for Unsupervised Domain Adaptation
Dawei Du, Libo Zhang 0001, Longyin Wen, Tiejian Luo, Pengfei Zhu 0001
ECCV (13)4
2020 Efficient Pig Counting in Crowds with Keypoints Tracking and Spatial-aware Temporal Response Filtering
abstract
Pig counting is a crucial task for large-scale pig farming. Pigs are usually visually counted by human. But this process is very time-consuming and error-prone. Few studies in literature developed automated pig counting method. The existing works only focused on pig counting using single image, and its level of accuracy faced challenges due to pig movements, occlusion and overlapping. Especially, the field of view of a single image is very limited, and could not meet the needs of pig counting for large pig grouping houses. Towards addressing these challenges, we presented a real-time automated pig counting system in crowds using only one monocular fisheye camera with an inspection robot. Our system showed that it achieved performance superior to human. Our pipeline began with a novel bottom-up pig detection algorithm to avoid false negatives due to overlapping, occlusion and deformable pig shapes. This detection included a deep convolution neural network (CNN) for pig body part keypoints detection and the keypoints association method to identify individual pigs. It then employed an efficient on-line tracking method to associate pigs across image frames. Finally, pig counts were estimated by a novel spatial-aware temporal response filtering (STRF) method to suppress false positives caused by pig or camera movements or tracking failures. The whole pipeline has been deployed in an edge computing device, and demonstrated the effectiveness.
Shiwen Shen, Longyin Wen, Si Luo, Liefeng Bo
ICRA3
2020 Guided Attention Network for Object Detection and Counting on Drones
abstract
Object detection and counting are related but challenging problems, especially for drone based scenes with small objects and cluttered background. In this paper, we propose a new Guided Attention network (GAnet) to deal with both object detection and counting tasks based on the feature pyramid. Different from the previous methods relying on unsupervised attention modules, we fuse different scales of feature maps by using the proposed weakly-supervised Background Attention (BA) between the background and objects for more semantic feature representation. Then, the Foreground Attention (FA) module is developed to consider both global and local appearance of the object to facilitate accurate localization. Moreover, the new data argumentation strategy is designed to train a robust model in the drone based scenes with various illumination conditions. Extensive experiments on three challenging benchmarks (i.e., UAVDT, CARPK and PUCPR+) show the state-of-the-art detection and counting performance of the proposed method compared with existing methods. Code can be found at https://isrc.iscas.ac.cn/gitlab/research/ganet.
Yuanqiang Cai, Dawei Du, Libo Zhang 0001, Longyin Wen, Weiqiang Wang 0001, Siwei Lyu
ACM Multimedia4
2020 UA-DETRAC: A new benchmark and protocol for multi-object detection and tracking
Longyin Wen, Dawei Du, Zhaowei Cai, Zhen Lei 0001, Ming-Ching Chang, Honggang Qi, Jongwoo Lim, Ming-Hsuan Yang 0001, Siwei Lyu
Comput. Vis. Image Underst.1
2019 Learning Non-Uniform Hypergraph for Multi-Object Tracking
abstract
The majority of Multi-Object Tracking (MOT) algorithms based on the tracking-by-detection scheme do not use higher order dependencies among objects or tracklets, which makes them less effective in handling complex scenarios. In this work, we present a new near-online MOT algorithm based on non-uniform hypergraph, which can model different degrees of dependencies among tracklets in a unified objective. The nodes in the hypergraph correspond to the tracklets and the hyperedges with different degrees encode various kinds of dependencies among them. Specifically, instead of setting the weights of hyperedges with different degrees empirically, they are learned automatically using the structural support vector machine algorithm (SSVM). Several experiments are carried out on various challenging datasets (i.e., PETS09, ParkingLot sequence, SubwayFace, and MOT16 benchmark), to demonstrate that our method achieves favorable performance against the state-of-the-art MOT methods.
Longyin Wen, Dawei Du, Shengkun Li, Xiao Bian, Siwei Lyu
AAAI1
2019 Graph-to-Graph Energy Minimization for Video Object Segmentation
abstract
We describe a new unsupervised video object segmentation (VOS) method based on the graph-to-graph energy minimization, which focuses on exploiting the mutual bootstrapping information between bottom-up (i.e., using pixel/superpixel attributes) and top-down (i.e., using learned appearance and motion cues) processes in a unified framework. Specifically, we construct a graph-to-graph energy function to encode the spatial similarities among superpixels (superpixel-graph) and temporal consistency among regions (region-graph). An efficient heuristic iterative algorithm is used to minimize the energy function to get the optimal assignment of superpixel and region labels to complete the VOS task. Experiments on two challenging benchmarks (i.e., SegTrack v2 and DAVIS) show that the proposed method achieves favorable performance against the state-of-the-art unsupervised VOS methods and comparable performance with the state-of-the-art semi-supervised methods.
Yuezun Li, Longyin Wen, Ming-Ching Chang, Siwei Lyu
AVSS2
2019 Spatiotemporal CNN for Video Object Segmentation
abstract
In this paper, we present a unified, end-to-end trainable spatiotemporal CNN model for VOS, which consists of two branches, i.e., the temporal coherence branch and the spatial segmentation branch. Specifically, the temporal coherence branch pretrained in an adversarial fashion from unlabeled video data, is designed to capture the dynamic appearance and motion cues of video sequences to guide object segmentation. The spatial segmentation branch focuses on segmenting objects accurately based on the learned appearance and motion cues. To obtain accurate segmentation results, we design a coarse-to-fine process to sequentially apply a designed attention module on multi-scale feature maps, and concatenate them to produce the final prediction. In this way, the spatial segmentation branch is enforced to gradually concentrate on object regions. These two branches are jointly fine-tuned on video segmentation sequences in an end-to-end manner. Several experiments are carried out on three challenging datasets (i.e., DAVIS-2016, DAVIS-2017 and Youtube-Object) to show that our method achieves favorable performance against the state-of-the-arts. Code is available at https://github.com/longyin880815/STCNN.
Kai Xu 0013, Longyin Wen, Guorong Li, Liefeng Bo, Qingming Huang
CVPR2
2019 ScratchDet: Training Single-Shot Object Detectors From Scratch
abstract
Current state-of-the-art object objectors are fine-tuned from the off-the-shelf networks pretrained on large-scale classification dataset ImageNet, which incurs some additional problems: 1) The classification and detection have different degrees of sensitivity to translation, resulting in the learning objective bias; 2) The architecture is limited by the classification network, leading to the inconvenience of modification. To cope with these problems, training detectors from scratch is a feasible solution. However, the detectors trained from scratch generally perform worse than the pretrained ones, even suffer from the convergence issue in training. In this paper, we explore to train object detectors from scratch robustly. By analysing the previous work on optimization landscape, we find that one of the overlooked points in current trained-from-scratch detector is the BatchNorm. Resorting to the stable and predictable gradient brought by BatchNorm, detectors can be trained from scratch stably while keeping the favourable performance independent to the network architecture. Taking this advantage, we are able to explore various types of networks for object detection, without suffering from the poor convergence. By extensive experiments and analyses on downsampling factor, we propose the Root-ResNet backbone network, which makes full use of the information from original images. Our ScratchDet achieves the state-of-the-art accuracy on PASCAL VOC 2007, 2012 and MS COCO among all the train-from-scratch detectors and even performs better than several one-stage pretrained methods. Codes will be made publicly available at https://github.com/KimSoybean/ScratchDet.
Rui Zhu 0014, Xiaobo Wang 0001, Longyin Wen, Hailin Shi, Liefeng Bo, Tao Mei 0001
CVPR4
2019 Data Priming Network for Automatic Check-Out
abstract
Automatic Check-Out (ACO) receives increased interests in recent years. An important component of the ACO system is the visual item counting, which recognizes the categories and counts of the items chosen by the customers. However, the training of such a system is challenged by the domain adaptation problem, in which the training data are images from isolated items while the testing images are for collections of items. Existing methods solve this problem with data augmentation using synthesized images, but the image synthesis leads to unreal images that affect the training process. In this paper, we propose a new data priming method to solve the domain adaptation problem. Specifically, we first use pre-augmentation data priming, in which we remove distracting background from the training images using the coarse-to-fine strategy and select images with realistic view angles by the pose pruning method. In the post-augmentation step, we train a data priming network using detection and counting collaborative learning, and select more reliable images from testing data to fine-tune the final visual item tallying network. Experiments on the large scale Retail Product Checkout (RPC) dataset demonstrate the superiority of the proposed method, i.e., we achieve 80.51% checkout accuracy compared with 56.68% of the baseline methods. The source codes can be found in https://isrc.iscas.ac.cn/gitlab/research/acm-mm-2019-ACO.
Dawei Du, Libo Zhang 0001, Tiejian Luo, Qi Tian 0001, Longyin Wen, Siwei Lyu
ACM Multimedia7
2019 Single-Shot Scale-Aware Network for Real-Time Face Detection
Longyin Wen, Hailin Shi, Zhen Lei 0001, Siwei Lyu, Stan Z. Li
Int. J. Comput. Vis.2
2018 Evolvement Constrained Adversarial Learning for Video Style Transfer
Wenbo Li 0001, Longyin Wen, Xiao Bian, Siwei Lyu
ACCV (1)2
2018 Pixel Offset Regression (POR) for Single-shot Instance Segmentation
abstract
State-of-the-art instance segmentation methods including Mask-RCNN and MNC are multi-shot, as multiple region of interest (ROI) forward passes are required to distinguish candidate regions. Multi-shot architectures usually achieve good performance on public benchmarks. However, hundreds of ROI forward passes in sequel limits their running efficiency, which is a critical point in several utilities such as vehicle surveillance. As such, we arrange our focus on seeking a well trade-off between performance and efficiency. In this paper, we introduce a novel Pixel Offset Regression (POR) scheme which can simply extend single-shot object detector to single-shot instance segmentation system, i.e., segmenting all instances in a single pass. Our framework is based on VGG161with following four parts: (1) a single-shot detection branch to generate object detections, (2) a segmentation branch to estimate foreground masks, (3) a pixel offset regression branch to effectively estimate the distance and orientation from each pixel to the respective object center and (4) a merging process combining output of each branch to obtain instances. Our framework is evaluated on Berkeley-BDD, KITTI and PASCAL VOC2012 validation set, with comparison against several VGG16 based multi-shot methods. Without whistles and bells, our framework exhibits decent performance, which shows good potential for fast speed required applications.
Yuezun Li, Xiao Bian, Ming-Ching Chang, Longyin Wen, Siwei Lyu
AVSS4
2018 Single-Shot Refinement Neural Network for Object Detection
abstract
For object detection, the two-stage approach (e.g., Faster R-CNN) has been achieving the highest accuracy, whereas the one-stage approach (e.g., SSD) has the advantage of high efficiency. To inherit the merits of both while overcoming their disadvantages, in this paper, we propose a novel single-shot based detector, called RefineDet, that achieves better accuracy than two-stage methods and maintains comparable efficiency of one-stage methods. RefineDet consists of two inter-connected modules, namely, the anchor refinement module and the object detection module. Specifically, the former aims to (1) filter out negative anchors to reduce search space for the classifier, and (2) coarsely adjust the locations and sizes of anchors to provide better initialization for the subsequent regressor. The latter module takes the refined anchors as the input from the former to further improve the regression accuracy and predict multi-class label. Meanwhile, we design a transfer connection block to transfer the features in the anchor refinement module to predict locations, sizes and class labels of objects in the object detection module. The multitask loss function enables us to train the whole network in an end-to-end way. Extensive experiments on PASCAL VOC 2007, PASCAL VOC 2012, and MS COCO demonstrate that RefineDet achieves state-of-the-art detection accuracy with high efficiency. Code is available at https://github.com/sfzhang15/RefineDet.
Longyin Wen, Xiao Bian, Zhen Lei 0001, Stan Z. Li
CVPR2
2018 Occlusion-Aware R-CNN: Detecting Pedestrians in a Crowd
Longyin Wen, Xiao Bian, Zhen Lei 0001, Stan Z. Li
ECCV (3)2
2018 Iterative Graph Seeking for Object Tracking
abstract
To effectively solve the challenges in object tracking, such as large deformation and severe occlusion, many existing methods use graph-based models to capture target part relations, and adopt a sequential scheme of target part selection, part matching, and state estimation. However, such methods have two major drawbacks: 1) inaccurate part selection leads to performance deterioration of part matching and state estimation and 2) there are insufficient effective global constraints for local part selection and matching. In this paper, we propose a new object tracking method based on iterative graph seeking, which integrate target part selection, part matching, and state estimation using a unified energy minimization framework. Our method also incorporates structural information in local parts variations using the global constraint. We devise an alternative iteration scheme to minimize the energy function for searching the most plausible target geometric graph. Experimental results on several challenging benchmarks (i.e., VOT2015, OTB2013, and OTB2015) demonstrate improved performance and robustness in comparison with existing algorithms.
Dawei Du, Longyin Wen, Honggang Qi, Qingming Huang, Qi Tian 0001, Siwei Lyu
IEEE Trans. Image Process.2
2018 Contrast Enhancement Estimation for Digital Image Forensics
abstract
Inconsistency in contrast enhancement can be used to expose image forgeries. In this work, we describe a new method to estimate contrast enhancement operations from a single image. Our method takes advantage of the nature of contrast enhancement as a mapping between pixel values and the distinct characteristics it introduces to the image pixel histogram. Our method recovers the original pixel histogram and the contrast enhancement simultaneously from a single image with an iterative algorithm. Unlike previous works, our method is robust in the presence of additive noise perturbations that are used to hide the traces of contrast enhancement. Furthermore, we also develop an effective method to detect image regions undergone contrast enhancement transformations that are different from the rest of the image, and we use this method to detect composite images. We perform extensive experimental evaluations to demonstrate the efficacy and efficiency of our method.
Longyin Wen, Honggang Qi, Siwei Lyu
ACM Trans. Multim. Comput. Commun. Appl.1
2017 Unsupervised Learning of Multi-Level Descriptors for Person Re-Identification
abstract
In this paper, we propose a novel coding method named weighted linear coding (WLC) to learn multi-level (e.g., pixel-level, patch-level and image-level) descriptors from raw pixel data in an unsupervised manner. It guarantees the property of saliency with a similarity constraint. The resulting multi-level descriptors have a good balance between the robustness and distinctiveness. Based on WLC, all data from the same region can be jointly encoded. Consequently, when we extract the holistic image features, it is able to preserve the spatial consistency. Furthermore, we apply PCA to these features and compact person representations are then achieved. During the stage of matching persons, we exploit the complementary information resided in multi-level descriptors via a score-level fusion strategy. Experiments on the challenging person re-identification datasets - VIPeR and CUHK 01, demonstrate the effectiveness of our method.
Yang Yang 0062, Longyin Wen, Siwei Lyu, Stan Z. Li
AAAI2
2017 UA-DETRAC 2017: Report of AVSS2017 & IWT4S Challenge on Advanced Traffic Monitoring
abstract
The rapid advances of transportation infrastructure have led to a dramatic increase in the demand for smart systems capable of monitoring traffic and street safety. Fundamental to these applications are a community-based evaluation platform and benchmark for object detection and multi-object tracking. To this end, we organize the AVSS2017 Challenge on Advanced Traffic Monitoring, in conjunction with the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S), to evaluate the state-of-the-art object detection and multi-object tracking algorithms in the relevance of traffic surveillance. Submitted algorithms are evaluated using the large-scale UA-DETRAC benchmark and evaluation protocol. The benchmark, the evaluation toolkit and the algorithm performance are publicly available from the website http://detrac-db.rit.albany.edu.
Siwei Lyu, Ming-Ching Chang, Dawei Du, Longyin Wen, Honggang Qi, Yuezun Li, Yi Wei 0006, Lipeng Ke, Tao Hu 0011, Marco Del Coco, Pierluigi Carcagnì, Dmitriy Anisimov, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Hao Ye 0005, Hong Wang 0014, Kannappan Palaniappan, Koray Ozcan, Li Wang 0033, Liang Wang 0001, Martin Lauer, Nattachai Watcharapinchai, Nenghui Song, Noor Al-Shakarji, Sikandar Amin, Sitapa Watcharapinchai, Tatiana Khanova, Thomas Sikora, Tino Kutschbach, Volker Eiselein, Wei Tian 0001, Xiangyang Xue 0001, Xiaoyi Yu, Yao Lu 0028, Yingbin Zheng, Yongzhen Huang, Yuqi Zhang 0001
AVSS4
2017 Adaptive RNN Tree for Large-Scale Human Action Recognition
abstract
In this work, we present the RNN Tree (RNN-T), an adaptive learning framework for skeleton based human action recognition. Our method categorizes action classes and uses multiple Recurrent Neural Networks (RNNs) in a treelike hierarchy. The RNNs in RNN-T are co-trained with the action category hierarchy, which determines the structure of RNN-T. Actions in skeletal representations are recognized via a hierarchical inference process, during which individual RNNs differentiate finer-grained action classes with increasing confidence. Inference in RNN-T ends when any RNN in the tree recognizes the action with high confidence, or a leaf node is reached. RNN-T effectively addresses two main challenges of large-scale action recognition: (i) able to distinguish fine-grained action classes that are intractable using a single network, and (ii) adaptive to new action classes by augmenting an existing model. We demonstrate the effectiveness of RNN-T/ACH method and compare it with the state-of-the-art methods on a large-scale dataset and several existing benchmarks.
Wenbo Li 0001, Longyin Wen, Ming-Ching Chang, Ser-Nam Lim, Siwei Lyu
ICCV2
2017 Hybrid structure hypergraph for online deformable object tracking
abstract
Recent advances in visual tracking field design part-based model to handle the deformation and occlusion challenges. Previous methods only consider the sole degree of dependencies (e.g., pairwise or high-order dependencies) between object parts in consecutive frames. However, the degree of dependencies of different object parts in consecutive frames are not consistent, especially when large deformation and occlusion happen. To that end, we design a hybrid structure hypergraph based tracker, which use a non-uniform hypergraph to model the dependencies among object parts. The tracking task is further formulated as the dense structures extracting problem on the non-uniform hypergraph, which is solved by an approximate algorithm efficiently. Several experiments are carried out on publicly available online deformable object tracking dataset, i.e., Deform-SOT dataset, to demonstrate the favorable performance of the proposed method against the state-of-the-art online tracking methods.
Shengkun Li, Dawei Du, Longyin Wen, Ming-Ching Chang, Siwei Lyu
ICIP3
2017 Multi-Camera Multi-Target Tracking with Space-Time-View Hyper-graph
Longyin Wen, Zhen Lei 0001, Ming-Ching Chang, Honggang Qi, Siwei Lyu
Int. J. Comput. Vis.1
2017 Geometric Hypergraph Learning for Visual Tracking
abstract
Graph-based representation is widely used in visual tracking field by finding correct correspondences between target parts in different frames. However, most graph-based trackers consider pairwise geometric relations between local parts. They do not make full use of the target's intrinsic structure, thereby making the representation easily disturbed by errors in pairwise affinities when large deformation or occlusion occurs. In this paper, we propose a geometric hypergraph learning-based tracking method, which fully exploits high-order geometric relations among multiple correspondences of parts in different frames. Then visual tracking is formulated as the mode-seeking problem on the hypergraph in which vertices represent correspondence hypotheses and hyperedges describe high-order geometric relations among correspondences. Besides, a confidence-aware sampling method is developed to select representative vertices and hyperedges to construct the geometric hypergraph for more robustness and scalability. The experiments are carried out on three challenging datasets (VOT2014, OTB100, and Deform-SOT) to demonstrate that our method performs favorably against other existing trackers.
Dawei Du, Honggang Qi, Longyin Wen, Qi Tian 0001, Qingming Huang, Siwei Lyu
IEEE Trans. Cybern.3
2016 Efficient large-scale photometric reconstruction using Divide-Recon-Fuse 3D Structure from Motion
abstract
We propose an efficient framework for large-scale 3D reconstruction from a large set of photos following the Structure-from-Motion (SfM) paradigm with divide-conquer and fusion. Our main novelty is to ensure commonality from overlaps between image sets corresponding to their reconstructions, which facilitates effective stitching and fusion. Specifically, such commonality is ensured by selecting a set of duplicated images (which are termed anchor images) in adjacent image sets prior to the 3D reconstruction. The anchor images can assist accurate fusion of the 3D point clouds. We describe an efficient RANSAC scheme for pairwise stitching. Our method is intuitively scalable to large site reconstruction via subdivision and fusion following a graph construct. We further describe another RANSAC algorithm to improve loop closure in our anchor image approach. Experimental results on reconstructing a large portion of a university campus demonstrate the efficacy of our method.
Yueming Yang, Ming-Ching Chang, Longyin Wen, Peter H. Tu, Honggang Qi, Siwei Lyu
AVSS3
2016 Stochastic Online AUC Maximization
abstract
Area under ROC (AUC) is a metric which is widely used for measuring the classification performance for imbalanced data. It is of theoretical and practical interest to develop online learning algorithms that maximizes AUC for large-scale data. A specific challenge in developing online AUC maximization algorithm is that the learning objective function is usually defined over a pair of training examples of opposite classes, and existing methods achieves on-line processing with higher space and time complexity. In this work, we propose a new stochastic online algorithm for AUC maximization. In particular, we show that AUC optimization can be equivalently formulated as a convex-concave saddle point problem. From this saddle representation, a stochastic online algorithm (SOLAM) is proposed which has time and space complexity of one datum. We establish theoretical convergence of SOLAM with high probability and demonstrate its effectiveness and efficiency on standard benchmark datasets.
Yiming Ying, Longyin Wen, Siwei Lyu
NIPS2
2016 Exploiting Hierarchical Dense Structures on Hypergraphs for Multi-Object Tracking
abstract
Most multi-object tracking algorithms are developed within the tracking-by-detection framework that consider the pairwise appearance similarities between detection responses or tracklets within a limited temporal window, and thus less effective in handling long-term occlusions or distinguishing spatially close targets with similar appearance in crowded scenes. In this work, we propose an algorithm that formulates the multi-object tracking task as one to exploit hierarchical dense structures on an undirected hypergraph constructed based on tracklet affinity. The dense structures indicate a group of vertices that are inter-connected with a set of hyperedges with high affinity values. The appearance and motion similarities among multiple tracklets across the spatio-temporal domain are considered globally by exploiting high-order similarities rather than pairwise ones, thereby facilitating distinguish spatially close targets with similar appearance. In addition, the hierarchical design of the optimization process helps the proposed tracking algorithm handle long-term occlusions robustly. Extensive experiments on various challenging datasets of both multi-pedestrian and multi-face tracking tasks, demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods.
Longyin Wen, Zhen Lei 0001, Siwei Lyu, Stan Z. Li, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 Online Deformable Object Tracking Based on Structure-Aware Hyper-Graph
abstract
Recent advances in online visual tracking focus on designing part-based model to handle the deformation and occlusion challenges. However, previous methods usually consider only the pairwise structural dependences of target parts in two consecutive frames rather than the higher order constraints in multiple frames, making them less effective in handling large deformation and occlusion challenges. This paper describes a new and efficient method for online deformable object tracking. Different from most existing methods, this paper exploits higher order structural dependences of different parts of the tracking target in multiple consecutive frames. We construct a structure-aware hyper-graph to capture such higher order dependences, and solve the tracking problem by searching dense subgraphs on it. Furthermore, we also describe a new evaluating data set for online deformable object tracking (the Deform-SOT data set), which includes 50 challenging sequences with full annotations that represent realistic tracking challenges, such as large deformations and severe occlusions. The experimental result of the proposed method shows considerable improvement in performance over the state-of-the-art tracking methods.
Dawei Du, Honggang Qi, Wenbo Li 0001, Longyin Wen, Qingming Huang, Siwei Lyu
IEEE Trans. Image Process.4
2015 JOTS: Joint Online Tracking and Segmentation
abstract
We present a novel Joint Online Tracking and Segmentation (JOTS) algorithm which integrates the multi-part tracking and segmentation into a unified energy optimization framework to handle the video segmentation task. The multi-part segmentation is posed as a pixel-level label assignment task with regularization according to the estimated part models, and tracking is formulated as estimating the part models based on the pixel labels, which in turn is used to refine the model. The multi-part tracking and segmentation are carried out iteratively to minimize the proposed objective function by a RANSAC-style approach. Extensive experiments on the SegTrack and SegTrack v2 databases demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods.
Longyin Wen, Dawei Du, Zhen Lei 0001, Stan Z. Li, Ming-Hsuan Yang 0001
CVPR1
2015 Category-Blind Human Action Recognition: A Practical Recognition System
abstract
Existing human action recognition systems for 3D sequences obtained from the depth camera are designed to cope with only one action category, either single-person action or two-person interaction, and are difficult to be extended to scenarios where both action categories co-exist. In this paper, we propose the category-blind human recognition method (CHARM) which can recognize a human action without making assumptions of the action category. In our CHARM approach, we represent a human action (either a single-person action or a two-person interaction) class using a co-occurrence of motion primitives. Subsequently, we classify an action instance based on matching its motion primitive co-occurrence patterns to each class representation. The matching task is formulated as maximum clique problems. We conduct extensive evaluations of CHARM using three datasets for single-person actions, two-person interactions, and their mixtures. Experimental results show that CHARM performs favorably when compared with several state-of-the-art single-person action and two-person interaction based methods without making explicit assumptions of action category.
Wenbo Li 0001, Longyin Wen, Mooi Choo Chuah, Siwei Lyu
ICCV2
2015 Online Visual Tracking Using Temporally Coherent Part Cluster
abstract
Recent advances in visual tracking have focused on handling deformations and occlusions using the part-based appearance model. However, it remains a challenge to come up with a reliable target representation using local parts, and hence existing trackers continue to face drifting problems. To deal with this challenge, we propose a robust online model, formulating the tracking task as a problem of identifying Temporally Coherent Part (TCP) clusters. Specifically, we pose the TCP clusters identification task as a dense neighborhoods searching problem using a relational hyper graph in which the relationship among multiple temporal local parts is encoded as the affinity value of a hyper edge connecting them. Such high-order relations ships among multiple local parts across the temporal domain make our tracker more robust towards deformations and occlusions. Extensive experiments on various challenging video sequences demonstrate that our TCP-based method performs better than the state-of-the-art methods.
Wenbo Li 0001, Longyin Wen, Mooi Choo Chuah, Yi Zhang 0070, Zhen Lei 0001, Stan Z. Li
WACV2
2014 Multiple Target Tracking Based on Undirected Hierarchical Relation Hypergraph
abstract
Multi-target tracking is an interesting but challenging task in computer vision field. Most previous data association based methods merely consider the relationships (e.g. appearance and motion pattern similarities) between detections in local limited temporal domain, leading to their difficulties in handling long-term occlusion and distinguishing the spatially close targets with similar appearance in crowded scenes. In this paper, a novel data association approach based on undirected hierarchical relation hypergraph is proposed, which formulates the tracking task as a hierarchical dense neighborhoods searching problem on the dynamically constructed undirected affinity graph. The relationships between different detections across the spatiotemporal domain are considered in a high-order way, which makes the tracker robust to the spatially close targets with similar appearance. Meanwhile, the hierarchical design of the optimization process fuels our tracker to long-term occlusion with more robustness. Extensive experiments on various challenging datasets (i.e. PETS2009 dataset, ParkingLot), including both low and high density sequences, demonstrate that the proposed method performs favorably against the state-of-the-art methods.
Longyin Wen, Wenbo Li 0001, Zhen Lei 0001, Dong Yi, Stan Z. Li
CVPR1
2014 The Fastest Deformable Part Model for Object Detection
abstract
This paper solves the speed bottleneck of deformable part model (DPM), while maintaining the accuracy in detection on challenging datasets. Three prohibitive steps in cascade version of DPM are accelerated, including 2D correlation between root filter and feature map, cascade part pruning and HOG feature extraction. For 2D correlation, the root filter is constrained to be low rank, so that 2D correlation can be calculated by more efficient linear combination of 1D correlations. A proximal gradient algorithm is adopted to progressively learn the low rank filter in a discriminative manner. For cascade part pruning, neighborhood aware cascade is proposed to capture the dependence in neighborhood regions for aggressive pruning. Instead of explicit computation of part scores, hypotheses can be pruned by scores of neighborhoods under the first order approximation. For HOG feature extraction, look-up tables are constructed to replace expensive calculations of orientation partition and magnitude with simpler matrix index operations. Extensive experiments show that (a) the proposed method is 4 times faster than the current fastest DPM method with similar accuracy on Pascal VOC, (b) the proposed method achieves state-of-the-art accuracy on pedestrian and face detection task with frame-rate speed.
Zhen Lei 0001, Longyin Wen, Stan Z. Li
CVPR3
2014 A Probabilistic Framework for Multitarget Tracking with Mutual Occlusions
abstract
Mutual occlusions among targets can cause track loss or target position deviation, because the observation likelihood of an occluded target may vanish even when we have the estimated location of the target. This paper presents a novel probability framework for multitarget tracking with mutual occlusions. The primary contribution of this work is the introduction of a vectorial occlusion variable as part of the solution. The occlusion variable describes occlusion states of the targets. This forms the basis of the proposed probability framework, with the following further contributions: 1) Likelihood: A new observation likelihood model is presented, in which the likelihood of an occluded target is computed by referring to both of the occluded and oc-cluding targets. 2) Priori: Markov random field (MRF) is used to model the occlusion priori such that less likely "circular" or "cascading" types of occlusions have lower priori probabilities. Both the occlusion priori and the motion priori take into consideration the state of occlusion. 3) Optimization: A realtime RJMCMC-based algorithm with a newmove type called "occlusion state update" is presented. Experimental results show that the proposed framework can handle occlusions well, even including long-duration full occlusions, which may cause tracking failures in the traditional methods.
Menglong Yang, Yiguang Liu, Longyin Wen, Zhisheng You, Stan Z. Li
CVPR3
2014 Robust Deformable and Occluded Object Tracking With Dynamic Graph
abstract
While some efforts have been paid to handle deformation and occlusion in visual tracking, they are still great challenges. In this paper, a dynamic graph-based tracker (DGT) is proposed to address these two challenges in a unified framework. In the dynamic target graph, nodes are the target local parts encoding appearance information, and edges are the interactions between nodes encoding inner geometric structure information. This graph representation provides much more information for tracking in the presence of deformation and occlusion. The target tracking is then formulated as tracking this dynamic undirected graph, which is also a matching problem between the target graph and the candidate graph. The local parts within the candidate graph are separated from the background with Markov random field, and spectral clustering is used to solve the graph matching. The final target state is determined through a weighted voting procedure according to the reliability of part correspondence, and refined with recourse to a foreground/background segmentation. An effective online updating mechanism is proposed to update the model, allowing DGT to robustly adapt to variations of target structure. Experimental results show improved performance over several state-of-the-art trackers, in various challenging scenarios.
Zhaowei Cai, Longyin Wen, Zhen Lei 0001, Nuno Vasconcelos, Stan Z. Li
IEEE Trans. Image Process.2
2014 Robust Online Learned Spatio-Temporal Context Model for Visual Tracking
abstract
Visual tracking is an important but challenging problem in the computer vision field. In the real world, the appearances of the target and its surroundings change continuously over space and time, which provides effective information to track the target robustly. However, enough attention has not been paid to the spatio-temporal appearance information in previous works. In this paper, a robust spatio-temporal context model based tracker is presented to complete the tracking task in unconstrained environments. The tracker is constructed with temporal and spatial appearance context models. The temporal appearance context model captures the historical appearance of the target to prevent the tracker from drifting to the background in a long-term tracking. The spatial appearance context model integrates contributors to build a supporting field. The contributors are the patches with the same size of the target at the key-points automatically discovered around the target. The constructed supporting field provides much more information than the appearance of the target itself, and thus, ensures the robustness of the tracker in complex environments. Extensive experiments on various challenging databases validate the superiority of our tracker over other state-of-the-art trackers.
Longyin Wen, Zhaowei Cai, Zhen Lei 0001, Dong Yi, Stan Z. Li
IEEE Trans. Image Process.1
2012 Structured Visual Tracking with Dynamic Graph
Zhaowei Cai, Longyin Wen, Zhen Lei 0001, Stan Z. Li
ACCV (3)2
2012 A New Projection Space for Separation of Specular-Diffuse Reflection Components in Color Images
Zhaowei Cai, Longyin Wen, Zhen Lei 0001, Guodong Guo, Stan Z. Li
ACCV (4)3
2012 Online Multiple Instance Joint Model for Visual Tracking
abstract
Although numerous online learning strategies have been proposed to handle the appearance variation in visual tracking, the existing methods just perform well in certain cases since they lack effective appearance learning mechanism. In this paper, a joint model tracker (JMT) is presented, which consists of a generative model based on Multiple Subspaces and a discriminative model based on improved Multiple Instance Boosting (MIBoosting). The generative model utilizes a series of local constructed subspaces to update the Multiple Subspaces model and considers the energy dissipation of dimension reduction in updating step. The discriminative model adopts the Gaussian Mixture Model (GMM) to estimate the posterior probability of the likelihood function. These two parts supervise each other to update in multiple instance way which helps our tracker recover from drift. Extensive experiments on various databases validate the effectiveness of our proposed method over other state-of-the-art trackers.
Longyin Wen, Zhaowei Cai, Menglong Yang, Zhen Lei 0001, Dong Yi, Stan Z. Li
AVSS1
2012 Online Spatio-temporal Structural Context Learning for Visual Tracking
Longyin Wen, Zhaowei Cai, Zhen Lei 0001, Dong Yi, Stan Z. Li
ECCV (4)1