Dawei Du

dblp:51/6958 · DBLP profile ↗
← Back
57ranked-venue papers
11as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 26 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 2 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Sem-Mask Enhanced Cross-modal Alignment for Event Causality Identification
Dawei Du, Luyuan Chen
ICIC (22)1
2025 Human Activity Recognition in an Open World (Abstract Reprint)
abstract
Managing novelty in perception-based human activity recognition (HAR) is critical in realistic settings to improve task performance over time and ensure solution generalization outside of prior seen samples. Novelty manifests in HAR as unseen samples, activities, objects, environments, and sensor changes, among other ways. Novelty may be task-relevant, such as a new class or new features, or task-irrelevant resulting in nuisance novelty, such as never before seen noise, blur, or distorted video recordings. To perform HAR optimally, algorithmic solutions must be tolerant to nuisance novelty, and learn over time in the face of novelty. This paper 1) formalizes the definition of novelty in HAR building upon the prior definition of novelty in classification tasks, 2) proposes an incremental open world learning (OWL) protocol and applies it to the Kinetics datasets to generate a new benchmark KOWL-718, 3) analyzes the performance of current stateof-the-art HAR models when novelty is introduced over time, 4) provides a containerized and packaged pipeline for reproducing the OWL protocol and for modifying for any future updates to Kinetics. The experimental analysis includes an ablation study of how the different models perform under various conditions as annotated by Kinetics-AVA. The code may be used to analyze different annotations and subsets of the Kinetics datasets in an incremental open world fashion, as well as be extended as further updates to Kinetics are released.
Derek S. Prijatelj, Samuel Grieggs, Dawei Du, Ameya Shringi, Christopher Funk, Adam Kaufman, Eric Robertson 0001, Walter J. Scheirer
IJCAI4
2025 Re-identifying People in Video via Learned Temporal Attention and Multi-modal Foundation Models
abstract
Biometric recognition from security camera video is a challenging problem when the individuals change clothes or when they are partly occluded. Others have recently demonstrated that CLIP's visual encoder performs well in this domain, but existing methods fail to make use of the model's text encoder or temporal information available in video. In this paper, we present VCLIP, a method for person identification in videos captured in challenging poses and with changes to a person's clothing. Harnessing the power of pre-trained vision-language models, we Jointly train a temporal fusion network while fine-tuning the visual encoder. To leverage the cross-modal embedding space, we use learned biometric pedestrian attribute features to further enhance our model's person re-identification (Re-ID) ability. We demonstrate significant performance improvements via experiments with the MEVID and CCVID datasets, particularly in the more challenging clothes-changing conditions. In support of this and future methods that use textual attributes for Re-ID with multimodal models, we release a dataset of annotated pedestrian attributes for the popular MEVID dataset [4].
Cole Hill, Florence Yellin, Krishna Regmi, Dawei Du, Scott McCloskey
WACV4
2024 Human Activity Recognition in an Open World
abstract
Managing novelty in perception-based human activity recognition (HAR) is critical in realistic settings to improve task performance over time and ensure solution generalization outside of prior seen samples. Novelty manifests in HAR as unseen samples, activities, objects, environments, and sensor changes, among other ways. Novelty may be task-relevant, such as a new class or new features, or task-irrelevant resulting in nuisance novelty, such as never before seen noise, blur, or distorted video recordings. To perform HAR optimally, algorithmic solutions must be tolerant to nuisance novelty, and learn over time in the face of novelty. This paper 1) formalizes the definition of novelty in HAR building upon the prior definition of novelty in classification tasks, 2) proposes an incremental open world learning (OWL) protocol and applies it to the Kinetics datasets to generate a new benchmark KOWL-718, 3) analyzes the performance of current stateof-the-art HAR models when novelty is introduced over time, 4) provides a containerized and packaged pipeline for reproducing the OWL protocol and for modifying for any future updates to Kinetics. The experimental analysis includes an ablation study of how the different models perform under various conditions as annotated by Kinetics-AVA. The code may be used to analyze different annotations and subsets of the Kinetics datasets in an incremental open world fashion, as well as be extended as further updates to Kinetics are released.
Derek S. Prijatelj, Samuel Grieggs, Dawei Du, Ameya Shringi, Christopher Funk, Adam Kaufman, Eric Robertson 0001, Walter J. Scheirer
J. Artif. Intell. Res.4
2024 Multiview Deep Subspace Clustering Networks
abstract
Multiview subspace clustering aims to discover the inherent structure of data by fusing multiple views of complementary information. Most existing methods first extract multiple types of handcrafted features and then learn a joint affinity matrix for clustering. The disadvantage of this approach lies in two aspects: 1) multiview relations are not embedded into feature learning and 2) the end-to-end learning manner of deep learning is not suitable for multiview clustering. Even when deep features have been extracted, it is a nontrivial problem to choose a proper backbone for clustering on different datasets. To address these issues, we propose the multiview deep subspace clustering networks (MvDSCNs), which learns a multiview self-representation matrix in an end-to-end manner. The MvDSCN consists of two subnetworks, i.e., a diversity network (Dnet) and a universality network (Unet). A latent space is built using deep convolutional autoencoders, and a self-representation matrix is learned in the latent space using a fully connected layer. Dnet learns view-specific self-representation matrices, whereas Unet learns a common self-representation matrix for all views. To exploit the complementarity of multiview representations, the Hilbert-Schmidt independence criterion (HSIC) is introduced as a diversity regularizer that captures the nonlinear, high-order interview relations. Because different views share the same label space, the self-representation matrices of each view are aligned to the common one by universality regularization. The MvDSCN also unifies multiple backbones to boost clustering performance and avoid the need for model selection. Experiments demonstrate the superiority of the MvDSCN.
Pengfei Zhu 0001, Xinjie Yao, Yu Wang 0106, Binyuan Hui, Dawei Du, Qinghua Hu
IEEE Trans. Cybern.5
2023 Open Set Action Recognition via Multi-Label Evidential Learning
abstract
Existing methods for open set action recognition focus on novelty detection that assumes video clips show a single action, which is unrealistic in the real world. We propose a new method for open set action recognition and novelty detection via MUlti-Label Evidential learning (MULE), that goes beyond previous novel action detection methods by addressing the more general problems of single or multiple actors in the same scene, with simultaneous action(s) by any actor. Our Beta Evidential Neural Network estimates multi-action uncertainty with Beta densities based on actor-context-object relation representations. An evidence debiasing constraint is added to the objective function for optimization to reduce the static bias of video representations, which can incorrectly correlate predictions and static cues. We develop a primal-dual average scheme update-based learning algorithm to optimize the proposed problem and provide corresponding theoretical analysis. Besides, uncertainty and belief-based novelty estimation mechanisms are formulated to detect novel actions. Extensive experiments on two real-world video datasets show that our proposed approach achieves promising performance in single/multi-actor, single/multi-action settings. Our code and models are released at https://github.com/charliezhaoyinpeng/mule.
Chen Zhao 0010, Dawei Du, Anthony Hoogs, Christopher Funk
CVPR2
2023 DOERS: Distant Observation Enhancement and Recognition System
abstract
In order to recognize people across long distances and from elevated viewpoints, biometric systems must handle the challenges of imaging through atmospheric turbulence and non-frontal presentations, in addition to the traditional A-PIE challenges of aging, pose, illumination, and expression. While individual biometric modalities such as facial appearance, gait, and whole body appearance each have a role to play, no single modality can address all of these challenges. This paper describes a novel multi-modal biometric recognition system that addresses the challenges of atmospheric turbulence, occlusions, and elevated viewpoints by combining these modalities. We demonstrate our system on both $R G B$ video-based identity verification and both open and closed-world search.
Dawei Du, Cole Hill, Gabriel Bertocco, Maurício Pamplona Segundo, Wes Robbins, Brandon RichardWebster, Roderic Collins, Sudeep Sarkar, Terrance E. Boult, Scott McCloskey
IJCB1
2023 AG-ReID 2023: Aerial-Ground Person Re-identification Challenge Results
abstract
Person re-identification (Re-ID) on aerial-ground platforms has emerged as an intriguing topic within computer vision, presenting a plethora of unique challenges. Highflying altitudes of aerial cameras make persons appear differently in terms of viewpoints, poses, and resolution compared to the images of the same person viewed from ground cameras. Despite its potential, few algorithms have been developed for person re-identification on aerial-ground data, mainly due to the absence of comprehensive datasets. In response, we have collected a large-scale dataset and organized the Aerial-Ground person Re-IDentification Challenge (AG-ReID2023) to foster advancements in the field. The dataset comprises 100,502 images with 1,615 unique identities, including 51,530 training images featuring 807 identities. The test set is divided into two subsets: Aerial to Ground (808 ids, 4,348 query images, 19,259 gallery images) and Ground to Aerial (808 ids, 4,151 query images, 21,214 gallery images). In addition, we manually annotate individuals with their matching IDs across cameras and provide 15 soft attribute labels. The AG-ReID2023 Challenge in conjunction with the 7thIEEE International Joint Conference on Biometrics (IJCB) has garnered interest from numerous institutes, resulting in the submission of five distinct algorithms. We provide an in-depth examination of the evaluation outcomes and present our findings from the contest. For additional details, kindly refer to the official website1.1https://agreid23.github.io.
Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Dana Michalski, Debayan Deb, Mahak Kothari, Manisha Saini, Dawei Du, Scott McCloskey, Gabriel Bertocco, Fernanda A. Andaló, Terrance E. Boult, Anderson Rocha 0001, Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia, Zaigham A. Randhawa, Sinan Sabri, Gianfranco Doretto
IJCB12
2023 Novel Object Detection in Remote Sensing Imagery
abstract
Novel object detection in remote sensing is challenging due to small objects, background clutter and open-set recognition. To discover and identify objects that are not within the set of known classes in training, we developed the extreme value theory-based novel object detection framework. Specifically, we first employed the state-of-the-art object detector to extract object detections. Then, the extreme value model (EVM) is trained based on the features of extracted detections. Thus our method can characterize the distribution of outliers in the known classes-based distributions to classify novel objects. If the novelty score is larger than the pre-set threshold, we assign this sample to novel classes; otherwise the sample is classified by the original object detector. To adapt to our task, we hold out a novel set of 18 of overall 60 classes in the xView dataset in satellite imagery. The experimental results on the xView dataset show the effectiveness of our proposed approach over traditional softmax thresholding.
Dawei Du, Christopher Funk, Katarina Doctor, Anthony Hoogs
IGARSS1
2023 MEVID: Multi-view Extended Videos with Identities for Video Person Re-Identification
abstract
In this paper, we present the Multi-view Extended Videos with Identities (MEVID) dataset for large-scale, video person re-identification (ReID) in the wild. To our knowledge, MEVID represents the most-varied video person ReID dataset, spanning an extensive indoor and outdoor environment across nine unique dates in a 73-day window, various camera viewpoints, and entity clothing changes. Specifically, we label the identities of 158 unique people wearing 598 outfits taken from 8, 092 tracklets, average length of about 590 frames, seen in 33 camera views from the very-large-scale MEVA person activities dataset. While other datasets have more unique identities, MEVID emphasizes a richer set of information about each individual, such as: 4 outfits/identity vs. 2 outfits/identity in CCVID, 33 viewpoints across 17 locations vs. 6 in 5 simulated locations for MTA, and 10 million frames vs. 3 million for LS-VID. Being based on the MEVA video dataset, we also inherit data that is intentionally demographically balanced to the continental United States. To accelerate the annotation process, we developed a semi-automatic annotation framework and GUI that combines state-of-the-art real-time models for object detection, pose estimation, person ReID, and multi-object tracking. We evaluate several state-of-the-art methods on MEVID challenge problems and comprehensively quantify their robustness in terms of changes of outfit, scale, and background location. Our quantitative analysis on the realistic, unique aspects of MEVID shows that there are significant remaining challenges in video person ReID and indicates important directions for future research.
Daniel Davila, Dawei Du, Bryon Lewis, Christopher Funk, Joseph VanPelt, Roderic Collins, Kellie Corona, Matt S. Brown, Scott McCloskey, Anthony Hoogs, Brian Clipp
WACV2
2023 Reconstructing Humpty Dumpty: Multi-feature Graph Autoencoder for Open Set Action Recognition
abstract
Most action recognition datasets and algorithms assume a closed world, where all test samples are instances of the known classes. In open set problems, test samples may be drawn from either known or unknown classes. Existing open set action recognition methods are typically based on extending closed set methods by adding post hoc analysis of classification scores or feature distances and do not capture the relations among all the video clip elements. Our approach uses the reconstruction error to determine the novelty of the video since unknown classes are harder to put back together and thus have a higher reconstruction error than videos from known classes. We refer to our solution to the open set action recognition problem as "Humpty Dumpty", due to its reconstruction abilities. Humpty Dumpty is a novel graph-based autoencoder that accounts for contextual and semantic relations among the clip pieces for improved reconstruction. A larger reconstruction error leads to an increased likelihood that the action can not be reconstructed, i.e., can not put Humpty Dumpty back together again, indicating that the action has never been seen before and is novel/unknown. Extensive experiments are performed on two publicly available action recognition datasets including HMDB-51 and UCF-101, showing the state-of-the-art performance for open set action recognition.
Dawei Du, Ameya Shringi, Anthony Hoogs, Christopher Funk
WACV1
2022 Cascade Transformers for End-to-End Person Search
abstract
The goal of person search is to localize a target person from a gallery set of scene images, which is extremely challenging due to large scale variations, pose/viewpoint changes, and occlusions. In this paper, we propose the Cascade Occluded Attention Transformer (COAT) for end-to-end person search. Our three-stage cascade design focuses on detecting people in the first stage, while later stages simultaneously and progressively refine the representation for person detection and re-identification. At each stage the occluded attention transformer applies tighter intersection over union thresholds, forcing the network to learn coarse-to-fine pose/scale invariant features. Meanwhile, we calculate each detection's occluded attention to differentiate a person's tokens from other people or the background. In this way, we simulate the effect of other objects occluding a person of interest at the token-level. Through comprehensive experiments, we demonstrate the benefits of our method by achieving state-of-the-art performance on two benchmark datasets.
Rui Yu 0002, Dawei Du, Rodney LaLonde, Daniel Davila, Christopher Funk, Anthony Hoogs, Brian Clipp
CVPR2
2022 Multi-Granularity Alignment Domain Adaptation for Object Detection
abstract
Domain adaptive object detection is challenging due to distinctive data distribution between source domain and target domain. In this paper, we propose a unified multi-granularity alignment based object detection framework towards domain-invariant feature learning. To this end, we encode the dependencies across different granularity perspectives including pixel-, instance-, and category-levels simultaneously to align two domains. Based on pixel-level feature maps from the backbone network, we first develop the omniscale gated fusion module to aggregate discriminative representations of instances by scale-aware convolutions, leading to robust multi-scale object detection. Meanwhile, the multi-granularity discriminators are proposed to identify which domain different granularities of samples (i.e., pixels, instances, and categories) come from. Notably, we leverage not only the instance discriminability in different categories but also the category consistency between two domains. Extensive experiments are carried out on multiple domain adaptation scenarios, demonstrating the effectiveness of our framework over state-of-the-art algorithms on top of anchor-free FCOS and anchor-based Faster R-CNN detectors with different backbones.
Wenzhang Zhou, Dawei Du, Libo Zhang 0001, Tiejian Luo
CVPR2
2022 Unbiased Multi-modality Guidance for Image Inpainting
Dawei Du, Libo Zhang 0001, Tiejian Luo
ECCV (16)2
2022 Novelty Detection in Remote Sensing Imagery
abstract
Object detection and classification in remote sensing imagery have been studied for decades, and has had a resurgence recently with significant improvements from deep learning. Most approaches follow the standard target recognition paradigm by assuming a fixed set of known object classes. The detector/classifier is trained on these, and attempts to disregard everything else. However, the real-world is complicated and unpredictable; often, there are new, interesting objects that are similar to known classes, but sufficiently different such that the system will (correctly) ignore them. The goal of novelty detection is to detect instances of new object types rather than misclassifying them as known types or background, while continuing to correctly classify instances of known object types. The primary challenge in novelty detection is determining how different a new image should be in order to be novel, vs. a new condition or variant of a known class. To address this, our method performs novelty detection in imagery using extreme value theory (EVT) operating in a CNN-based feature space. EVT characterizes the distribution of outliers in long-tailed distributions to identify novelties. We conducted experiments on the xView dataset for object detection and classification in satellite imagery, reducing it to a classification dataset by using its annotated bounding boxes on objects and holding out a set of 18 of its 60 classes as novelties. Our results indicate that EVT is effective at distinguishing novel from known object classes, even when novel classes are similar to known ones.
Dawei Du, Christopher Funk, Anthony Hoogs
IGARSS1
2022 Detection and Tracking Meet Drones Challenge
abstract
Drones, or general UAVs, equipped with cameras have been fast deployed with a wide range of applications, including agriculture, aerial photography, and surveillance. Consequently, automatic understanding of visual data collected from drones becomes highly demanding, bringing computer vision and drones more and more closely. To promote and track the developments of object detection and tracking algorithms, we have organized three challenge workshops in conjunction with ECCV 2018, ICCV 2019 and ECCV 2020, attracting more than 100 teams around the world. We provide a large-scale drone captured dataset, VisDrone, which includes four tracks, i.e., (1) image object detection, (2) video object detection, (3) single object tracking, and (4) multi-object tracking. In this paper, we first present a thorough review of object detection and tracking datasets and benchmarks, and discuss the challenges of collecting large-scale drone-based object detection and tracking datasets with fully manual annotations. After that, we describe our VisDrone dataset, which is captured over various urban/suburban areas of 14 different cities across China from North to South. Being the largest such dataset ever published, VisDrone enables extensive evaluation and investigation of visual analysis algorithms for the drone platform. We provide a detailed analysis of the current state of the field of large-scale object detection and tracking on drones, and conclude the challenge as well as propose future directions. We expect the benchmark largely boost the research and development in video analysis on drone platforms. All the datasets and experimental results can be downloaded from https://github.com/VisDrone/VisDrone-Dataset.
Pengfei Zhu 0001, Longyin Wen, Dawei Du, Xiao Bian, Heng Fan 0001, Qinghua Hu, Haibin Ling
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Rethinking Object Detection in Retail Stores
abstract
The conventional standard for object detection uses a bounding box to represent each individual object instance. However, it is not practical in the industry-relevant applications in the context of warehouses due to severe occlusions among groups of instances of the same categories. In this paper, we propose a new task, i.e., simultaneously object localization and counting, abbreviated as Locount, which requires algorithms to localize groups of objects of interest with the number of instances. However, there does not exist a dataset or benchmark designed for such a task. To this end, we collect a large-scale object localization and counting dataset with rich annotations in retail stores, which consists of 50,394 images with more than 1.9 million object instances in 140 categories. Together with this dataset, we provide a new evaluation protocol and divide the training and testing subsets to fairly evaluate the performance of algorithms for Locount, developing a new benchmark for the Locount task. Moreover, we present a cascaded localization and counting network as a strong baseline, which gradually classifies and regresses the bounding boxes of objects with the predicted numbers of instances enclosed in the bounding boxes, trained in an end-to-end manner. Extensive experiments are conducted on the proposed dataset to demonstrate its significance and the analysis is provided to indicate future directions. Dataset is available at https://isrc.iscas.ac.cn/gitlab/research/locount-dataset.
Yuanqiang Cai, Longyin Wen, Libo Zhang 0001, Dawei Du, Weiqiang Wang 0001
AAAI4
2021 Detection, Tracking, and Counting Meets Drones in Crowds: A Benchmark
abstract
To promote the developments of object detection, tracking and counting algorithms in drone-captured videos, we construct a benchmark with a new drone-captured large-scale dataset, named as DroneCrowd, formed by 112 video clips with 33, 600 HD frames in various scenarios. Notably, we annotate 20, 800 people trajectories with 4.8 million heads and several video-level attributes. Meanwhile, we design the Space-Time Neighbor-Aware Network (STNNet) as a strong baseline to solve object detection, tracking and counting jointly in dense crowds. STNNet is formed by the feature extraction module, followed by the density map estimation heads, and localization and association subnets. To exploit the context information of neighboring objects, we design the neighboring context loss to guide the association subnet training, which enforces consistent relative position of nearby objects in temporal domain. Extensive experiments on our DroneCrowd dataset demonstrate that STNNet performs favorably against the state-of-the-arts.
Longyin Wen, Dawei Du, Pengfei Zhu 0001, Qinghua Hu, Qilong Wang 0001, Liefeng Bo, Siwei Lyu
CVPR2
2021 Non-deterministic and emotional chatting machine: learning emotional conversation generation using conditional variational autoencoders
Kaichun Yao, Libo Zhang 0001, Tiejian Luo, Dawei Du
Neural Comput. Appl.4
2021 Scale-Residual Learning Network for Scene Text Detection
abstract
Detecting incidentally captured text in the wild remains an open problem due to challenging factors including unconstrained scenarios and large scale variation. In this paper, we establish a large-scale scene text detection dataset (LS-Text), containing 36, 000 images and 270, 783 text instances with various scales and complex scenarios, to promote the research of text detection. We propose a Scale-residual Learning Network (SLN) to deal with the scale variation problem in a progressive optimization manner. Specifically, we integrate both learnable feature concatenation and feature up-sampling operator. It can effectively eliminate the residuals between the outputs of SLN and ground-truth text instances by processing both the Feature Fusion Residuals (FFR) and the Scale Transformation Residuals (STR), simultaneously. By stacking multi-scale feature maps in a deep-to-shallow manner, SLN continuously optimizes feature representation by accumulating strong semantic information and rich texture details in a scale-residual learning way. Extensive experimental results on five challenging datasets demonstrate the state-of-the-art performance of the proposed SLN model, and the challenging aspects related to real-world scenarios of the proposed LS-Text dataset. Both the source code of SLN and the LS-Text dataset are available athttps://github.com/SLN-Text-Detection.
Yuanqiang Cai, Chang Liu 0047, Peirui Cheng, Dawei Du, Libo Zhang 0001, Weiqiang Wang 0001, Qixiang Ye
IEEE Trans. Circuits Syst. Video Technol.4
2021 Multi-Drone-Based Single Object Tracking With Agent Sharing Network
abstract
Drones equipped with cameras (UAVs) can dynamically track the target in the air from a broader view compared with static cameras or moving sensors over the ground. However, it is still challenging to accurately track the target using a single drone due to several factors such as appearance variations and severe occlusions. To this end, we collect a newMulti-Drone singleObjectTracking (MDOT) dataset that consists of 92 groups of video clips with 113, 918 high resolution frames taken by two drones and 63 groups of video clips with 145, 875 high resolution frames taken by three drones. Besides, two evaluation metrics are specially designed for multi-drone single object tracking,i.e., automatic fusion score (AFS) and ideal fusion score (IFS). Moreover, the agent sharing network (ASNet) is proposed by integrating self-supervised template sharing, target re-detection, and view-aware fusion of the target from multiple drones into a unified framework, which can improve the tracking accuracy significantly compared with single drone tracking. Extensive experiments on MDOT show that our ASNet significantly outperforms recent state-of-the-art trackers. The dataset can be found inhttps://github.com/VisDrone/MultiDrone.
Pengfei Zhu 0001, Jiayu Zheng, Dawei Du, Longyin Wen, Yiming Sun 0003, Qinghua Hu
IEEE Trans. Circuits Syst. Video Technol.3
2021 Embedding Perspective Analysis Into Multi-Column Convolutional Neural Network for Crowd Counting
abstract
The crowd counting is challenging for deep networks due to several factors. For instance, the networks can not efficiently analyze the perspective information of arbitrary scenes, and they are naturally inefficient to handle the scale variations. In this work, we deliver a simple yet efficient multi-column network, which integrates the perspective analysis method with the counting network. The proposed method explicitly excavates the perspective information and drives the counting network to analyze the scenes. More concretely, we explore the perspective information from the estimated density maps and quantify the perspective space into several separate scenes. We then embed the perspective analysis into the multi-column framework with a recurrent connection. Therefore, the proposed network matches various scales with the different receptive fields efficiently. Secondly, we share the parameters of the branches with various receptive fields. This strategy drives the convolutional kernels to be sensitive to the instances with various scales. Furthermore, to improve the evaluation accuracy of the column with a large receptive field, we propose a transform dilated convolution. The transform dilated convolution breaks the fixed sampling structure of the deep network. Moreover, it needs no extra parameters and training, and the offsets are constrained in a local region, which is designed for the congested scenes. The proposed method achieves state-of-the-art performance on five datasets (ShanghaiTech, UCF CC 50, WorldEXPO'10, UCSD, and TRANCOS).
Guorong Li, Dawei Du, Qingming Huang, Nicu Sebe
IEEE Trans. Image Process.3
2021 SiamCAN: Real-Time Visual Tracking Based on Siamese Center-Aware Network
abstract
In this article, we present a novel Siamese center-aware network (SiamCAN) for visual tracking, which consists of the Siamese feature extraction subnetwork, followed by the classification, regression, and localization branches in parallel. The classification branch is used to distinguish the target from background, and the regression branch is introduced to regress the bounding box of the target. To reduce the impact of manually designed anchor boxes to adapt to different target motion patterns, we design the localization branch to localize the target center directly to assist the regression branch generating accurate results. Meanwhile, we introduce the global context module into the localization branch to capture long-range dependencies for more robustness to large displacements of the target. A multi-scale learnable attention module is used to guide these three branches to exploit discriminative features for better performance. Extensive experiments on 9 challenging benchmarks, namely VOT2016, VOT2018, VOT2019, OTB100, LTB35, LaSOT, TC128, UAV123 and VisDrone-SOT2019 demonstrate that SiamCAN achieves leading accuracy with high efficiency. Our source code is available at https://isrc.iscas.ac.cn/gitlab/research/siamcan.
Wenzhang Zhou, Longyin Wen, Libo Zhang 0001, Dawei Du, Tiejian Luo
IEEE Trans. Image Process.4
2021 Graph Regularized Flow Attention Network for Video Animal Counting From Drones
abstract
In this paper, we propose a large-scale video based animal counting dataset collected by drones (AnimalDrone) for agriculture and wildlife protection. The dataset consists of two subsets, i.e., PartA captured on site by drones and PartB collected from the Internet, with rich annotations of more than 4 million objects in 53, 644 frames and corresponding attributes in terms of density, altitude and view. Moreover, we develop a new graph regularized flow attention network (GFAN) to perform density map estimation in dense crowds of video clips with arbitrary crowd density, perspective, and flight altitude. Specifically, our GFAN method leverages optical flow to warp the multi-scale feature maps in sequential frames to exploit the temporal relations, and then combines the enhanced features to predict the density maps. Moreover, we introduce the multi-granularity loss function including pixel-wise density loss and region-wise count loss to enforce the network to concentrate on discriminative features for different scales of objects. Meanwhile, the graph regularizer is imposed on the density maps of multiple consecutive frames to maintain temporal coherency. Extensive experiments are conducted to demonstrate the effectiveness of the proposed method, compared with several state-of-the-art counting algorithms. The AnimalDrone dataset is available at https://github.com/VisDrone/AnimalDrone.
Pengfei Zhu 0001, Dawei Du, Libo Zhang 0001, Qinghua Hu
IEEE Trans. Image Process.3
2021 Iterative Knowledge Distillation for Automatic Check-Out
abstract
Automatic Check-Out (ACO) provides an object detection based mechanism for retailers to process the purchases of customers automatically. However, it suffers a lot from the domain shift problem because of different data distribution between the single item in training exemplar images and mixed items in testing checkout images. In this paper, we propose a new iterative knowledge distillation method to solve the domain adaptation problem for this task. First, we develop a new augmentation data strategy to generate synthesized checkout images. It can extract segmented items from the training images by the coarse-to-fine strategy and filter items with unrealistic poses by pose pruning. Second, we propose a dual pyramid scale network (DPSNet) to exploit the multi-scale feature representation in joint detection and counting views. Third, the iterative knowledge distillation training strategy is developed to make full use of both image-level and instance-level samples to narrow the semantic gap between source domain and target domain. Extensive experiments on the large-scale Retail Product Checkout (RPC) dataset show the proposed DPSNet can achieve state-of-the-art performance compared with existing methods. The source codes can be found athttps://isrc.iscas.ac.cn/gitlab/research/dpsnet.
Libo Zhang 0001, Dawei Du, Tiejian Luo
IEEE Trans. Multim.2
2020 Attention Convolutional Binary Neural Tree for Fine-Grained Visual Categorization
abstract
Fine-grained visual categorization (FGVC) is an important but challenging task due to high intra-class variances and low inter-class variances caused by deformation, occlusion, illumination, etc. An attention convolutional binary neural tree architecture is presented to address those problems for weakly supervised FGVC. Specifically, we incorporate convolutional operations along edges of the tree structure, and use the routing functions in each node to determine the root-to-leaf computational paths within the tree. The final decision is computed as the summation of the predictions from leaf nodes. The deep convolutional operations learn to capture the representations of objects, and the tree structure characterizes the coarse-to-fine hierarchical feature learning process. In addition, we use the attention transformer module to enforce the network to capture discriminative features. The negative log-likelihood loss is used to train the entire network in an end-to-end fashion by SGD with back-propagation. Several experiments on the CUB-200-2011, Stanford Cars and Aircraft datasets demonstrate that the proposed method performs favorably against the state-of-the-arts.
Ruyi Ji, Longyin Wen, Libo Zhang 0001, Dawei Du, Chen Zhao 0024, Xianglong Liu 0001, Feiyue Huang
CVPR4
2020 Learning Semantic Neural Tree for Human Parsing
Ruyi Ji, Dawei Du, Libo Zhang 0001, Longyin Wen, Chen Zhao 0024, Feiyue Huang, Siwei Lyu
ECCV (13)2
2020 Spatial Attention Pyramid Network for Unsupervised Domain Adaptation
Dawei Du, Libo Zhang 0001, Longyin Wen, Tiejian Luo, Pengfei Zhu 0001
ECCV (13)2
2020 Guided Attention Network for Object Detection and Counting on Drones
abstract
Object detection and counting are related but challenging problems, especially for drone based scenes with small objects and cluttered background. In this paper, we propose a new Guided Attention network (GAnet) to deal with both object detection and counting tasks based on the feature pyramid. Different from the previous methods relying on unsupervised attention modules, we fuse different scales of feature maps by using the proposed weakly-supervised Background Attention (BA) between the background and objects for more semantic feature representation. Then, the Foreground Attention (FA) module is developed to consider both global and local appearance of the object to facilitate accurate localization. Moreover, the new data argumentation strategy is designed to train a robust model in the drone based scenes with various illumination conditions. Extensive experiments on three challenging benchmarks (i.e., UAVDT, CARPK and PUCPR+) show the state-of-the-art detection and counting performance of the proposed method compared with existing methods. Code can be found at https://isrc.iscas.ac.cn/gitlab/research/ganet.
Yuanqiang Cai, Dawei Du, Libo Zhang 0001, Longyin Wen, Weiqiang Wang 0001, Siwei Lyu
ACM Multimedia2
2020 UA-DETRAC: A new benchmark and protocol for multi-object detection and tracking
Longyin Wen, Dawei Du, Zhaowei Cai, Zhen Lei 0001, Ming-Ching Chang, Honggang Qi, Jongwoo Lim, Ming-Hsuan Yang 0001, Siwei Lyu
Comput. Vis. Image Underst.2
2020 The Unmanned Aerial Vehicle Benchmark: Object Detection, Tracking and Baseline
Hongyang Yu 0001, Guorong Li, Weigang Zhang, Qingming Huang, Dawei Du, Qi Tian 0001, Nicu Sebe
Int. J. Comput. Vis.5
2020 Detecting Small Objects Using a Channel-Aware Deconvolutional Network
abstract
Detecting small objects is a challenging task due to their low resolution and noisy representation even using deep learning methods. In this paper, we propose a novel object detection method based on the channel-aware deconvolutional network (CADNet) for accurate small object detection. Specifically, we develop the channel-aware deconvolution (ChaDeConv) layer to exploit the correlations of feature maps in different channels across deeper layers, improving the recall rate of small objects at low additional computational costs. Following the ChaDeConv layer, the multiple region proposal sub-network (Multi-RPN) is employed to supervise and optimize multiple detection layers simultaneously to achieve better accuracy. The Multi-RPN module is only used in the training phase and does not increase the computation cost of the inference. In addition, we design a new anchor matching strategy based on the center point translation (CPTMatching) of anchors to select more extending anchors as positive samples in the training phase. The extensive experiments on the PASCAL VOC 2007/2012, MS COCO, and UAVDT datasets show that the proposed CADNet achieves state-of-the-art performance compared to the existing methods.
Kaiwen Duan, Dawei Du, Honggang Qi, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.2
2020 Dual Encoding for Abstractive Text Summarization
abstract
Recurrent neural network-based sequence-to-sequence attentional models have proven effective in abstractive text summarization. In this paper, we model abstractive text summarization using a dual encoding model. Different from the previous works only using a single encoder, the proposed method employs a dual encoder including the primary and the secondary encoders. Specifically, the primary encoder conducts coarse encoding in a regular way, while the secondary encoder models the importance of words and generates more fine encoding based on the input raw text and the previously generated output text summarization. The two level encodings are combined and fed into the decoder to generate more diverse summary that can decrease repetition phenomenon for long sequence generation. The experimental results on two challenging datasets (i.e., CNN/DailyMail and DUC 2004) demonstrate that our dual encoding model performs against existing methods.
Kaichun Yao, Libo Zhang 0001, Dawei Du, Tiejian Luo, Lili Tao
IEEE Trans. Cybern.3
2020 Robust Obstacle Detection and Recognition for Driver Assistance Systems
abstract
This paper proposes a robust obstacle detection and recognition method for driver assistance systems. Unlike existing methods, our method aims to detect and recognize obstacles on the road rather than all the obstacles in the view. The proposed method involves two stages aiming at an increased quality of the results. The first stage is to locate the positions of obstacles on the road. In order to accurately locate the on-road obstacles, we propose an obstacle detection method based on the U-V disparity map generated from a stereo vision system. The proposed U-V disparity algorithm makes use of the V-disparity map that provides a good representation of the geometric content of the road region to extract the road features, and then detects the on-road obstacles using our proposed realistic U-disparity map that eliminates the foreshortening effects caused by the perspective projection of pinhole imaging. The proposed realistic U-disparity map greatly improves the detection accuracy of the distant obstacles compared with the conventional U-disparity map. Second, the detection results of our proposed U-V disparity algorithm are put into a context-aware Faster-RCNN that combines the interior and contextual features to improve the recognition accuracy of small and occluded obstacles. Specifically, we propose a context-aware module and apply it into the architecture of Faster-RCNN. The experimental results on two public datasets show that our proposed method achieves state-of-the-art performance under various driving conditions.
Jiaxu Leng, Ying Liu 0039, Dawei Du, Pei Quan
IEEE Trans. Intell. Transp. Syst.3
2019 Scale Invariant Fully Convolutional Network: Detecting Hands Efficiently
abstract
Existing hand detection methods usually follow the pipeline of multiple stages with high computation cost, i.e., feature extraction, region proposal, bounding box regression, and additional layers for rotated region detection. In this paper, we propose a new Scale Invariant Fully Convolutional Network (SIFCN) trained in an end-to-end fashion to detect hands efficiently. Specifically, we merge the feature maps from high to low layers in an iterative way, which handles different scales of hands better with less time overhead comparing to concatenating them simply. Moreover, we develop the Complementary Weighted Fusion (CWF) block to make full use of the distinctive features among multiple layers to achieve scale invariance. To deal with rotated hand detection, we present the rotation map to get rid of complex rotation and derotation layers. Besides, we design the multi-scale loss scheme to accelerate the training process significantly by adding supervision to the intermediate layers of the network. Compared with the state-of-the-art methods, our algorithm shows comparable accuracy and runs a 4.23 times faster speed on the VIVA dataset and achieves better average precision on Oxford hand detection dataset at a speed of 62.5 fps.
Dawei Du, Libo Zhang 0001, Tiejian Luo, Feiyue Huang, Siwei Lyu
AAAI2
2019 Learning Non-Uniform Hypergraph for Multi-Object Tracking
abstract
The majority of Multi-Object Tracking (MOT) algorithms based on the tracking-by-detection scheme do not use higher order dependencies among objects or tracklets, which makes them less effective in handling complex scenarios. In this work, we present a new near-online MOT algorithm based on non-uniform hypergraph, which can model different degrees of dependencies among tracklets in a unified objective. The nodes in the hypergraph correspond to the tracklets and the hyperedges with different degrees encode various kinds of dependencies among them. Specifically, instead of setting the weights of hyperedges with different degrees empirically, they are learned automatically using the structural support vector machine algorithm (SSVM). Several experiments are carried out on various challenging datasets (i.e., PETS09, ParkingLot sequence, SubwayFace, and MOT16 benchmark), to demonstrate that our method achieves favorable performance against the state-of-the-art MOT methods.
Longyin Wen, Dawei Du, Shengkun Li, Xiao Bian, Siwei Lyu
AAAI2
2019 Deep Correlated Predictive Subspace Learning for Incomplete Multi-View Semi-Supervised Classification
abstract
Incomplete view information often results in failure cases of the conventional multi-view methods. To address this problem, we propose a Deep Correlated Predictive Subspace Learning (DCPSL) method for incomplete multi-view semi-supervised classification. Specifically, we integrate semi-supervised deep matrix factorization, correlated subspace learning, and multi-view label prediction into a unified framework to jointly learn the deep correlated predictive subspace and multi-view shared and private label predictors. DCPSL is able to learn proper subspace representation that is suitable for class label prediction, which can further improve the performance of classification. Extensive experimental results on various practical datasets demonstrate that the proposed method performs favorably against the state-of-the-art methods.
Zhe Xue, Junping Du 0001, Dawei Du, Wenqi Ren, Siwei Lyu
IJCAI3
2019 Data Priming Network for Automatic Check-Out
abstract
Automatic Check-Out (ACO) receives increased interests in recent years. An important component of the ACO system is the visual item counting, which recognizes the categories and counts of the items chosen by the customers. However, the training of such a system is challenged by the domain adaptation problem, in which the training data are images from isolated items while the testing images are for collections of items. Existing methods solve this problem with data augmentation using synthesized images, but the image synthesis leads to unreal images that affect the training process. In this paper, we propose a new data priming method to solve the domain adaptation problem. Specifically, we first use pre-augmentation data priming, in which we remove distracting background from the training images using the coarse-to-fine strategy and select images with realistic view angles by the pose pruning method. In the post-augmentation step, we train a data priming network using detection and counting collaborative learning, and select more reliable images from testing data to fine-tune the final visual item tallying network. Experiments on the large scale Retail Product Checkout (RPC) dataset demonstrate the superiority of the proposed method, i.e., we achieve 80.51% checkout accuracy compared with 56.68% of the baseline methods. The source codes can be found in https://isrc.iscas.ac.cn/gitlab/research/acm-mm-2019-ACO.
Dawei Du, Libo Zhang 0001, Tiejian Luo, Qi Tian 0001, Longyin Wen, Siwei Lyu
ACM Multimedia2
2019 Deep low-rank subspace ensemble for multi-view clustering
Zhe Xue, Junping Du 0001, Dawei Du, Siwei Lyu
Inf. Sci.3
2019 Deep Constrained Low-Rank Subspace Learning for Multi-View Semi-Supervised Classification
abstract
Semi-supervised classification receives increasing interests because it can predict class labels based on both limited labeled and sufficient unlabeled data. In this letter, we propose a deep constrained low-rank subspace learning (DCLSL) method for multi-view semi-supervised classification. Specifically, we integrate deep constrained matrix factorization, low-rank subspace learning, and class label learning into a unified objective function to jointly learn data similarity matrices and class label matrix. DCLSL is able to obtain the discriminative subspace representation of each view and effectively aggregate similarity matrices of multiple views, resulting in better classification performance. Experimental results on various datasets demonstrate the effectiveness of our method.
Zhe Xue, Junping Du 0001, Dawei Du, Guorong Li, Qingming Huang, Siwei Lyu
IEEE Signal Process. Lett.3
2018 UA-DETRAC 2018: Report of AVSS2018 & IWT4S Challenge on Advanced Traffic Monitoring
abstract
A desirable smart traffic-monitoring and street-safety system can elicit and support the intervention of law enforcement agencies or medical staff. Recently, there has been a dramatically higher demand for such smart systems. To this end, the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S) was organized in conjunction with the 15th IEEE International Conference on Advanced Video and Signal-based Surveillance (AVSS 2018). Our goal is to advance the state-of-the-art detection and tracking algorithms and provide a comprehensive performance evaluation for them. We evaluate 5 submitted detection and 7 submitted tracking methods on the large-scale UA-DETRAC benchmark, and the results are shared publicly on the website http://detrac-db. rit.albany.edu. We expect this challenge to advance the research and development of new detection and tracking methods for transportation applications.
Siwei Lyu, Ming-Ching Chang, Dawei Du, Wenbo Li 0001, Yi Wei 0006, Marco Del Coco, Pierluigi Carcagnì, Arne Schumann, Bharti Munjal, Dinh-Quoc-Trung Dang, Doo-Hyun Choi, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Guna Seetharaman, Jang-Woon Baek, Jong Taek Lee, Kannappan Palaniappan, Kil-Taek Lim, Kiyoung Moon, Kwang-Ju Kim, Lars Wilko Sommer, Meltem Brandlmaier, Minsung Kang, Moongu Jeon, Noor Al-Shakarji, Oliver Acatay, Pyong-Kun Kim, Sikandar Amin, Thomas Sikora, Tien Ba Dinh, Tobias Senst, Vu-Gia-Hy Che, Young-Chul Lim, Yun-Su Chung
AVSS3
2018 The Unmanned Aerial Vehicle Benchmark: Object Detection and Tracking
Dawei Du, Yuankai Qi, Hongyang Yu 0001, Kaiwen Duan, Guorong Li, Weigang Zhang, Qingming Huang, Qi Tian 0001
ECCV (10)1
2018 Global Contrast Enhancement Detection via Deep Multi-Path Network
abstract
Identifying global contrast enhancement in an image is an important task in forensics estimation. Several previous methods analyze the “peak-gap” fingerprints in graylevel histograms. However, images in real scenarios are often stored in the JPEG format with middle/low compression quality, resulting in less obvious “peak-gap” effect and then unsatisfactory performance. In this paper, we propose a novel deep Multi-Path Network (MPNet) based approach to learn discriminative features from graylevel histograms. Specifically, given the histograms, their high-level peaks and gaps information can be exploited effectively after several shared convolutional layers in the network, even in middle/low quality compressed images. Moreover, the proposed multi-path module is able to focus on dealing with specific forensics operations for more robustness on image compression. The experiments on three challenging datasets (i.e., Dresden, RAISE and UCID) demonstrate the effectiveness of the proposed method compared to existing methods.
Dawei Du, Lipeng Ke, Honggang Qi, Siwei Lyu
ICPR2
2018 Iterative Graph Seeking for Object Tracking
abstract
To effectively solve the challenges in object tracking, such as large deformation and severe occlusion, many existing methods use graph-based models to capture target part relations, and adopt a sequential scheme of target part selection, part matching, and state estimation. However, such methods have two major drawbacks: 1) inaccurate part selection leads to performance deterioration of part matching and state estimation and 2) there are insufficient effective global constraints for local part selection and matching. In this paper, we propose a new object tracking method based on iterative graph seeking, which integrate target part selection, part matching, and state estimation using a unified energy minimization framework. Our method also incorporates structural information in local parts variations using the global constraint. We devise an alternative iteration scheme to minimize the energy function for searching the most plausible target geometric graph. Experimental results on several challenging benchmarks (i.e., VOT2015, OTB2013, and OTB2015) demonstrate improved performance and robustness in comparison with existing algorithms.
Dawei Du, Longyin Wen, Honggang Qi, Qingming Huang, Qi Tian 0001, Siwei Lyu
IEEE Trans. Image Process.1
2017 UA-DETRAC 2017: Report of AVSS2017 & IWT4S Challenge on Advanced Traffic Monitoring
abstract
The rapid advances of transportation infrastructure have led to a dramatic increase in the demand for smart systems capable of monitoring traffic and street safety. Fundamental to these applications are a community-based evaluation platform and benchmark for object detection and multi-object tracking. To this end, we organize the AVSS2017 Challenge on Advanced Traffic Monitoring, in conjunction with the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S), to evaluate the state-of-the-art object detection and multi-object tracking algorithms in the relevance of traffic surveillance. Submitted algorithms are evaluated using the large-scale UA-DETRAC benchmark and evaluation protocol. The benchmark, the evaluation toolkit and the algorithm performance are publicly available from the website http://detrac-db.rit.albany.edu.
Siwei Lyu, Ming-Ching Chang, Dawei Du, Longyin Wen, Honggang Qi, Yuezun Li, Yi Wei 0006, Lipeng Ke, Tao Hu 0011, Marco Del Coco, Pierluigi Carcagnì, Dmitriy Anisimov, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Hao Ye 0005, Hong Wang 0014, Kannappan Palaniappan, Koray Ozcan, Li Wang 0033, Liang Wang 0001, Martin Lauer, Nattachai Watcharapinchai, Nenghui Song, Noor Al-Shakarji, Sikandar Amin, Sitapa Watcharapinchai, Tatiana Khanova, Thomas Sikora, Tino Kutschbach, Volker Eiselein, Wei Tian 0001, Xiangyang Xue 0001, Xiaoyi Yu, Yao Lu 0028, Yingbin Zheng, Yongzhen Huang, Yuqi Zhang 0001
AVSS3
2017 Hybrid structure hypergraph for online deformable object tracking
abstract
Recent advances in visual tracking field design part-based model to handle the deformation and occlusion challenges. Previous methods only consider the sole degree of dependencies (e.g., pairwise or high-order dependencies) between object parts in consecutive frames. However, the degree of dependencies of different object parts in consecutive frames are not consistent, especially when large deformation and occlusion happen. To that end, we design a hybrid structure hypergraph based tracker, which use a non-uniform hypergraph to model the dependencies among object parts. The tracking task is further formulated as the dense structures extracting problem on the non-uniform hypergraph, which is solved by an approximate algorithm efficiently. Several experiments are carried out on publicly available online deformable object tracking dataset, i.e., Deform-SOT dataset, to demonstrate the favorable performance of the proposed method against the state-of-the-art online tracking methods.
Shengkun Li, Dawei Du, Longyin Wen, Ming-Ching Chang, Siwei Lyu
ICIP2
2017 Geometric Hypergraph Learning for Visual Tracking
abstract
Graph-based representation is widely used in visual tracking field by finding correct correspondences between target parts in different frames. However, most graph-based trackers consider pairwise geometric relations between local parts. They do not make full use of the target's intrinsic structure, thereby making the representation easily disturbed by errors in pairwise affinities when large deformation or occlusion occurs. In this paper, we propose a geometric hypergraph learning-based tracking method, which fully exploits high-order geometric relations among multiple correspondences of parts in different frames. Then visual tracking is formulated as the mode-seeking problem on the hypergraph in which vertices represent correspondence hypotheses and hyperedges describe high-order geometric relations among correspondences. Besides, a confidence-aware sampling method is developed to select representative vertices and hyperedges to construct the geometric hypergraph for more robustness and scalability. The experiments are carried out on three challenging datasets (VOT2014, OTB100, and Deform-SOT) to demonstrate that our method performs favorably against other existing trackers.
Dawei Du, Honggang Qi, Longyin Wen, Qi Tian 0001, Qingming Huang, Siwei Lyu
IEEE Trans. Cybern.1
2016 Trajectory optimization with memetic algorithms: Time-to-torque minimization of turbocharged engines
abstract
A general memetic trajectory optimization method is introduced. The method is comprised of an evolutionary algorithm (EA) for global optimization, followed by local optimization. The global optimization algorithm is biogeography-based optimization (BBO), which is an EA motivated by the migratory behavior of biological organisms. For local optimization, we start with identifying a local linearized model within the region of the BBO solution by approximating the linear model with Jacobian matrix, and then optimize trajectory using gradient method. The process iterates Jacobian learning and optimization until an optimal trajectory is identified. We apply this memetic algorithm to a time-to-torque minimization problem for a gasoline turbocharged direct injection automotive engine. The optimized trajectory demonstrates significant improvement over the intuitive bang-bang controls that were originally thought to deliver the fastest transient torque response. Simulation results show that BBO decreases time-to-torque by 48% relative to bang-bang controls, and adaptive optimization decreases time-to-torque by an additional 26%. These results have significant implications for improved automotive engine performance.
Dan Simon, Yan Wang 0075, Oliver Tiber, Dawei Du, Dimitar P. Filev, John Michelini
SMC4
2016 Online Deformable Object Tracking Based on Structure-Aware Hyper-Graph
abstract
Recent advances in online visual tracking focus on designing part-based model to handle the deformation and occlusion challenges. However, previous methods usually consider only the pairwise structural dependences of target parts in two consecutive frames rather than the higher order constraints in multiple frames, making them less effective in handling large deformation and occlusion challenges. This paper describes a new and efficient method for online deformable object tracking. Different from most existing methods, this paper exploits higher order structural dependences of different parts of the tracking target in multiple consecutive frames. We construct a structure-aware hyper-graph to capture such higher order dependences, and solve the tracking problem by searching dense subgraphs on it. Furthermore, we also describe a new evaluating data set for online deformable object tracking (the Deform-SOT data set), which includes 50 challenging sequences with full annotations that represent realistic tracking challenges, such as large deformations and severe occlusions. The experimental result of the proposed method shows considerable improvement in performance over the state-of-the-art tracking methods.
Dawei Du, Honggang Qi, Wenbo Li 0001, Longyin Wen, Qingming Huang, Siwei Lyu
IEEE Trans. Image Process.1
2015 JOTS: Joint Online Tracking and Segmentation
abstract
We present a novel Joint Online Tracking and Segmentation (JOTS) algorithm which integrates the multi-part tracking and segmentation into a unified energy optimization framework to handle the video segmentation task. The multi-part segmentation is posed as a pixel-level label assignment task with regularization according to the estimated part models, and tracking is formulated as estimating the part models based on the pixel labels, which in turn is used to refine the model. The multi-part tracking and segmentation are carried out iteratively to minimize the proposed objective function by a RANSAC-style approach. Extensive experiments on the SegTrack and SegTrack v2 databases demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods.
Longyin Wen, Dawei Du, Zhen Lei 0001, Stan Z. Li, Ming-Hsuan Yang 0001
CVPR2
2013 Abnormal event detection in crowded scenes based on Structural Multi-scale Motion Interrelated Patterns
abstract
Detecting abnormal events in crowded scenes remains challenging due to the diversity of events defined by various applications. Among the many application situations, motion analysis for event representation is suited for crowded scenes. In this paper, we propose a novel abnormal event detection method via likelihood estimation of dynamic-texture motion representation, called Structural Multi-scale Motion Interrelated Patterns (SMMIP). SMMIP combines both original motion patterns and their structural spatio-temporal information, which effectively represents localized events by different resolutions of motion patterns. To model normal events, the Gaussian mixture model is trained with the observed normal events, then the likelihood estimation for testing events is computed to judge whether they are abnormal. Meanwhile, the proposed model can be learned online by updating the parameters incrementally. The proposed approach is evaluated on several publicly available datasets and outperforms several other methods proposed before, which is shown that the structural spatio-temporal information added in motion representation helps increasing the anomalies detection rate.
Dawei Du, Honggang Qi, Qingming Huang, Wei Zeng 0006, Changhua Zhang
ICME1
2013 Recover image details from LDR photographs
abstract
In this paper, a novel self-adaptive curve, based on human visual model (HVM), is proposed for recovering details from low dynamic range (LDR) digital photographs, which are under-exposed or over-exposed or both. In order to improve the perceptual visibility, we utilize HVM to construct our method, which is able to take advantage of entire dynamic range to enhance the contrast of images. Extensive experiments demonstrate that our method consistently achieves satisfying results for unwell-exposed LDR photographs.
Kui Fan, Honggang Qi, Dawei Du, Changhua Zhang
ISCAS3
2011 Analytical and numerical comparisons of biogeography-based optimization and genetic algorithms
Dan Simon, Richard A. Rarick, Mehmet Ergezer, Dawei Du
Inf. Sci.4
2011 Markov Models for Biogeography-Based Optimization
abstract
Biogeography-based optimization (BBO) is a population-based evolutionary algorithm that is based on the mathematics of biogeography. Biogeography is the science and study of the geographical distribution of biological organisms. In BBO, problem solutions are analogous to islands, and the sharing of features between solutions is analogous to the migration of species. This paper derives Markov models for BBO with selection, migration, and mutation operators. Our models give the theoretically exact limiting probabilities for each possible population distribution for a given problem. We provide simulation results to confirm the Markov models.
Dan Simon, Mehmet Ergezer, Dawei Du, Richard A. Rarick
IEEE Trans. Syst. Man Cybern. Part B3
2009 Biogeography-Based Optimization Combined with Evolutionary Strategy and Immigration Refusal
abstract
Biogeography-based optimization (BBO) is a recently developed heuristic algorithm which has shown impressive performance on many well known benchmarks. In order to improve BBO, this paper incorporates distinctive features from other successful heuristic algorithms into BBO. In this paper, features from evolutionary strategy (ES) are used for BBO modification. Also, a new immigration refusal approach is added to BBO. After the modification of BBO, F-tests and T-tests are used to demonstrate the differences between different implementations of BBOs.
Dawei Du, Dan Simon, Mehmet Ergezer
SMC1
2009 Oppositional Biogeography-Based Optimization
abstract
We propose a novel variation to biogeography-based optimization (BBO), which is an evolutionary algorithm (EA) developed for global optimization. The new algorithm employs opposition-based learning (OBL) alongside BBO's migration rates to create oppositional BBO (OBBO). Additionally, a new opposition method named quasi-reflection is introduced. Quasi-reflection is based on opposite numbers theory and we mathematically prove that it has the highest expected probability of being closer to the problem solution among all OBL methods. The oppositional algorithm is further revised by the addition of dynamic domain scaling and weighted reflection. Simulations have been performed to validate the performance of quasi-opposition as well as a mathematical analysis for a single-dimensional problem. Empirical results demonstrate that with the assistance of quasi-reflection, OBBO significantly outperforms BBO in terms of success rate and the number of fitness function evaluations required to find an optimal solution.
Mehmet Ergezer, Dan Simon, Dawei Du
SMC3
2009 Population Distributions in Biogeography-Based Optimization Algorithms with Elitism
abstract
Biogeography-based optimization (BBO) is an evolutionary algorithm that is based on the science of biogeography. Biogeography is the study of the geographical distribution of organisms. In BBO, problem solutions are represented as islands, and the sharing of features between solutions is represented as migration between islands. This paper develops a Markov analysis of BBO, including the option of elitism. Our analysis gives the probability of BBO convergence to each possible population distribution for a given problem. We compare our BBO Markov analysis with a similar genetic algorithm (GA) Markov analysis. Analytical comparisons on three simple problems show that with high mutation rates the performance of GAs and BBO is similar, but with low mutation rates BBO outperforms GAs. Our analysis also shows that elitism is not necessary for all problems, but for some problems it can significantly improve performance.
Dan Simon, Mehmet Ergezer, Dawei Du
SMC3