EDBT 2026 Demo / reviewers in the wild / expert
Qilin Zhang 0004
dblp:96/10020-4
· DBLP profile ↗
31ranked-venue papers
2as first author
11since 2021 · last 2024
0000-0002-7917-9749ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 19 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Stepwise Multi-grained Boundary Detector for Point-Supervised Temporal Action Localization
Mengnan Liu 0001, Le Wang 0003, Sanping Zhou, Qilin Zhang 0004, Gang Hua 0001 |
ECCV (7) | 6 |
| 2024 | Adversarial Attack and Defense in Deep RankingabstractDeep Neural Network classifiers are vulnerable to adversarial attacks, where an imperceptible perturbation could result in misclassification. However, the vulnerability of DNN-based image ranking systems remains under-explored. In this paper, we propose two attacks against deep ranking systems, i.e., Candidate Attack and Query Attack, that can raise or lower the rank of chosen candidates by adversarial perturbations. Specifically, the expected ranking order is first represented as a set of inequalities. Then a triplet-like objective function is designed to obtain the optimal perturbation. Conversely, an anti-collapse triplet defense is proposed to improve the ranking model robustness against all proposed attacks, where the model learns to prevent the adversarial attack from pulling the positive and negative samples close to each other. To comprehensively measure the empirical adversarial robustness of a ranking model with our defense, we propose an empirical robustness score, which involves a set of representative attacks against ranking models. Our adversarial ranking attacks and defenses are evaluated on MNIST, Fashion-MNIST, CUB200-2011, CARS196, and Stanford Online Products datasets. Experimental results demonstrate that our attacks can effectively compromise a typical deep ranking system. Nevertheless, our defense can significantly improve the ranking system's robustness and simultaneously mitigate a wide range of attacks. Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Adaptive Two-Stream Consensus Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (W-TAL) aims to classify and localize all action instances in untrimmed videos under only video-level supervision. Without frame-level annotations, it is challenging for W-TAL methods to clearly distinguish actions and background, which severely degrades the action boundary localization and action proposal scoring. In this paper, we present an adaptive two-stream consensus network (A-TSCN) to address this problem. Our A-TSCN features an iterative refinement training scheme: a frame-level pseudo ground truth is generated and iteratively updated from a late-fusion activation sequence, and used to provide frame-level supervision for improved model training. Besides, we introduce an adaptive attention normalization loss, which adaptively selects action and background snippets according to video attention distribution. By differentiating the attention values of the selected action snippets and background snippets, it forces the predicted attention to act as a binary selection and promotes the precise localization of action boundaries. Furthermore, we propose a video-level and a snippet-level uncertainty estimator, and they can mitigate the adverse effect caused by learning from noisy pseudo ground truth. Experiments conducted on the THUMOS14, ActivityNet v1.2, ActivityNet v1.3, and HACS datasets show that our A-TSCN outperforms current state-of-the-art methods, and even achieves comparable performance with several fully-supervised methods. Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Nanning Zheng 0001, David S. Doermann, Junsong Yuan 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Adaptive Ladder Loss for Learning Coherent Visual-Semantic EmbeddingabstractFor visual-semantic embedding, the existing methods normally treat the relevance between queries and candidates in a bipolar way – relevant or irrelevant, and all “irrelevant” candidates are uniformly pushed away from the query by an equal margin in the embedding space, regardless of their various proximity to the query. This practice disregards relatively discriminative information and could lead to suboptimal ranking in the retrieval results and poorer user experience, especially in the long-tail query scenario where a matching candidate may not necessarily exist. In this paper, we introduce a continuous variable to model the relevance degree between queries and multiple candidates, and propose to learn a coherent embedding space, where candidates with higher relevance degrees are mapped closer to the query than those with lower relevance degrees. In particular, the new ladder loss is proposed by extending the triplet loss inequality to a more general inequality chain, which implements variable push-away margins according to respective relevance degrees. To adapt to the varying mini-batch statistics and improve the efficiency of the ladder loss, we also propose a Silhouette score-based method to adaptively decide the ladder level and hence the underlying inequality chain. In addition, a proper Coherent Score metric is proposed to better measure the ranking results including those “irrelevant” candidates. Extensive experiments on multiple datasets validate the efficacy of our proposed method, which achieves significant improvement over existing state-of-the-art methods. Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksabstractGiven only video-level action categorical labels during training, weakly-supervised temporal action localization (WS-TAL) learns to detect action instances and locates their temporal boundaries in untrimmed videos. Compared to its fully supervised counterpart, WS-TAL is more cost-effective in data labeling and thus favorable in practical applications. However, the coarse video-level supervision inevitably incurs ambiguities in action localization, especially in untrimmed videos containing multiple action instances. To overcome this challenge, we observe that significant temporal contrasts among video snippets, e.g., caused by temporal discontinuities and sudden changes, often occur around true action boundaries. This motivates us to introduce a Contrast-based Localization EvaluAtioN Network (CleanNet), whose core is a new temporal action proposal evaluator, which provides fine-grained pseudo supervision by leveraging the temporal contrasts among snippet-level classification predictions. As a result, the uncertainty in locating action instances can be resolved via evaluating their temporal contrast scores. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Besides, we also explore the usage of temporal contrast on temporal action proposal (TAP) generation task, which we believe is the first attempt with the weak supervision setting. Experiments on the THUMOS14, ActivityNet v1.2 and v1.3 datasets validate the efficacy of our method against existing state-of-the-art WS-TAL algorithms. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Action Coherence Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised Temporal Action Localization (W-TAL) aims at simultaneously classifying and locating all action instances with only video-level supervision. However, current W-TAL methods have two limitations. First, they ignore the difference in video representations between an action instance and its surrounding background when generating and scoring action proposals. Second, the unique characteristics of the RGB frames and optical flow are largely ignored when fusing these two modalities. To address these problems, an Action Coherence Network (ACN) is proposed in this paper. Its core is a new coherence loss which exploits both classification predictions and video content representations to supervise action boundary regression and thus leads to more accurate action localization results. Besides, the proposed ACN explicitly takes into account the specific characteristics of RGB frames and optical flow by training two separate sub-networks, each of which is able to generate modality-specific action proposals independently. Finally, to take advantage of the complementary action proposals generated by two streams, a novel fusion module is introduced to reconcile them and obtain the final action localization results. Experiments on the THUMOS14 and ActivityNet datasets show that our ACN outperforms the state-of-the-art W-TAL methods, and is even comparable to some recent fully-supervised methods. Particularly, ACN achieves a mean average precision of 26.4% on the THUMOS14 dataset under the IoU threshold 0.5. Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action LocalizationabstractThe object of Weakly-supervised Temporal Action Localization (WS-TAL) is to localize all action instances in an untrimmed video with only video-level supervision. Due to the lack of frame-level annotations during training, current WS-TAL methods rely on attention mechanisms to localize the foreground snippets or frames that contribute to the video-level classification task. This strategy frequently confuse context with the actual action, in the localization result. Separating action and context is a core problem for precise WS-TAL, but it is very challenging and has been largely ignored in the literature. In this paper, we introduce an Action-Context Separation Network (ACSNet) that explicitly takes into account context for accurate action localization. It consists of two branches (i.e., the Foreground-Background branch and the Action-Context branch). The Foreground-Background branch first distinguishes foreground from background within the entire video while the Action-Context branch further separates the foreground as action and context. We associate video snippets with two latent components (i.e., a positive component and a negative component), and their different combinations can effectively characterize foreground, action and context. Furthermore, we introduce extended labels with auxiliary context categories to facilitate the learning of action-context separation. Experiments on THUMOS14 and ActivityNet v1.2/v1.3 datasets demonstrate the ACSNet outperforms existing state-of-the-art WS-TAL methods by a large margin. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Wei Tang 0016, Junsong Yuan 0001, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 3 |
| 2021 | Practical Relative Order Attack in Deep RankingabstractRecent studies unveil the vulnerabilities of deep ranking models, where an imperceptible perturbation can trigger dramatic changes in the ranking result. While previous attempts focus on manipulating absolute ranks of certain candidates, the possibility of adjusting their relative order remains under-explored. In this paper, we formulate a new adversarial attack against deep ranking systems, i.e., the Order Attack, which covertly alters the relative order among a selected set of candidates according to an attacker-specified permutation, with limited interference to other unrelated candidates. Specifically, it is formulated as a triplet-style loss imposing an inequality chain reflecting the specified permutation. However, direct optimization of such white-box objective is infeasible in a real-world attack scenario due to various black-box limitations. To cope with them, we propose a Short-range Ranking Correlation metric as a surrogate objective for black-box Order Attack to approximate the white-box method. The Order Attack is evaluated on the Fashion-MNIST and Stanford-Online-Products datasets under both white-box and black-box threat models. The black-box attack is also successfully implemented on a major e-commerce platform. Comprehensive experimental evaluations demonstrate the effectiveness of the proposed methods, revealing a new type of ranking model vulnerability. Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 4 |
| 2021 | Graph-based temporal action co-localization from an untrimmed video
Le Wang 0003, Changbo Zhai, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
Neurocomputing | 3 |
| 2021 | Giant Panda IdentificationabstractThe lack of automatic tools to identify giant panda makes it hard to keep track of and manage giant pandas in wildlife conservation missions. In this paper, we introduce a new Giant Panda Identification (GPID) task, which aims to identify each individual panda based on an image. Though related to the human re-identification and animal classification problem, GPID is extraordinarily challenging due to subtle visual differences between pandas and cluttered global information. In this paper, we propose a new benchmark dataset iPanda-50 for GPID. The iPanda-50 consists of 6, 874 images from 50 giant panda individuals, and is collected from panda streaming videos. We also introduce a new Feature-Fusion Network with Patch Detector (FFN-PD) for GPID. The proposed FFN-PD exploits the patch detector to detect discriminative local patches without using any part annotations or extra location sub-networks, and builds a hierarchical representation by fusing both global and local features to enhance the inter-layer patch feature interactions. Specifically, an attentional cross-channel pooling is embedded in the proposed FFN-PD to improve the identify-specific patch detectors. Experiments performed on the iPanda-50 datasets demonstrate the proposed FFN-PD significantly outperforms competing methods. Besides, experiments on other fine-grained recognition datasets (i.e., CUB-200-2011, Stanford Cars, and FGVC-Aircraft) demonstrate that the proposed FFN-PD outperforms existing state-of-the-art methods. Le Wang 0003, Rizhi Ding, Yuanhao Zhai 0001, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Object Cosegmentation in Noisy Videos With Multilevel HypergraphabstractWith the target of simultaneously segmenting semantically related videos to identify the common objects, video object cosegmentation has attracted the attention of researchers in recent years. Existing methods are primarily based on pair-wise relations between adjacent pixels and regions, which are susceptible to performance degradation from object entries/exists or occlusions. Specifically, we refer these video frames without the common objects present as the “empty” frames. In this paper, we propose a multilevel hypergraph-based full Video object CoSegmentation (VCS) method, which incorporates high-level semantics and low-level appearance/motion/saliency to construct the hyperedge among multiple spatially and temporally adjacent regions. Specifically, the high-level semantic model fuses multiple object proposals from each frame instead of relying on a single object proposal per frame. A hypergraph cut is subsequently utilized to calculate the object cosegmentation. Experiments on four video object segmentation/cosegmentation datasets against state-of-the-art methods with both objective and subjective results manifest the effectiveness of the proposed VCS method, including the SegTrack and VCoSeg datasets without “empty” frames, the XJTU-Stevens dataset with 3.7% “empty” frames, and the Noisy-ViCoSeg dataset proposed together with our method with 30.3% “empty” frames. Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | Ladder Loss for Coherent Visual-Semantic EmbeddingabstractFor visual-semantic embedding, the existing methods normally treat the relevance between queries and candidates in a bipolar way – relevant or irrelevant, and all “irrelevant” candidates are uniformly pushed away from the query by an equal margin in the embedding space, regardless of their various proximity to the query. This practice disregards relatively discriminative information and could lead to suboptimal ranking in the retrieval results and poorer user experience, especially in the long-tail query scenario where a matching candidate may not necessarily exist. In this paper, we introduce a continuous variable to model the relevance degree between queries and multiple candidates, and propose to learn a coherent embedding space, where candidates with higher relevance degrees are mapped closer to the query than those with lower relevance degrees. In particular, the new ladder loss is proposed by extending the triplet loss inequality to a more general inequality chain, which implements variable push-away margins according to respective relevance degrees. In addition, a proper Coherent Score metric is proposed to better measure the ranking results including those “irrelevant” candidates. Extensive experiments on multiple datasets validate the efficacy of our proposed method, which achieves significant improvement over existing state-of-the-art methods. Zhenxing Niu, Le Wang 0003, Zhanning Gao, Qilin Zhang 0004, Gang Hua 0001 |
AAAI | 5 |
| 2020 | Multi-label X-Ray Imagery Classification via Bottom-Up Attention and Meta Fusion
Benyi Hu, Chi Zhang 0020, Le Wang 0003, Qilin Zhang 0004, Yuehu Liu |
ACCV (6) | 4 |
| 2020 | Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization
Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Junsong Yuan 0001, Gang Hua 0001 |
ECCV (6) | 4 |
| 2020 | Adversarial Ranking Attack and Defense
Zhenxing Niu, Le Wang 0003, Qilin Zhang 0004, Gang Hua 0001 |
ECCV (14) | 4 |
| 2020 | Fine-Grained Giant Panda IdentificationabstractThe image-based fine-grained identification of individual giant pandas (Ailuropoda melanoleuca) is an emerging technology, and it is extraordinarily challenging due to the extremely subtle visual differences between individual giant pandas and limited annotated training data. To address these challenges, we propose the Feature-Fusion Convolutional Neural Network with Patch Detector (FFCNN-PD) algorithm, which exploits the discriminative local patches and builds a hierarchical representation generated by fusing both global and local features. Specifically, an attentional cross-channel pooling is embedded in the FFCNN-PD to improve the class- specific patch detectors. In addition, we propose a new giant panda identification dataset (iPanda-30) to establish a benchmark. Experiments on the proposed iPanda-30 dataset and other fine-grained recognition datasets demonstrate the effectiveness of the FFCNN-PD algorithm against the existing state-of-the-arts. Rizhi Ding, Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICASSP | 3 |
| 2020 | Action Co-localization in an Untrimmed Video by Graph Neural Networks
Changbo Zhai, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
MMM (1) | 3 |
| 2019 | Video Imprint Segmentation for Temporal Action Detection in Untrimmed VideosabstractWe propose a temporal action detection by spatial segmentation framework, which simultaneously categorize actions and temporally localize action instances in untrimmed videos. The core idea is the conversion of temporal detection task into a spatial semantic segmentation task. Firstly, the video imprint representation is employed to capture the spatial/temporal interdependences within/among frames and represent them as spatial proximity in a feature space. Subsequently, the obtained imprint representation is spatially segmented by a fully convolutional network. With such segmentation labels projected back to the video space, both temporal action boundary localization and per-frame spatial annotation can be obtained simultaneously. The proposed framework is robust to variable lengths of untrimmed videos, due to the underlying fixed-size imprint representations. The efficacy of the framework is validated in two public action detection datasets. Zhanning Gao, Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 3 |
| 2019 | Object Affordances Graph Network for Action Recognition
Haoliang Tan, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Nanning Zheng 0001, Gang Hua 0001 |
BMVC | 3 |
| 2019 | Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksabstractWeakly-supervised temporal action localization (WS-TAL) is a promising but challenging task with only video-level action categorical labels available during training. Without requiring temporal action boundary annotations in training data, WS-TAL could possibly exploit automatically retrieved video tags as video-level labels. However, such coarse video-level supervision inevitably incurs confusions, especially in untrimmed videos containing multiple action instances. To address this challenge, we propose the Contrast-based Localization EvaluAtioN Network (CleanNet) with our new action proposal evaluator, which provides pseudo-supervision by leveraging the temporal contrast in snippet-level action classification predictions. Essentially, the new action proposal evaluator enforces an additional temporal contrast constraint so that high-evaluation-score action proposals are more likely to coincide with true action instances. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Experiments on THUMOS14 and ActivityNet datasets validate the efficacy of CleanNet against existing state-ofthe- art WS-TAL algorithms. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 3 |
| 2019 | Action Coherence Network for Weakly Supervised Temporal Action LocalizationabstractMost prominent temporal action localization methods are of the fully-supervised type, which rely heavily on frame-level labels, which could be prohibitively expensive to annotate. Thanks to recent developments on the Weakly-supervised Temporal Action Localization (W-TAL), this alternative paradigm requires only video-level labels in training, alleviating such annotation efforts. Specifically, we present Action Coherence Network (ACN) for W-TAL, which features a new coherence loss that better supervises action boundary learning and facilitate proposal regression. In addition, a purpose-built fusion module is proposed for localization inference based on features extracted by two streams of convolutional neural network. Overall, the proposed ACN achieves state-of-the-art W-TAL performance on two challenging datasets (THU-MOS14 and ActivityNet1.2, particularly ACN attains mAP of 24.2% on THUMOS14 under IoU threshold 0.5), which is approaching some recent fully-supervised TAL methods. Yuanhao Zhai 0001, Le Wang 0003, Ziyi Liu 0001, Qilin Zhang 0004, Gang Hua 0001, Nanning Zheng 0001 |
ICIP | 4 |
| 2018 | Video-Based Sign Language Recognition Without Temporal SegmentationabstractMillions of hearing impaired people around the world routinely use some variants of sign languages to communicate, thus the automatic translation of a sign language is meaningful and important. Currently, there are two sub-problems in Sign Language Recognition (SLR), i.e., isolated SLR that recognizes word by word and continuous SLR that translates entire sentences. Existing continuous SLR methods typically utilize isolated SLRs as building blocks, with an extra layer of preprocessing (temporal segmentation) and another layer of post-processing (sentence synthesis). Unfortunately, temporal segmentation itself is non-trivial and inevitably propagates errors into subsequent steps. Worse still, isolated SLR methods typically require strenuous labeling of each word separately in a sentence, severely limiting the amount of attainable training data. To address these challenges, we propose a novel continuous sign recognition framework, the Hierarchical Attention Network with Latent Space (LS-HAN), which eliminates the preprocessing of temporal segmentation. The proposed LS-HAN consists of three components: a two-stream Convolutional Neural Network (CNN) for video feature representation generation, a Latent Space (LS) for semantic gap bridging, and a Hierarchical Attention Network (HAN) for latent space based recognition. Experiments are carried out on two large scale datasets. Experimental results demonstrate the effectiveness of the proposed framework. Jie Huang 0011, Wengang Zhou 0001, Qilin Zhang 0004, Houqiang Li, Weiping Li 0003 |
AAAI | 3 |
| 2018 | Joint Spatio-Temporal Action Localization in Untrimmed Videos with Per-Frame SegmentationabstractInspired by the recent spatio-temporal action localization efforts with tubelets (sequences of bounding boxes), we present a new spatio-temporal action detector Segment-tube, which consists of sequences of per-frame segmentation masks. The proposed Segment-tube detector can temporally pinpoint the starting/ending frame of each action class in the presence of preceding/subsequent interference actions in untrimmed videos. Simultaneously, the Segment-tube detector produces per-frame segmentation masks instead of bounding boxes, offering superior spatial accuracy to tubelets. This is achieved by alternating iterative optimization between temporal action localization and spatial action segmentation. Experimental results on multiple datasets validate the efficacy of the proposed detector. Xuhuan Duan, Le Wang 0003, Changbo Zhai, Nanning Zheng 0001, Qilin Zhang 0004, Zhenxing Niu, Gang Hua 0001 |
ICIP | 5 |
| 2018 | Video Object Co-Segmentation from Noisy Videos by a Multi-Level Hypergraph ModelabstractDefined as simultaneously segmenting a set of related videos to identify the common objects, video co-segmentation has attracted the attention of researchers in recent years. Existing methods are primarily based on pair-wise relations between adjacent pixels/regions, which are susceptible to performance degradation from “empty” video frames (e.g., due to transient/intermittent common objects). In this paper, a new multilevel hypergraph based method, termed the full Video object Co-Segmentation method (VCS), is proposed, which incorporates both a high-level semantics object model and a low-level appearance/motion/saliency object model to construct the hyperedge among multiple spatially and temporally adjacent regions. Specifically, the high-level semantic model fuses multiple object proposals from each frame instead of relying on a single object proposal per frame. A hypergraph cut is subsequently utilized to calculate the object co-segmentation. Experiments on three datasets demonstrate the efficacy of the proposed VCS method. Le Wang 0003, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
ICIP | 3 |
| 2018 | Traffic Sensory Data Classification by Quantifying Scenario ComplexityabstractFor unmanned ground vehicle (UGV) off-line testing and performance evaluation, massive amount of traffic scenario data is often required. The annotations in current off-line traffic sensory dataset typically include I) types of roadways II) scene types III) specific characteristics that are generally considered challenging for cognitive algorithms. While such annotations are helpful in manual selection of data, they are insufficient for comprehensive and quantitate measurement of per-roadway-segment scenario complexity. To resolve such limitations, we propose a traffic sensory data classification paradigm based on quantifying the scenario complexity for each roadway segment, where such quantification is jointly based on road semantic complexity and traffic element complexity. The road semantic complexity is a proposed measurement of the complexity incurred by the static elements such as curvy roads, intersections, merges and splits, which is predicted with a Support Vector Regression (SVR). The traffic element complexity is a measurement of complexity due to dynamic traffic elements, such as nearby vehicles and pedestrians. Experimental results and a case study verify the efficacy of the proposed method. Chi Zhang 0020, Yuehu Liu, Qilin Zhang 0004 |
Intelligent Vehicles Symposium | 4 |
| 2018 | Multi-model Traffic Scene Simulation with Road Image Sequences and GIS InformationabstractIn this paper, a new multi-modal traffic scene simulation framework with combined inputs of road image sequences and road information from Geographic Information Systems (GIS) is proposed. The proposed framework contains two major steps, with the first one being a preprocessing step, including 3D road model extraction, camera location and orientation estimation and lane extraction from both GIS and road image sequences. After such preprocessing, the traffic scene reconstruction is reformulated into a 6-degree of freedom (6DoF) pose estimation in the 3D road model. Subsequently, the iterative closest point (ICP) algorithm is exploited for coarse point registration by estimating the pose in the road model. In addition, an objective function is established to incorporate the image features (e.g., lanes) into the road model and to refine the pose estimation. In the experiments with the publicly available KITTI dataset, the proposed method achieves high average Intersection-over-Union (IoU) scores as compared to the ground truth image sequences. Zhichao Cui, Yuehu Liu, Fuji Ren, Qilin Zhang 0004 |
Intelligent Vehicles Symposium | 4 |
| 2018 | A Graded Offline Evaluation Framework for Intelligent Vehicle's Cognitive AbilityabstractCognitive ability evaluation in intelligent vehicles is conventionally evaluated by classical autonomous driving dataset, which lacks comprehensive annotations of driving difficulty. Realistically, different driving conditions require vast different level of cognitive ability, e.g., driving in highly congested traffic is much more challenging than driving on limited access highway; driving in a blizzard/hurricane requires much more robust environmental cognition abilities than driving under ordinary conditions. Different datasets contain different proportions of various driving conditions, rendering intelligent vehicle evaluation susceptible to dataset variations. To overcome such limitations, we propose to first benchmark the driving difficulty with the proposed “Cascaded Tanks Model” and obtain a fine-grained per-segment difficulty rating based on our proposed Semantic Descriptor. With the proposed Graded Offline Evaluation (GOE) framework, it is demonstrated that offline validation of the cognitive abilities in Intelligent Vehicles (IV) is more consistent regardless of dataset choice. Chi Zhang 0020, Yuehu Liu, Qilin Zhang 0004, Le Wang 0003 |
Intelligent Vehicles Symposium | 3 |
| 2018 | Convolutional Neural Networks with Generalized Attentional Pooling for Action RecognitionabstractInspired by the recent advance in attentional pooling techniques in image classification and action recognition tasks, we propose the Generalized Attentional Pooling (GAP) based Convolutional Neural Network (CNN) algorithm for action recognition in still images. The proposed GAP-CNN can be formulated as a new approximation of the second-order/bilinear pooling techniques widely used in fine-grained image classification. Unlike the existing rank-1 approximation, a generalized factoring (with non-linear functions) is introduced to exploit the intrinsic structural information of the sample covariance matrices of convolutional layer outputs. Without requiring preprocessing steps such as object (e.g., human body) bounding boxes detection, the proposed GAP-CNN automatically focuses on the most informative part in still images. With the additional guidance of keypoints of human pose, the proposed GAP-CNN algorithm achieves the state-of-the-art action recognition accuracy on the large-scale MPII still image dataset. Wengang Zhou 0001, Qilin Zhang 0004, Houqiang Li |
VCIP | 3 |
| 2018 | Joint Video Object Discovery and Segmentation by Coupled Dynamic Markov NetworksabstractIt is a challenging task to extract segmentation mask of a target from a single noisy video, which involves object discovery coupled with segmentation. To solve this challenge, we present a method to jointly discover and segment an object from a noisy video, where the target disappears intermittently throughout the video. Previous methods either only fulfill video object discovery, or video object segmentation presuming the existence of the object in each frame. We argue that jointly conducting the two tasks in a unified way will be beneficial. In other words, video object discovery and video object segmentation tasks can facilitate each other. To validate this hypothesis, we propose a principled probabilistic model, where two dynamic Markov networks are coupled-one for discovery and the other for segmentation. When conducting the Bayesian inference on this model using belief propagation, the bi-directional message passing reveals a clear collaboration between these two inference tasks. We validated our proposed method in five data sets. The first three video data sets, i.e., the SegTrack data set, the YouTube-objects data set, and the Davis data set, are not noisy, where all video frames contain the objects. The two noisy data sets, i.e., the XJTU-Stevens data set, and the Noisy-ViDiSeg data set, newly introduced in this paper, both have many frames that do not contain the objects. When compared with state of the art, it is shown that although our method produces inferior results on video data sets without noisy frames, we are able to obtain better results on video data sets with noisy frames. Ziyi Liu 0001, Le Wang 0003, Gang Hua 0001, Qilin Zhang 0004, Zhenxing Niu, Ying Wu 0001, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 4 |
| 2015 | Multi-View Visual Recognition of Imperfect Testing DataabstractA practical yet under-explored problem often encountered by multimedia researchers is the recognition of imperfect testing data, where multiple sensing channels are deployed but interference or transmission distortion corrupts some of them. Typical cases of imperfect testing data include missing features and feature misalignments. To address these challenges, we choose the latent space model and introduce a new similarity learning canonical-correlation analysis (SLCCA) method to capture the semantic consensus between views. The consensus information is preserved by projection matrices learned with modified canonical-correlation analysis (CCA) optimization terms with new, explicit class-similarity constraints. To make it computationally tractable, we propose to combine a practical relaxation and an alternating scheme to solve the optimization problem. Experiments on four challenging multi-view visual recognition datasets demonstrate the efficacy of the proposed method. Qilin Zhang 0004, Gang Hua 0001 |
ACM Multimedia | 1 |
| 2014 | Can Visual Recognition Benefit from Auxiliary Information in Training?
Qilin Zhang 0004, Gang Hua 0001, Wei Liu 0005, Zicheng Liu 0001, Zhengyou Zhang |
ACCV (1) | 1 |