Zhiwen Fang

dblp:83/10447 · DBLP profile ↗
← Back
33ranked-venue papers
8as first author
18since 2021 · last 2025
0000-0002-9314-5262ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Security and privacy · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2025 Multilingual-Prompt-Guided Directional Feature Learning for Weakly Supervised Video Anomaly Detection
abstract
Weakly supervised video anomaly detection has gained attention for its effective performance and cost-efficient annotation, using video-level labels to distinguish between normal and abnormal patterns. However, challenges arise from the diversity and incompleteness of anomalous events, complicating feature learning. Vision-language models offer promising approaches, but designing precise prompts remains difficult. This is because accommodating the diverse range of normal and anomalous scenarios in real-world settings is challenging, and the workload is significant. To tackle these issues, we propose integrating multilingualism and multiple prompts to improve feature learning. By utilizing prompts in various languages to define "anomaly" and "normalcy," we tackle these concepts across different linguistic domains. In each domain, multiple prompts are employed for adaptive top-K prompt selection of snippets. To enhance visual feature learning, a multi-granularity attention module combining Transformer and Mamba is designed. Mamba's long-range adaptation selection builds fine-grained temporal correlations among coarse-grained snippets, while Transformer enhances fine-grained information guided by coarse-grained information. Alongside a multilingual prompt guidance loss, we introduce a gradual directional loss to jointly optimize visual feature distribution and the top-K prompt selection. Our method demonstrates effectiveness on four video datasets and provides generalizability analyses on two medical datasets, including EMG and ECG temporal data.
Chizhuo Xiao, Yang Xiao 0007, Joey Tianyi Zhou, Zhiwen Fang
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Rhythmer: Ranking-Based Skill Assessment With Rhythm-Aware Transformer
abstract
Ranking-based skill assessment is an essential component of video understanding. In this task lacking precise procedure annotations, existing methods place greater emphasis on evaluating the procedure quality via manually normalizing the execution duration. However, the inherent duration-related procedural patterns will undergo alteration. Experimentally, we discover that distinct duration biases are prevalent in duration-sensitive skills, such as those in medical and everyday life. Hence, duration information is crucial for ranking-based skill assessment when dealing with varying durations. Additionally, similar execution processes tend to have closer execution durations. Thus, another critical factor lies in extracting duration-related procedural information alongside similar durations. It is defined as mining rhythm patterns, which are inspired by music rhythms including various duration and duration-related procedures. In our work, a rhythm-aware transformer is proposed to mine the rhythm patterns adaptively. Given pairwise inputs, a co-attention module is designed to mutually highlight duration-related procedure information when comparing pairwise input videos with similar durations, and adaptively attenuate the efficacy when confronted with pairwise inputs featuring significantly different durations. A rhythm-encoding module further embeds duration information into the concatenation of raw features and co-attention features. Following these features, the transformer decoder is designed to learn duration-related queries supervised by a novel duration grouping loss among various duration groups. The experimental results demonstrate that the rhythm-aware transformer is effective for ranking-based skill assessment.
Zhuang Luo, Yang Xiao 0007, Feng Yang 0012, Joey Tianyi Zhou, Zhiwen Fang
IEEE Trans. Circuits Syst. Video Technol.5
2024 CrossGLG: LLM Guides One-Shot Skeleton-Based 3D Action Recognition in a Cross-Level Manner
Tingbing Yan, Wenzheng Zeng, Yang Xiao 0007, Xingyu Tong 0002, Zhiwen Fang, Zhiguo Cao 0001, Joey Tianyi Zhou
ECCV (20)6
2024 C2Net: content-dependent and -independent cross-attention network for anomaly detection in videos
Jiafei Liang, Yang Xiao 0007, Joey Tianyi Zhou, Feng Yang 0012, Zhiwen Fang
Appl. Intell.6
2024 TaiChiNet: Negative-Positive Cross-Attention Network for Breast Lesion Segmentation in Ultrasound Images
abstract
Breast lesion segmentation in ultrasound images is essential for computer-aided breast-cancer diagnosis. To improve the segmentation performance, most approaches design sophisticated deep-learning models by mining the patterns of foreground lesions and normal backgrounds simultaneously or by unilaterally enhancing foreground lesions via various focal losses. However, the potential of normal backgrounds is underutilized, which could reduce false positives by compacting the feature representation of all normal backgrounds. From a novel viewpoint of bilateral enhancement, we propose a negative-positive cross-attention network to concentrate on normal backgrounds and foreground lesions, respectively. Derived from the complementing opposites of bipolarity in TaiChi, the network is denoted as TaiChiNet, which consists of the negative normal-background and positive foreground-lesion paths. To transmit the information across the two paths, a cross-attention module, a complementary MLP-head, and a complementary loss are built for deep-layer features, shallow-layer features, and mutual-learning supervision, separately. To the best of our knowledge, this is the first work to formulate breast lesion segmentation as a mutual supervision task from the foreground-lesion and normal-background views. Experimental results have demonstrated the effectiveness of TaiChiNet on two breast lesion segmentation datasets with a lightweight architecture. Furthermore, extensive experiments on the thyroid nodule segmentation and retinal optic cup/disc segmentation datasets indicate the application potential of TaiChiNet.
Jinting Wang, Jiafei Liang, Yang Xiao 0007, Joey Tianyi Zhou, Zhiwen Fang, Feng Yang 0012
IEEE J. Biomed. Health Informatics5
2024 Multi-Template Meta-Information Regularized Network for Alzheimer's Disease Diagnosis Using Structural MRI
abstract
Structural magnetic resonance imaging (sMRI) has been widely applied in computer-aided Alzheimer's disease (AD) diagnosis, owing to its capabilities in providing detailed brain morphometric patterns and anatomical features in vivo. Although previous works have validated the effectiveness of incorporating metadata (e.g., age, gender, and educational years) for sMRI-based AD diagnosis, existing methods solely paid attention to metadata-associated correlation to AD (e.g., gender bias in AD prevalence) or confounding effects (e.g., the issue of normal aging and metadata-related heterogeneity). Hence, it is difficult to fully excavate the influence of metadata on AD diagnosis. To address these issues, we constructed a novel Multi-template Meta-information Regularized Network (MMRN) for AD diagnosis. Specifically, considering diagnostic variation resulting from different spatial transformations onto different brain templates, we first regarded different transformations as data augmentation for self-supervised learning after template selection. Since the confounding effects may arise from excessive attention to meta-information owing to its correlation with AD, we then designed the modules of weakly supervised meta-information learning and mutual information minimization to learn and disentangle meta-information from learned class-related representations, which accounts for meta-information regularization for disease diagnosis. We have evaluated our proposed MMRN on two public multi-center cohorts, including the Alzheimer's Disease Neuroimaging Initiative (ADNI) with 1,950 subjects and the National Alzheimer's Coordinating Center (NACC) with 1,163 subjects. The experimental results have shown that our proposed method outperformed the state-of-the-art approaches in both tasks of AD diagnosis, mild cognitive impairment (MCI) conversion prediction, and normal control (NC) vs. MCI vs. AD classification.
Kangfu Han, Gang Li 0001, Zhiwen Fang, Feng Yang 0012
IEEE Trans. Medical Imaging3
2024 GREnet: Gradually REcurrent Network With Curriculum Learning for 2-D Medical Image Segmentation
abstract
Medical image segmentation is a vital stage in medical image analysis. Numerous deep-learning methods are booming to improve the performance of 2-D medical image segmentation, owing to the fast growth of the convolutional neural network. Generally, the manually defined ground truth is utilized directly to supervise models in the training phase. However, direct supervision of the ground truth often results in ambiguity and distractors as complex challenges appear simultaneously. To alleviate this issue, we propose a gradually recurrent network with curriculum learning, which is supervised by gradual information of the ground truth. The whole model is composed of two independent networks. One is the segmentation network denoted as GREnet, which formulates 2-D medical image segmentation as a temporal task supervised by pixel-level gradual curricula in the training phase. The other is a curriculum-mining network. To a certain degree, the curriculum-mining network provides curricula with an increasing difficulty in the ground truth of the training set by progressively uncovering hard-to-segmentation pixels via a data-driven manner. Given that segmentation is a pixel-level dense-prediction challenge, to the best of our knowledge, this is the first work to function 2-D medical image segmentation as a temporal task with pixel-level curriculum learning. In GREnet, the naive UNet is adopted as the backbone, while ConvLSTM is used to establish the temporal link between gradual curricula. In the curriculum-mining network, UNet++ supplemented by transformer is designed to deliver curricula through the outputs of the modified UNet++ at different layers. Experimental results have demonstrated the effectiveness of GREnet on seven datasets, i.e., three lesion segmentation datasets in dermoscopic images, an optic disc and cup segmentation dataset and a blood vessel segmentation dataset in retinal images, a breast lesion segmentation dataset in ultrasound images, and a lung segmentation dataset in computed tomography (CT).
Jinting Wang, Yujiao Tang, Yang Xiao 0007, Joey Tianyi Zhou, Zhiwen Fang, Feng Yang 0012
IEEE Trans. Neural Networks Learn. Syst.5
2023 Real-time Multi-person Eyeblink Detection in the Wild for Untrimmed Video
abstract
Real-time eyeblink detection in the wild can widely serve for fatigue detection, face anti-spoofing, emotion analysis, etc. The existing research efforts generally focus on single-person cases towards trimmed video. However, multi-person scenario within untrimmed videos is also important for practical applications, which has not been well concerned yet. To address this, we shed light on this research field for the first time with essential contributions on dataset, theory, and practices. In particular, a large-scale dataset termed MPEblink that involves 686 untrimmed videos with 8748 eyeblink events is proposed under multi-person conditions. The samples are captured from uncon-strainedfilms to reveal “in the wild“ characteristics. Meanwhile, a real-time multi-person eyeblink detection method is also proposed. Being different from the existing counter-parts, our proposition runs in a one-stage spatio-temporal way with end-to-end learning capacity. Specifically, it simultaneously addresses the sub-tasks of face detection, face tracking, and human instance-level eyeblink detection. This paradigm holds 2 main advantages: (1) eyeblink features can be facilitated via the face's global context (e.g., head pose and illumination condition) with joint optimization and interaction, and (2) addressing these sub-tasks in parallel instead of sequential manner can save time remarkably to meet the real-time running requirement. Experiments on MPEblink verify the essential challenges of real-time multi-person eyeblink detection in the wild for untrimmed video. Our method also outperforms existing approaches by large margins and with a high inference speed.
Wenzheng Zeng, Yang Xiao 0007, Sicheng Wei, Jinfang Gan, Xintao Zhang, Zhiguo Cao 0001, Zhiwen Fang, Joey Tianyi Zhou
CVPR7
2023 Multi-spectral template matching based object detection in a few-shot learning manner
Chen Feng 0002, Zhiguo Cao 0001, Yang Xiao 0007, Zhiwen Fang, Joey Tianyi Zhou
Inf. Sci.4
2023 Eyelid's Intrinsic Motion-Aware Feature Learning for Real-Time Eyeblink Detection in the Wild
abstract
Real-time eyeblink detection in the wild is a recently emerged challenging task that suffers from dramatic variations in face attribute, pose, illumination, camera view and distance, etc. One key issue is to well characterize eyelid’s intrinsic motion (i.e., approaching and departure between upper and lower eyelid) robustly, under unconstrained conditions. Towards this, a novel eyelid’s intrinsic motion-aware feature learning approach is proposed. Our proposition lies in 3 folds. First, the feature extractor is led to focus on informative eye region adaptively via introducing visual attention in a coarse-to-fine way, to guarantee robustness and fine-grained descriptive ability jointly. Then, 2 constraints are proposed to make feature learning be aware of eyelid’s intrinsic motion. Particularly, one concerns the fact that the inter-frame feature divergence within eyeblink processes should be greater than non-eyeblink ones to better reveal eyelid’s intrinsic motion. The other constraint minimizes the inter-frame feature divergence of non-eyeblink samples, to suppress motion clues due to head or camera movement, illumination change, etc. Meanwhile, concerning the high ambiguity between eyeblink and non-eyeblink samples, soft sample labels are acquired via self-knowledge distillation to conduct feature learning with finer supervision than the hard ones. The experiments verify that, our proposition is significantly superior to the state-of-the-art ones (i.e., advantage on F1-score over 7%) and with real-time running efficiency. It is also of strong generalization capacity towards constrained conditions. The source code is available athttps://github.com/wenzhengzeng/blink_eyelid.
Wenzheng Zeng, Yang Xiao 0007, Guilei Hu, Zhiguo Cao 0001, Sicheng Wei, Zhiwen Fang, Joey Tianyi Zhou, Junsong Yuan 0001
IEEE Trans. Inf. Forensics Secur.6
2022 Locality-Aware Crowd Counting
abstract
Imbalanced data distribution in crowd counting datasets leads to severe under-estimation and over-estimation problems, which has been less investigated in existing works. In this paper, we tackle this challenging problem by proposing a simple but effective locality-based learning paradigm to produce generalizable features by alleviating sample bias. Our proposed method is locality-aware in two aspects. First, we introduce a locality-aware data partition (LADP) approach to group the training data into different bins via locality-sensitive hashing. As a result, a more balanced data batch is then constructed by LADP. To further reduce the training bias and enhance the collaboration with LADP, a new data augmentation method called locality-aware data augmentation (LADA) is proposed where the image patches are adaptively augmented based on the loss. The proposed method is independent of the backbone network architectures, and thus could be smoothly integrated with most existing deep crowd counting approaches in an end-to-end paradigm to boost their performance. We also demonstrate the versatility of the proposed method by applying it for adversarial defense. Extensive experiments verify the superiority of the proposed method over the state of the arts.
Joey Tianyi Zhou, Le Zhang 0001, Jiawei Du 0002, Xi Peng 0001, Zhiwen Fang, Hongyuan Zhu 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 MAT: Multianchor Visual Tracking With Selective Search Region
abstract
The core prerequisite of most modern trackers is a motion assumption, defined as predicting the current location in a limited search region centering at the previous prediction. For clarity, the central subregion of a search region is denoted as the tracking anchor (e.g., the location of the previous prediction in the current frame). However, providing accurate predictions in all frames is very challenging in the complex nature scenes. In addition, the target locations in consecutive frames often change violently under the attribute of fast motion. Both facts are likely to lead the previous prediction to an unbelievable tracking anchor, which will make the aforementioned prerequisite invalid and cause tracking drift. To enhance the reliability of tracking anchors, we propose a real-time multianchor visual tracking mechanism, called multianchor tracking (MAT). Instead of directly relying on the tracking anchor inherited from the previous prediction, MAT selects the best anchor from an anchor ensemble, which includes several objectness-based anchor proposals and the anchor inherited from the previous prediction. The objectness-based anchors provide several complementary selective search regions, and an entropy-minimization-based selection method is introduced to find the best anchor. Our approach offers two benefits: 1) selective search regions can increase the chance of tracking success with affordable computational load and 2) anchor selection introduces the best anchor for each frame, which breaks the limitation of solo depending on the previous prediction. The extensive experiments of nine base trackers upgraded by MAT on four challenging datasets demonstrate the effectiveness of MAT.
Zhiwen Fang, Zhiguo Cao 0001, Yang Xiao 0007, Kaicheng Gong, Junsong Yuan 0001
IEEE Trans. Cybern.1
2022 Person Re-Identification With Hierarchical Discriminative Spatial Aggregation
abstract
Practically, person re-identification (re-ID) may suffer from the critical spatial misalignment problem due to inaccurate human detection, variation on human pose and camera viewpoint, etc. To address this, a hierarchical discriminative spatial aggregation method is proposed. The key idea is to conduct spatial aggregation on local human parts via global average-pooling to acquire the strong spatial misalignment tolerance, with VALD encoding on the local parts for facilitating discriminative power jointly. This proposition is built on NetVLAD to ensure end-to-end deep learning capacity. Due to the fine-grained property of person re-ID task that has not been well concerned by the original NetVLAD model for scene recognition, a feature refinement layer that consists of 1 fully-connected (FC) layer and 2 batch normalization (BN) layers is added on top of the raw NetVLAD layer to enhance the discriminative power and training convergence. And, a human body occlusion and background component dropout manner is also proposed to resist the effect of serious occlusion. Technically, a refined codeword initialization manner is proposed to alleviate the potential codeword imbalance problem caused by naive random initialization. The proposed discriminative spatial aggregation approach is then conducted on multi-resolution convolutional feature map layers hierarchically via early feature fusion, to involve richer semantic and fine-grained visual clues jointly. Wide-range experiments on 6 datasets (i.e., CUHK03, DukeMTMC-reID, Occluded-DukeMTMC, Market-1501, MSMT17 and Occluded-REID) verifies the effectiveness of our proposition. The source code and supporting material is available athttps://github.com/zmyme/HDSA-reID.
Yang Xiao 0007, Fu Xiong, Zhiguo Cao 0001, Zhiwen Fang, Joey Tianyi Zhou
IEEE Trans. Inf. Forensics Secur.6
2022 Anomaly Detection With Bidirectional Consistency in Videos
abstract
The core component of most anomaly detectors is a self-supervised model, tasked with modeling patterns included in training samples and detecting unexpected patterns as the anomalies in testing samples. To cope with normal patterns, this model is typically trained with reconstruction constraints. However, the model has the risk of overfitting to training samples and being sensitive to hard normal patterns in the inference phase, which results in irregular responses at normal frames. To address this problem, we formulate anomaly detection as a mutual supervision problem. Due to collaborative training, the complementary information of mutual learning can alleviate the aforementioned problem. Based on this motivation, a SIamese generative network (SIGnet), including two subnetworks with the same architecture, is proposed to simultaneously model the patterns of the forward and backward frames. During training, in addition to traditional constraints on improving the reconstruction performance, a bidirectional consistency loss based on the forward and backward views is designed as the regularization term to improve the generalization ability of the model. Moreover, we introduce a consistency-based evaluation criterion to achieve stable scores at the normal frames, which will benefit detecting anomalies with fluctuant scores in the inference phase. The results on several challenging benchmark data sets demonstrate the effectiveness of our proposed method.
Zhiwen Fang, Jiafei Liang, Joey Tianyi Zhou, Yang Xiao 0007, Feng Yang 0012
IEEE Trans. Neural Networks Learn. Syst.1
2021 LPQ++: A discriminative blur-insensitive textural descriptor with spatial-channel interaction
Yang Xiao 0007, Zhiguo Cao 0001, Zhiwen Fang, Joey Tianyi Zhou
Inf. Sci.5
2021 Single-Image Dehazing via Compositional Adversarial Network
abstract
Single-image dehazing has been an important topic given the commonly occurred image degradation caused by adverse atmosphere aerosols. The key to haze removal relies on an accurate estimation of global air-light and the transmission map. Most existing methods estimate these two parameters using separate pipelines which reduces the efficiency and accumulates errors, thus leading to a suboptimal approximation, hurting the model interpretability, and degrading the performance. To address these issues, this article introduces a novel generative adversarial network (GAN) for single-image dehazing. The network consists of a novel compositional generator and a novel deeply supervised discriminator. The compositional generator is a densely connected network, which combines fine-scale and coarse-scale information. Benefiting from the new generator, our method can directly learn the physical parameters from data and recover clean images from hazy ones in an end-to-end manner. The proposed discriminator is deeply supervised, which enforces that the output of the generator to look similar to the clean images from low-level details to high-level structures. To the best of our knowledge, this is the first end-to-end generative adversarial model for image dehazing, which simultaneously outputs clean images, transmission maps, and air-lights. Extensive experiments show that our method remarkably outperforms the state-of-the-art methods. Furthermore, to facilitate future research, we create the HazeCOCO dataset which is currently the largest dataset for single-image dehazing.
Hongyuan Zhu 0002, Xi Peng 0001, Joey Tianyi Zhou, Zhao Kang 0001, Shijian Lu, Zhiwen Fang, Liyuan Li, Joo-Hwee Lim
IEEE Trans. Cybern.7
2021 Multi-Encoder Towards Effective Anomaly Detection in Videos
abstract
Given normal training samples, anomaly detection in videos can be regarded as a challenging problem of identifying unexpected events. The state-of-the-art approaches generally resort to the autoencoder model by using a single encoder to capture the motion and content patterns jointly. Nevertheless, due to the lack of accurate labels of normal and abnormal samples, how to detect anomalies is decided by the subjective understanding of models. It infers that different models will prefer to mine different patterns according to the characteristics of models. We call this problem as a pattern bias problem. To alleviate this problem, a novel Multi-Encoder Single-Decoder network, termed as MESDnet, is proposed in the spirit of encoding motion and content cues individually with multiple encoders. MESDnet is of end-to-end learning ability and real-time running speed. Particularly, the differences between adjacent frames and the raw frames are used as the motion and content sources, respectively. Then, a decoder takes charge of detecting anomalies in the way of observing reconstructing error towards the video frames by using the multi-stream encoded motion and content features simultaneously. The experiments on the CUHK Avenue dataset, the UCSD Pedestrian dataset, and the ShanghaiTech Campus dataset verify the effectiveness of MESDnet.
Zhiwen Fang, Joey Tianyi Zhou, Yang Xiao 0007, Yanan Li 0006, Feng Yang 0012
IEEE Trans. Multim.1
2021 Abrupt-motion-aware lightweight visual tracking for unmanned aerial vehicles
Kaicheng Gong, Zhiguo Cao 0001, Yang Xiao 0007, Zhiwen Fang
Vis. Comput.4
2020 Image Denoising for Efficient Anomaly Detection in Videos
abstract
Video anomaly detection is tasked with the identification of events that do not conform to expected events. Currently, most methods tackle this problem by mining common normal patterns from training data and minimizing the generative errors. In inference phase, a large generative error is assigned to an abnormal event and a small one is for a normal event. However, because these methods only focus on the error intensity but ignore the error pattern, partial abnormal events will own similar generative error intensities to the normal ones. Thus, we propose to tackle the anomaly detection within an efficient image denoising framework. In this framework, the generative errors are treated as a kind of artificial noise, which will be superimposed on the current frame. Then, the contaminated frame is fed into a denoising network, which is trained to output a frame close to the current frame. In the denoising network, the common patterns of training data and the error patterns of each training frame can be learned jointly. It will benefit anomaly detection by restraining the generative errors of normal frames. The results on several challenging benchmark datasets demonstrate the effectiveness of our proposed method.
Zhiwen Fang, Zhou Yue, Weiyuan Liu, Feng Yang 0012
ICDCS1
2020 Attention-Driven Loss for Anomaly Detection in Video Surveillance
abstract
Recent video anomaly detection methods focus on reconstructing or predicting frames. Under this umbrella, the long-standing inter-class data-imbalance problem resorts to the imbalance between foreground and stationary background objects in video anomaly detection and this has been less investigated by existing solutions. Naively optimizing the reconstructing loss yields a biased optimization towards background reconstruction rather than the objects of interest in the foreground. To solve this, we proposed a simple yet effective solution, termed attention-driven loss to alleviate the foreground-background imbalance problem in anomaly detection. Specifically, we compute a single mask map that summarizes the frame evolution of moving foreground regions and suppresses the background in the training video clips. After that, we construct an attention map through the combination of the mask map and background to give different weights to the foreground and background region respectively. The proposed attention-driven loss is independent of backbone networks and can be easily augmented in most existing anomaly detection models. Augmented with attention-driven loss, the model is able to achieve AUC 86.0% on Avenue, 83.9% on Ped1, 96% on Ped2 datasets. Extensive experimental results and ablation studies further validate the effectiveness of our model.
Joey Tianyi Zhou, Le Zhang 0001, Zhiwen Fang, Jiawei Du 0002, Xi Peng 0001, Yang Xiao 0007
IEEE Trans. Circuits Syst. Video Technol.3
2020 Towards Real-Time Eyeblink Detection in the Wild: Dataset, Theory and Practices
abstract
Effective and real-time eyeblink detection is of wide-range applications, such as deception detection, drive fatigue detection, face anti-spoofing. Despite previous efforts, most of existing focus on addressing the eyeblink detection problem under constrained indoor conditions with relative consistent subject and environment setup. Nevertheless, towards practical applications, eyeblink detection in the wild is highly preferred, and of greater challenges. In this paper, we shed the light to this research topic. A labelled eyeblink in the wild dataset (i.e., HUST-LEBW) of 673 eyeblink video samples (i.e., 381 positives, and 292 negatives) is first established. These samples are captured from the unconstrained movies, with the dramatic variation on face attribute, head pose, illumination condition, imaging configuration, etc. Then, we formulate eyeblink detection task as a binary spatial-temporal pattern recognition problem. After locating and tracking human eyes using SeetaFace engine and KCF (Kernelized Correlation Filters) tracker respectively, a modified LSTM model able to capture the multi-scale temporal information is proposed to verify eyeblink. A feature extraction approach that reveals the appearance and motion characteristics simultaneously is also proposed. The experiments on HUST-LEBW reveal the superiority and efficiency of our approach. The comparisons with the existing state-of-the-art methods validate the advantages of our manner for eyeblink detection in the wild.
Guilei Hu, Yang Xiao 0007, Zhiguo Cao 0001, Lubin Meng, Zhiwen Fang, Joey Tianyi Zhou, Junsong Yuan 0001
IEEE Trans. Inf. Forensics Secur.5
2019 Object Grasping of Humanoid Robot Based on YOLO
Nadia Magnenat-Thalmann, Daniel Thalmann, Zhiwen Fang, Jianmin Zheng
CGI4
2018 Understanding Human-Object Interaction in RGB-D videos for Human Robot Interaction
abstract
Detecting small hand-held objects plays a critical role for human-robot interaction, because the hand-held objects often reveal the intention of the human, e.g., use a cell phone to make a call or use a cup to drink, thus helps the robots understand the human behavior and response accordingly. Existing solutions relying on wearable sensor to detect hand-held objects often comprise the user experiences thus may not be preferred. With the development of commodity RGB-D sensors, e.g., Microsoft Kinect II, RGB and depth information have been used for the understanding of human actions and recognizing objects. Motivated by the previous success, we propose to detect hand-held objects using RGB-D sensor. However, instead of performing object detection alone, we propose to leverage human body pose as the context to achieve robust hand-held object detection in RGB-D videos. Our system demonstrates a person can interact with a humanoid social robot with hand-held object such as a cell phone or a cup. Experimental evaluations validate the effectiveness of this proposed method.
Zhiwen Fang, Junsong Yuan 0001, Nadia Magnenat-Thalmann
CGI1
2018 Incremental Upper Bound for the Maximum Clique Problem
abstract
The maximum clique problem (MaxClique for short) consists of searching for a maximum complete subgraph in a graph. A branch-and-bound (BnB) MaxClique algorithm computes an upper bound of the number of vertices of a maximum clique at every search tree node, to prune the subtree rooted at the node. Existing upper bounds are usually computed from scratch at every search tree node. In this paper, we define an incremental upper bound, called IncUB, which is derived efficiently from previous searches instead of from scratch. Then, we describe a new BnB MaxClique algorithm, called IncMC2, which uses graph coloring and MaxSAT reasoning to filter out the vertices that do not need to be branched on, and uses IncUB to prune the remaining branches. Our experimental results show that IncMC2 is significantly faster than algorithms such as BBMC and IncMaxCLQ. Finally, we carry out experiments to provide evidence that the performance of IncMC2 is due to IncUB. The online supplement is available at https://doi.org/10.1287/ijoc.2017.0770 .
Chu Min Li 0001, Zhiwen Fang, Ke Xu 0001
INFORMS J. Comput.2
2018 Toward Good Practices for Fine-Grained Maize Cultivar Identification With Filter-Specific Convolutional Activations
abstract
Crop cultivar identification is an important aspect in agricultural systems. Traditional solutions involve excessive human interventions, which is labor-intensive and timeconsuming. In addition, cultivar identification is a typical task of fine-grained visual categorization (FGVC). Compared with other common topics in FGVC, studies of this problem are somewhat lagging and limited. In this paper, targeting four Chinese maize cultivars of Jundan No.20, Wuyue No.3, Nongda No.108, and Zhengdan No.958, we first consider the problem of identifying the maize cultivar based on its tassel characteristics by computer vision. In particular, a novel fine-grained maize cultivar identification data set termed HUST-FG-MCI that contains 5000 images is first constructed. To better capture the textual differences in a weakly supervised manner, we proposed an effective deep convolutional neural network and Fisher vector (FV)based feature encoding mechanism. The mechanism tends to highlight subtle object patterns via filter-specific convolutional representations and thus provides strong discrimination for cultivar identification. Experimental results demonstrate that our method outperforms other state-of-the-art approaches. We show also that FV encoding can weaken the linear dependency between convolutional activations, redundant filters exist in the convolutional layer, and high accuracy can be maintained with relatively low-dimensional convolutional features and one or two Gaussian components in FV.
Hao Lu 0003, Zhiguo Cao 0001, Yang Xiao 0007, Zhiwen Fang, Yanjun Zhu
IEEE Trans Autom. Sci. Eng.4
2016 Fine-grained maize cultivar identification using filter-specific convolutional activations
abstract
Cultivar identification is an important aspect in agriculture and also a typical task of fine-grained visual categorization (FGVC). In comparison with other common topics in FGVC, studies on this problem are somewhat lagged and limited. In this paper, targeting four Chinese maize cultivars of Jundan No.20, Wuyue No.3, Nongda No.108, and Zhengdan No.958, we first consider the problem of identifying the maize cultivar based on its tassel characteristics. Technically, an effective convolutional neural network (CNN) based feature encoding pipeline that allows integration of deep CNN based column feature extraction, filter-specific Fisher vector (FV) encoding and mutual information (MI) based filter selection is proposed to better address this problem. In particular, a novel fine-grained maize cultivar identification dataset termed MCI-4000 that contains 4000 images is first constructed by our team. Experimental results demonstrate that our method outperforms other stat-of-the-art approaches by at least 5% in accuracy. We also show that, there exists redundant filters in the last convolutional layer, and high accuracy can be achieved with only relatively low-dimensional column features and a small number of Gaussian components in FV.
Hao Lu 0003, Zhiguo Cao 0001, Yang Xiao 0007, Zhiwen Fang, Yanjun Zhu
ICIP4
2016 An Exact Algorithm Based on MaxSAT Reasoning for the Maximum Weight Clique Problem
abstract
Recently, MaxSAT reasoning is shown very effective in computing a tight upper bound for a Maximum Clique (MC) of a (unweighted) graph. In this paper, we apply MaxSAT reasoning to compute a tight upper bound for a Maximum Weight Clique (MWC) of a wighted graph. We first study three usual encodings of MWC into weighted partial MaxSAT dealing with hard clauses, which must be satisfied in all solutions, and soft clauses, which are weighted and can be falsified. The drawbacks of these encodings motivate us to propose an encoding of MWC into a special weighted partial MaxSAT formalism, called LW (Literal-Weighted) encoding and dedicated for upper bounding an MWC, in which both soft clauses and literals in soft clauses are weighted. An optimal solution of the LW MaxSAT instance gives an upper bound for an MWC, instead of an optimal solution for MWC. We then introduce two notions called the Top-k literal failed clause and the Top-k empty clause to extend classical MaxSAT reasoning techniques, as well as two sound transformation rules to transform an LW MaxSAT instance. Successive transformations of an LW MaxSAT instance driven by MaxSAT reasoning give a tight upper bound for the encoded MWC. The approach is implemented in a branch-and-bound algorithm called MWCLQ. Experimental evaluations on the broadly used DIMACS benchmark, BHOSLIB benchmark, random graphs and the benchmark from the winner determination problem show that our approach allows MWCLQ to reduce the search space significantly and to solve MWC instances effectively. Consequently, MWCLQ outperforms state-of-the-art exact algorithms on the vast majority of instances. Moreover, it is surprisingly effective in solving hard and dense instances.
Zhiwen Fang, Chu Min Li 0001, Ke Xu 0001
J. Artif. Intell. Res.1
2016 Adobe Boxes: Locating Object Proposals Using Object Adobes
abstract
Despite the previous efforts of object proposals, the detection rates of the existing approaches are still not satisfactory enough. To address this, we propose Adobe Boxes to efficiently locate the potential objects with fewer proposals, in terms of searching the object adobes that are the salient object parts easy to be perceived. Because of the visual difference between the object and its surroundings, an object adobe obtained from the local region has a high probability to be a part of an object, which is capable of depicting the locative information of the proto-object. Our approach comprises of three main procedures. First, the coarse object proposals are acquired by employing randomly sampled windows. Then, based on local-contrast analysis, the object adobes are identified within the enlarged bounding boxes that correspond to the coarse proposals. The final object proposals are obtained by converging the bounding boxes to tightly surround the object adobes. Meanwhile, our object adobes can also refine the detection rate of most state-of-the-art methods as a refinement approach. The extensive experiments on four challenging datasets (PASCAL VOC2007, VOC2010, VOC2012, and ILSVRC2014) demonstrate that the detection rate of our approach generally outperforms the state-of-the-art methods, especially with relatively small number of proposals. The average time consumed on one image is about 48 ms, which nearly meets the real-time requirement.
Zhiwen Fang, Zhiguo Cao 0001, Yang Xiao 0007, Lei Zhu 0010, Junsong Yuan 0001
IEEE Trans. Image Process.1
2015 K-core-based attack to the internet: Is it more malicious than degree-based attack?
Jichang Zhao, Junjie Wu 0002, Zhiwen Fang, Ke Xu 0001
World Wide Web4
2014 Time-aware reciprocity prediction in trust network
abstract
Study of reciprocity helps to find influential factors for users building relationships, which greatly facilitates the social behavior understanding in trust networks. In the previous literature, the dynamics of both network structure and user generated content are rarely considered. Our investigation of the available timing information from a real-world network demonstrates that time delay has significant impact on reciprocity formation. In particular, we find structural factors possess greater effect on short-term reciprocity while factors based on user generated content become more important for long-term reciprocity. Based on the empirical analysis, we redefine the reciprocity prediction problem as a learning task specific to each pair of users with different reciprocal delays. Evaluations show that our time-aware framework eventually outperforms the conventional classifiers that ignore the temporal information. Meanwhile, we tackle the problem of concept drift through fitting the evolving trend of features for Naive Bayes and performing periodic retraining for Logistic Regression classifiers, respectively.
Jichang Zhao, Zhiwen Fang, Ke Xu 0001
ASONAM3
2014 Solving Maximum Weight Clique Using Maximum Satisfiability Reasoning
abstract
Satisfiability (SAT) and maximum satisfiability (MaxSAT) techniques are proved to be powerful in solving combinatorial optimization problems. In this paper, we encode the maximum weight clique (MWC) problem into weighted partial MaxSAT and use MaxSAT techniques to solve it. Concretely, we propose a new algorithm based on MaxSAT reasoning called Top-k failed literal detection to improve the upper bound for MWC, and implement an exact branch-and-bound solver for the MWC problem called MaxWClq based on the Top-k failed literal detection algorithm. To our best knowledge, this is the first time that MaxSat techniques are integrated to solve the MWC problem. Experimental evaluations on the broadly used DIMACS benchmark, BHOSLIB benchmark and random graphs show that MaxWClq outperforms state-of-the-art exact algorithms on the vast majority of instances. In particular, our algorithm is surprisingly powerful for dense and hard graphs.
Zhiwen Fang, Chu Min Li 0001, Kan Qiao, Ke Xu 0001
ECAI1
2013 Combining MaxSAT Reasoning and Incremental Upper Bound for the Maximum Clique Problem
abstract
Recently, MaxSAT reasoning has been shown to be powerful in computing upper bounds for the cardinality of a maximum clique of a graph. However, existing upper bounds based on MaxSAT reasoning have two drawbacks: (1)at every node of the search tree, MaxSAT reasoning has to be performed from scratch to compute an upper bound and is time-consuming, (2) due to the NP-hardness of the MaxSAT problem, MaxSAT reasoning generally cannot be complete at anode of a search tree, and may not give an upper bound tight enough for pruning search space. In this paper, we propose an incremental upper bound and combine it with MaxSAT reasoning to remedy the two drawbacks. The new approach is used to develop an efficient branch-and-bound algorithm for MaxClique, called IncMaxCLQ. We conduct experiments to show the complementarity of the incremental upper bound and MaxSAT reasoning and to compare IncMaxCLQ with several state-of-the-art algorithms for MaxClique.
Chu Min Li 0001, Zhiwen Fang, Ke Xu 0001
ICTAI2
2013 K-core-preferred Attack to the Internet: Is It More Malicious Than Degree Attack?
Jichang Zhao, Junjie Wu 0002, Zhiwen Fang, Ke Xu 0001
WAIM4