Jianzhe Lin

dblp:150/6771 · also Jianzhe Peter Lin · DBLP profile ↗
← Back
28ranked-venue papers
11as first author
17since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2024 BatchPrompt: Accomplish more with less
abstract
The ever-increasing token limits of large language models (LLMs) have enabled long context as input. Many LLMs are trained and fine-tuned to perform zero/few-shot inference using instruction-based prompts. Prompts typically include a detailed task instruction, several examples, and a single data point for inference. This baseline is referred to as “SinglePrompt” in this paper. In terms of token count, when the data input is small compared to instructions and examples, this results in lower token utilization, compared with encoder-based models like fine-tuned BERT. This cost inefficiency, affecting inference speed and compute budget, counteracts many of the benefits that LLMs offer. This paper aims to alleviate this problem by batching multiple data points in each prompt, a strategy we refer to as “BatchPrompt”. We improve token utilization by increasing the “density” of data points, however, this cannot be done naively. Simple batching can degrade performance, especially as batch size increases, and data points can yield different answers depending on their position within a prompt. To address the quality issue while retaining high token utilization, we introduce Batch Permutation and Ensembling (BPE) for BatchPrompt – a simple majority vote over repeated permutations of data, that recovers label quality at the cost of more token usage. To counterbalance this cost, we further propose Self-reflection-guided EArly Stopping (SEAS), which can terminate the voting process early for data points that the LLM handles confidently. Our comprehensive experimental evaluation demonstrates that BPE + SEAS can boost the performance of BatchPrompt by a striking margin on a range of popular NLP tasks, including question answering (Boolq), textual entailment (RTE), and duplicate questions identification (QQP). This performance is even competitive with/higher than single-data prompting (SinglePrompt), while using far fewer LLM calls and input tokens. At batch size 32, our BatchPrompt + BPE + SEAS uses 15.7% the number of LLM calls, and achieves: Boolq accuracy 90.6% → 90.9% with 27.4% tokens, QQP accuracy 87.2% → 88.4% with 18.6% tokens, RTE accuracy 91.5% → 91.1% with 30.8% tokens. We hope our simple yet effective approach will shed light on the future research of large language models. Code: github.com/microsoft/BatchPrompt
Jianzhe Lin, Maurice Diesendruck, Robin Abraham
ICLR1
2022 IntentVizor: Towards Generic Query Guided Interactive Video Summarization
abstract
The target of automatic video summarization is to create a short skim of the original long video while preserving the major content/events. There is a growing interest in the integration of user queries into video summarization or query-driven video summarization. This video summarization method predicts a concise synopsis of the original video based on the user query, which is commonly represented by the input text. However, two inherent problems exist in this query-driven way. First, the text query might not be enough to describe the exact and diverse needs of the user. Second, the user cannot edit once the summaries are produced, while we assume the needs of the user should be subtle and need to be adjusted interactively. To solve these two problems, we propose IntentVizor, an interactive video summarization framework guided by generic multi-modality queries. The input query that describes the user's needs are not limited to text but also the video snippets. We further represent these multi-modality finer-grained queries as user ‘intent’, which is interpretable, interactable, editable, and can better quantify the user's needs. In this paper, we use a set of the proposed intents to represent the user query and design a new interactive visual analytic interface. Users can interactively control and adjust these mixed-initiative intents to obtain a more satisfying summary through the interface. Also, to improve the summarization quality via video understanding, a novel Granularity-Scalable Ego-Graph Convolutional Networks (GSE-GCN) is proposed. We conduct our experiments on two benchmark datasets. Comparisons with the state-of-the-art methods verify the effectiveness of the proposed framework. Code and dataset are available at https://github.com/jnzs1836/intentvizor.
Guande Wu, Jianzhe Lin, Cláudio T. Silva
CVPR2
2022 Multilabel Aerial Image Classification With a Concept Attention Graph Neural Network
abstract
Compared with natural images, aerial images collected by satellite sensors/aerial cameras can provide a much larger field of view and often contain multiple objects of interest (multiple labels). There are certain limitations of existing multilabel aerial image classification methods. First, label correlations were often ignored in previous MAIC work, and thus, multilabel classifiers failed to be self-adapted. Second, existing multilabeled data sets for aerial images only cover limited images with fixed labels. Therefore, the underlying semantic correlations of labels cannot be fully included, while such correlation information is implicitly used as common knowledge by human beings. To tackle these concerns, we propose a novel multilabel classification method for aerial images. Our contributions are twofold. First, as the first attempt, label correlations are inferred from both the specific data set and ConceptNet (a popular knowledge graph for common sense). Second, based on graph neural network (GNN), we propose a novel end-to-end aerial image classification model, named the multiple label concept graph (ML-CG). ML-CG builds a concept graph to describe the semantic correlations from both the label set and the ConceptNet. We also incorporate both semantic attention and label attention in the GNN to better extract meaningful information of image labels. Compared with state-of-the-art methods, the effectiveness of the proposed method is demonstrated on both the commonly used UCM data set and a recently proposed DFC15 data set with high image resolution.
Dan Lin 0008, Jianzhe Lin, Liang Zhao 0005, Z. Jane Wang 0001, Zhikui Chen
IEEE Trans. Geosci. Remote. Sens.2
2022 Multilabel Aerial Image Classification With Unsupervised Domain Adaptation
abstract
Deep learning (DL) methods are promising for the multilabel aerial image classification (MAIC) task. However, current DL methods face a common problem: the need for large multilabeled datasets. Collecting and annotating raw aerial image datasets can be extremely time- and labor-consuming. To address this concern in MAIC, domain adaptation (DA) provides a novel solution by transferring the knowledge learned from a label-rich dataset (i.e., the source domain) to a label-scarce dataset (i.e., the target domain), while current DA models are mainly designed for single-labeled tasks. In this article, we propose a novel end-to-end MAIC model based on DA techniques, named DA-MAIC. To the best of our knowledge, this article for the first time integrates DA to tackle the label scarcity problem in the MAIC task. Specifically, the proposed DA-MAIC is composed of two main parts: the image classifier and the domain classifier. The image classifier captures task-discriminative features based on the graph convolutional network (GCN) to predict multiple image labels; and the domain classifier extracts domain-invariant representations, which mitigates the domain shift between two underlying distributions. We extensively evaluate the proposed DA-MAIC from different perspectives on three benchmark datasets, including the commonly used UCM dataset, the high-resolution AID dataset, and the recently proposed DFC15 dataset. Both quantitative and qualitative results support that the proposed DA-MAIC can generalize the source domain knowledge to new scenarios and substantially improve the classification performance on the target domain task.
Dan Lin 0008, Jianzhe Lin, Liang Zhao 0005, Z. Jane Wang 0001, Zhikui Chen
IEEE Trans. Geosci. Remote. Sens.2
2022 Rethinking Crowdsourcing Annotation: Partial Annotation With Salient Labels for Multilabel Aerial Image Classification
abstract
Annotated images are required for supervised model training and evaluation in aerial image classification. Manually annotating images is arduous and expensive, especially for aerial images, which often cover a large land area with multiple labels. A recent trend for conducting such annotation tasks is through crowdsourcing, where images are annotated by volunteers or paid workers (e.g., annotation volunteers for Open Street Map) online from scratch. However, for crowdsourcing image annotations, the quality cannot be guaranteed, and incompleteness and incorrectness are two major concerns. To address such concerns, we have a rethinking of crowdsourcing annotations: Our simple hypothesis is that if annotators only partially annotate multi-label images with salient labels they are confident in, there will be fewer annotation errors and annotators will spend less time on uncertain labels. As a pleasant surprise, with the same annotation budget, we show that a multi-label aerial image classifier supervised by images with salient annotations can outperform models supervised by fully annotated images. Our contributions are 2-fold: An active learning way is proposed to acquire salient labels for multi-label aerial images; and a novel Adaptive Temperature Associated Model (ATAM) specifically using partial annotations is proposed for multi-label aerial image classification. When tested on practical crowdsourcing aerial data, the Open Street Map (OSM) dataset, the proposed ATAM can achieve higher accuracy than state-of-the-art classification methods trained on fully annotated images. The proposed idea is promising for crowdsourcing aerial image annotation. Our code will be publicly available.
Jianzhe Lin, Tianze Yu, Z. Jane Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 SCIDA: Self-Correction Integrated Domain Adaptation From Single- to Multi-Label Aerial Images
abstract
Most publicly available datasets for image classification are with single labels, while images are inherently multilabeled in our daily life. Such an annotation gap makes many pretrained single-label classification models fail in practical scenarios. For aerial images, this annotation issue is more concerned: Aerial data naturally cover a relatively large land area with multiple labels, while annotated aerial datasets currently publicly available (e.g., UCM and AID) are single-labeled. As manually annotating multilabel aerial images (MAIs) would be time-/ labor-consuming, we propose a novel self-correction integrated domain adaptation (SCIDA) method for automatic multilabel learning. SCIDA is weakly supervised, i.e., automatically learning the multilabel image classification model from using massive, publicly available single-label images. To achieve this goal, we propose a novel labelwise self-correction (LWC) module to better explore underlying label correlations. This module also makes the unsupervised domain adaptation (UDA) from single-label to multilabel data possible. For model training, the proposed method uses single-label information yet requires no prior knowledge of multilabeled data and predicts labels for MAIs. Through extensive evaluations, the proposed model, which is trained with single-labeled MAI-AID-s and MAI-UCM-s datasets, achieves much better performances than comparative methods on our collected multiscene aerial image dataset. The code and data are available on GitHub (https://github.com/Ryan315/Single2multi-DA).
Tianze Yu, Jianzhe Lin, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001, Z. Jane Wang 0001
IEEE Trans. Geosci. Remote. Sens.2
2021 ERA: Entity-relationship Aware Video Summarization with Wasserstein GAN
Guande Wu, Jianzhe Lin, Cláudio T. Silva
BMVC2
2021 Reciprocal Landmark Detection and Tracking With Extremely Few Annotations
abstract
Localization of anatomical landmarks to perform two-dimensional measurements in echocardiography is part of routine clinical workflow in cardiac disease diagnosis. Automatic localization of those landmarks is highly desirable to improve workflow and reduce interobserver variability. Training a machine learning framework to perform such localization is hindered given the sparse nature of gold standard labels; only few percent of cardiac cine series frames are normally manually labeled for clinical use. In this paper, we propose a new end-to-end reciprocal detection and tracking model that is specifically designed to handle the sparse nature of echocardiography labels. The model is trained using few annotated frames across the entire cardiac cine sequence to generate consistent detection and tracking of landmarks, and an adversarial training for the model is proposed to take advantage of these annotated frames. The superiority of the proposed reciprocal model is demonstrated using a series of experiments.
Jianzhe Lin, Ghazal Sahebzamani, Christina Luong 0001, Fatemeh Taheri Dezaki, Mohammad H. Jafari 0001, Purang Abolmaesumi, Teresa Tsang
CVPR1
2021 ShuffleCount: Task-Specific Knowledge Distillation for Crowd Counting
abstract
One promising way to improve the performance of a small deep network is knowledge distillation. Performances of smaller student models with fewer parameters and lower computational cost can be comparable to that of larger teacher models in specific computer vision tasks. Knowledge distillation is especially attractive for the high-accuracy real-time crowd counting task in our daily lives, where the computational resource can be limited and the model efficiency is extremely important. In this paper, we propose a novel task-specific knowledge distillation framework for crowd counting, named ShuffleCount. Its main contributions are two-fold: First, different from existing frameworks, our task-specific ShuffleCount effectively learns from the teacher network through hierarchic feature regulation, and better avoids negative knowledge transferred from the teacher. Second, the proposed student network, i.e., the optimized Shufflenet, shows promising performances. When tested on the benchmark dataset Shanghai Tech A, it achieves a 15% higher accuracy yet keeps low computational cost when compared with the state-of-the-art MobileCount. Our code is available online at https://github.com/JiangMinyang/CC-KD.
Minyang Jiang, Jianzhe Lin, Z. Jane Wang 0001
ICIP2
2021 NYU-VPR: Long-Term Visual Place Recognition Benchmark with View Direction and Data Anonymization Influences
abstract
Visual place recognition (VPR) is critical in not only localization and mapping for autonomous driving vehicles, but also assistive navigation for the visually impaired population. To enable a long-term VPR system on a large scale, several challenges need to be addressed. First, different applications could require different image view directions, such as front views for self-driving cars while side views for the low vision people. Second, VPR in metropolitan scenes can often cause privacy concerns due to the imaging of pedestrian and vehicle identity information, calling for the need for data anonymization before VPR queries and database construction. Both factors could lead to VPR performance variations that are not well understood yet. To study their influences, we present the NYU-VPR dataset that contains more than 200,000 images over a 2km×2km area near the New York University campus, taken within the whole year of 2016. We present benchmark results on several popular VPR algorithms showing that side views are significantly more challenging for current VPR methods while the influence of data anonymization is almost negligible, together with our hypothetical explanations and in-depth analysis.
Diwei Sheng, Yuxiang Chai, Chen Feng 0002, Jianzhe Lin, Cláudio T. Silva, John-Ross Rizzo
IROS5
2021 Guest editorial: Graph learning for computer vision
abstract
Many fields in the real world involve a lot of structured data, such as social networks, transportation networks, communication networks etc., and their structures carry important information about the characteristics of the data. However, how to use its structural information to analyse and process the data efficiently has caused continuous research in the field. A graph provides an important means for dealing with structured data. It can describe the geometric structure of data intuitively and flexibly, especially in the representation of spatial irregular data. Graph learning refers to machine learning on graphs, which mainly utilises machine learning algorithms to extract the relevant features of graphs. In recent years, combined with specific applications, researchers have conducted in-depth research on graph learning and proposed various approaches. This Special Issue aims to introduce the latest studies in graph learning for computer vision and proposes new theories and approaches to solve the existing problems. It received a number of submissions from researchers in the field, which all went through a rigorous review process. After several rounds of review, six papers were accepted. These papers cover a variety of fields, such as medicine, remote sensing, and data mining. Specific tasks include image segmentation, knowledge graph reasoning, clustering, and image classification. These accepted papers are mainly divided into two categories. The first category covers the graph learning method guided by optimisation, which obtains the graph structure by establishing a clear model and solving the corresponding optimisation problem. The second category is the deep learning-oriented graph learning method, which combines convolutional neural network and graph neural network to construct the model. In the first paper of the Special Issue by Wang et al. entitled An Enhanced 3D U-Net with Graph-based Refining for Segmentation of Gastrointestinal Stromal Tumours, the authors propose a segmentation algorithm using an improved 3D U-Net to segment gastrointestinal stromal tumours. To enhance information transmission, multiple skip connections are attached into same size feature maps between an encoder and a decoder. Due to difficulties in tumour labelling and other reasons, the author transforms the small intestinal segmentation model into a gastrointestinal stromal tumour segmentation model. Since fully convolutional networks typically suffer from inaccuracies around the boundaries of small structures, the graph neural network is introduced to refine segmentation results. Experiments demonstrate that the proposed method presents superior performance over traditional U-Net. The second paper by Ma et al. entitled Hybrid Attention Mechanism for Few-Shot Relational Learning of Knowledge Graphs, develops a few-shot relationship learning framework. The authors first design an entity-enhanced encoder with weak attention networks and self-attention mechanisms to explore the influence of different levels for source entities. The local graph structure is then utilised to enhance the embedding of the source entity by combining explicit and implicit features. Finally, the model parameters are optimised to infer real entities in the candidate set of similar entities obtained by a loop-processing matching processor. The authors provide extensive experiments and confirm the excellent accuracy of the proposed model. The third paper by Zhao et al. entitled Incremental Multi-View Correlated Feature Learning Based on Non-Negative Matrix Factorization, studies multi-view data. The authors present an incremental multi-view correlated feature learning approach based on non-negative matrix factorization to analyse the uncorrelated items in each view. The algorithm separates uncorrelateditems across views and constructs incremental joint learning with uncorrelated and correlated features to study the common features for multi-view data. Subsequently, the authors design an incremental objective function and derive an effective updating scheme. The proposed method is proved to converge effectively, and its complexity is discussed. They evaluate the proposed solution on real-world datasets and report excellent performance in comparison with the existing state-of-the-art solutions. The fourth paper by Hu et al. entitled Complete/Incomplete Multi-view Subspace Clustering via Soft Block-Diagonal-Induced Regularizer, concentrates on complete and incomplete multi-view clustering problems. The proposed method adopts the self-representation model to individually construct the similarity graphs for each view. To fuse a shared affinity matrix for all views, the authors design the soft block-diagonal-induced regulariser to encourage the generation of a matrix with K diagonal blocks. Considering the incomplete multi-view data, the proposed method effectively utilises some indicator matrices to accurately mark the missing instances in each view. The authors analyse the complexity and convergence of the proposed method on four public datasets and demonstrate that it is better than the most advanced complete/incomplete clustering methods. The fifth paper by Gong et al. entitled Few-shot Learning with Relation Propagation and Constraint, aims to extract valuable information of pair-wise correlation between sparse training samples. The authors state that transductive relation propagation simply propagates the pair-wise relation without relation constraints. Thus, the paper develops a constrained relation–propagation network to capture the accurate relation so as to generate discriminative relational representations. To constrain the pair-wise relation, the proposed method introduces a relation constraint module to regularise the distilled relations between samples, which helps to calibrate the propagated correlation information. Extensive experiments conducted on several benchmark datasets indicate that the proposed method achieves remarkable performance compared to few-shot learning methods. The last paper by Guo et al. entitled CNN-Combined Graph Residual Network with Multilevel Feature Fusion for Hyperspectral Image Classification introduces graph convolutional networks to obtain more superpixel-level features with a topological structure. This paper develops an effective CNN-combined graph residual network with a multilevel feature fusion strategy. The main idea is to learn superpixeltopological information by using the graph residual network and pixel information by using the convolutional neural network. The strategy can adequately leverage the superpixel level and pixel-level features and capture the class boundary features, which further enhances the generalisation performance. Experiments report highly competitive performance in comparison to existing hyperspectral image classification approaches. The papers selected in this Special Issue highlight the extensive study of graph learning in computer vision. We hope that these papers can promote the theoretical study of graph learning as well as provide new ideas for more researchers who are committed to graph learning.
Qi Wang 0009, Hongkai Yu, Song Wang 0002, Jianzhe Lin
IET Comput. Vis.4
2021 Crowd understanding and analysis
abstract
With the development of modern life, all kinds of social activities have become more frequent. These social activities are often attended by a wide range of people, which puts forward high requirements for effective management and ensures the safety of the people involved in the activities. As an effective auxiliary measure, crowd understanding and analysis has been concerned by more and more researchers. The basic idea is to extract the key information from video sequences/images, and use digital image processing technology to study and analyse the behaviour characteristics and patterns of people in the region of interest. Currently, it has a wide range of applications in the fields of economy, public security, and so forth. This Special Issue aims to introduce the latest works in crowd understanding and analysis, and proposes new theory and approaches to solve the existing problems. The Special Issue received a number of submissions from researchers in the field. Finally, 27 papers were accepted after careful peer reviews and revisions. The accepted papers can be broadly divided into six sections according to the types of tasks, including Section 1: Crowd behaviour detection and recognition, Section 2: Crowd counting, Section 3: Object detection and recognition, and Section 4: Object tracking, and Section 5: Other tasks. “Human behaviour recognition with mid-level representations for crowd understanding and analysis” of Sun et al. uses mid-level semantic concepts to represent human actions from videos and argues that these semantic attributes enable the construction of more descriptive methods for human action recognition. The idea is verified on three challenging datasets, and the experimental results demonstrate that their method achieves better results than the baseline methods on human action recognition. “Class structure-aware adversarial loss for cross-domain human action recognition” of Chen et al. proposes a class structure-aware adversarial loss, which aims to address the issue that the existing adversarial-based approaches ignore the underlying coherence of class structure across domains. This paper incorporates category information into the adversarial learning branch to capture the fine-grained alignment of each class, effectively avoiding the false mix-up of samples from different categories in the embedding space. Experiments show significant improvement compared to the baseline. “Dual-view 3D human pose estimation without camera parameters for action recognition” of Liu et al. proposes a dual-view single-person 3D pose estimation method without camera parameters. This method first uses the 2D pose estimation network to estimate the 2D joint point coordinates from two images with different views, and then inputs them into the 3D regression network to generate the final 3D joint point coordinates. Experiments show that this method is effective. “Deep social force network for anomaly event detection” of Yang et al. develops a deep social force network by exploiting both social force extraction and deep motion coding. This network can discover the interaction force of particles to learn the deep social force features. The experiments on UCF-Crime and ShanghaiTech datasets demonstrate that this method can predict the temporal localization of anomaly events and outperform the state-of-the-art methods. “Anomaly detection in video sequences: a benchmark and computational model” of Wan et al. contributes a new Large-scale Anomaly Detection (LAD) dataset as the benchmark for anomaly detection in video sequences. This paper formulates anomaly detection as a fully supervised learning problem and proposes a multi-task deep neural network to solve it. Experimental results show that the proposed method outperforms state-of-the-art anomaly detection competitors on the proposed dataset and other public datasets. “Behaviour detection in crowded classroom scenes via enhancing features robust to scale and perspective variations” of Liu et al. proposes two modules to tackle the large variations of humans in scale and pose perspective, namely Scale Attention Aggregation module and RoI Spatial Transformation module. Besides, this paper constructs a new classroom human behaviour detection dataset with 1500 images based on a third-person perspective. The proposed method is verified on the proposed dataset. Experimental results demonstrate the effectiveness of the proposed method with better mAP values. “Crowd activity recognition in live video streaming via 3D-ResNet and region graph convolution network” of Kang et al. presents a crowd activity recognition method to identify and supervise the crowd activity from live videos. The method utilizes 3D-ResNet and ReGCN to extract deep spatiotemporal features and the correlation or external knowledge between crowd content, respectively. In addition to the above work, the authors construct a real-world video dataset called BJUT-CAD that includes eight kinds of crowd activity videos collected from live video websites. Experiments on BJUT-CAD and CAE datasets verify the effectiveness of this work. “Latent label mining for group activity recognition in basketball videos” of Wu et al. proposes a latent label mining strategy for group activity recognition in basketball videos. This paper aims at mining the latent labels of motion patterns from the frames and further combining two levels of supervision signal to obtain effective spatio-temporal representation. Experimental results demonstrate that the proposed algorithm achieves state-of-the-art performance. “A deep learning method for video-based action recognition” of Zhang et al. aims to address a key issue, that is, how to convert spatial and temporal information into an effective representation to infer actions. This paper employs boundary compensation on the basis of a deep neural network to achieve action proposal. Based on the resultant action proposals, a two-stream network with a spatio-temporal structure is adopted for the action recognition task. The experimental results show the competitive performance of the proposed method over the state-of-the-art methods. “MSR-FAN: Multi-scale residual feature-aware network for crowd counting” of Zhao et al. proposes a framework that combines the multi-scale features using multiple receptive field sizes and learns the feature-aware information on each image. This method effectively alleviates perspective distortion and the varying scales in congested scene images, which helps the algorithm crowd counting correctly. Experiment results on benchmark datasets indicate that the proposed approach outperforms the existing competitors. “MFP-Net: Multi-scale feature pyramid network for crowd counting” of Lei et al. introduces a feature pyramid fusion module and a feature attention-aware module. Two modules can extract different levels of fine-grained information; local and global context information, respectively. It enhances the correlation of different features and improves robustness to background noise effectively. Experiments show that the proposed method not only provides better crowd counting results than comparative models, but also requires fewer parameters. “Multi-level features extraction network with gating mechanism for crowd counting” of Zeng et al. designs a novel crowd counting model, which integrates multi-level information from multiple levels, such as appearance, scale, and context. To avoid interference from confusing information, a simple and effective multi-channel gated unit is proposed to adaptively select features at different levels of the network. Extensive experiments and evaluations clearly illustrate that the approach is superior. “Learn from object counting: crowd counting with meta-learning” of Zan et al. develops an efficient algorithm to extract the meta-information via utilizing object counting data in few-shot scenes. This method successfully explores the shared information between the crowd counting task and the object counting task, thus improving the performance and convergence rate of the crowd counting task. Comprehensive experiments on two datasets verify the effectiveness of the proposed approach. “Crowd estimation using key-point matching with support vector regression” of Ekanayake et al. proposes a novel key-point-based moving object detection in noisy backgrounds. This work mainly uses the key-point matching of continuous frames and optical flow density to identify moving objects in video sequences. The moving blobs are then generated via morphological operations. The experiments are compared with recent regression-based and CNN methods, and it is verified that the proposed work is superior. “A novel face recognition method based on fusion of LBP and HOG” of Chen et al. proposes an improved fusion local feature extraction algorithm called CS-NWALBP+HOG. This study not only smoothens noise sensitivity of the LBP operator, but also reduces the original computational complexity, as well as strengthens the description ability for image gradient direction information. Several experiments eventually demonstrate that the designed algorithm shows more robust performance under complex illumination conditions. “Multi-view intrinsic low-rank representation for robust face recognition and clustering” of Shen et al. considers the problem that the most existing methods ignore the specific local structure of different views. To address this issue, the paper proposes a multi-view low-rank representation method which exploits both intrinsic relationships and specific local structures of different views simultaneously. Experiments on several datasets demonstrate the effectiveness of this method in classification and clustering. “Multi-dimensional weighted cross-attention network in crowded scenes” of Xie et al. proposes an end-to-end anchor-free network, namely Multidimensional Weighted Cross-Attention Network, which can perform real-time human detection in crowded scenes. The designed model does not require manual intervention, and reduces the sizeable computational resource cost due to the anchor boxes mapping during the training process. Experiments reveal that the improved strategy achieves state-of-the-art results in the anchor-free methods. “Part-level attention networks for cross-domain person re-identification” of Zhao et al. uses the diversified spatial semantic feature in pixel-level learning in the target domain to improve the generality and adaptability of the model. Combined with partial branches, this method proves that it is effective to add an attention cascade module to the backbone network. Experiments indicate that the proposed approach has better recognition ability and robustness in cross-domain aspect. “MFNet-LE: Multilevel fusion network with Laplacian embedding for face presentation attacks detection” of Niu et al. proposes a face presentation attack detection method by incorporating a multilevel fusion structure and Laplacian loss into shallow CNNs. This allows the proposed model to learn more discriminative features under the joint supervision of softmax and Laplacian loss, which improves the detection ability. Experiments demonstrate the effectiveness of the proposed method. “Real-time automatic helmet detection of motorcyclists in urban traffic using improved YOLOv5 detector” of Jia et al. presents an automatic helmet detection that contains two steps. This work first utilizes improved YOLOv5 detector to detect motorcycles. Then, the motorcycle detected in the above step is input, and the improved YOLOv5 detector is used again to detect whether the rider is wearing a helmet. The proposed method is evaluated on constructed dataset, and the results show that it is superior to other detection methods. “Multi-label learning based target detecting from multi-frame data” of Mei et al. regards target detection from time series data as a multi-label problem to design the model. In this method, a background subtraction tracker is presented to track the slightly moving object in videos, which are based on Gaussian mixture model background subtraction and integral image. Experimental results show that the proposed method attains better performance. “Contrastive learning of graph encoder for accelerating pedestrian trajectory prediction training” of Yao et al. proposes a graph contrastive accelerating encoder. It accelerates the pedestrian trajectory prediction training process of spatiotemporal graph transformer networks. This method makes the pedestrian trajectory prediction error the lowest in the obviously early training steps, and makes the final performance reach the state-of-the-art level. “Multiple object tracking based on multi-task learning with strip attention” of Song et al. believes that it is difficult to strike a balance between accuracy and efficiency by embedding the re-identification (re-ID) model into the target tracking task. To enhance the overall tracking performance, a one-shot multiple object tracking is proposed based on multi-task learning, which contains two homogeneous branches of object detection and re-ID. By the fine-grained features extraction in pedestrian recognition, it benefits overall tracking framework improvement in both speed and robustness. The experiments show that the proposed method attains superior performance in more evaluation metrics. “Cross-modal semantic correlation learning by Bi-CNN network” of Wang et al. presents a novel cross modal retrieval framework, which integrates feature learning and latent space embedding. It aims to generate specific representations consistent with cross-modal tasks. It helps to reduce the differences in the distribution of categories in different modalities. Experiments on three real-word datasets show that the proposed work is superior to the popular methods. “Adaptive colour restoration and detail retention for image enhancement” of He et al. aims to overcome the problem of colour distortion caused by low illumination and fog. Considering the issue, this paper develops a multi-channel fusion-based adaptive image colour restoration method. To generate human-consistent observations, the detailed retention-based method is applied to enhance the details. Experiments demonstrate that the results are effective and outperform the compared methods both in visual and objective evaluations. “Image encryption algorithm for crowd data based on a new hyperchaotic system and Bernstein polynomial” of Jiang et al. designs a new two-dimensional chaotic system with hyperchaotic behaviour based on the Chebyshev system and the infinite collapse system. To protect the crowd image data, an image cryptosystem combined with the SVD and Bernstein polynomial is proposed. Security analyses indicate that this method has higher encryption efficiency and the visual quality of steganography image can reach 39 dB. “CA-PMG: Channel attention and progressive multi-granularity training network for fine-grained visual classification” of Zhao et al. designs a framework for visual classification for the subtle intra-class object variations. This model can be trained efficiently in an end-to-end manner without bounding box or part annotations. Extensive experiments on three challenging fine-grained datasets demonstrate that the approach obtains state-of-the-art performance. The papers selected in this Special Issue highlight the extensive study of crowd understanding and analysis. It is hoped that these papers can play a role in promoting theoretical research. Meanwhile, there are many challenges in this field that need to be further studied, such as the robustness of algorithms in cross-scenarios, the interaction between groups and individuals in different scenarios, and so forth. In addition, key problems such as the running time of the algorithm also need to be considered. These works will be helpful when applied to research in the real world. Lead Guest Editor Professor Qi Wang, Northwestern Polytechnical University, Xi'an, China Guest Editors Assistant Professor Bo Liu, Auburn University, USA Dr. Jianzhe Lin, University of British Columbia, Canada
Qi Wang 0009, Bo Liu 0006, Jianzhe Lin
IET Image Process.3
2021 A smartly simple way for joint crowd counting and localization
Minyang Jiang, Jianzhe Lin, Z. Jane Wang 0001
Neurocomputing2
2021 Spectral-Spatial Hyperspectral Image Classification Using Dual-Channel Capsule Networks
abstract
Deep learning methods have shown their marvel performance on hyperspectral image (HSI) classification tasks. In particular, algorithms based on convolution neural network (CNN) outperformed most of the conventional machine learning-based algorithms and have become the mainstream of the current HSI classification research works. Recently, a newly proposed neural network called capsule network (CapsNet) showed its potential to replace the CNNs in various classification tasks with its amazing performance. In this letter, we proposed a new network architecture based on the CapsNet for HSI classification tasks, called dual-channel capsule network (DCCapsNet). Our DCCapsNet model extracts the features from spectral and spatial domains, respectively, with two separate convolution channels and then concatenates and feeds them into the following capsule layers to classify each of the HSI pixels. The model was trained and validated on four real HSI data sets and achieved high accuracy. We also compared our network with some of the state-of-the-art models and found that our model outperformed these competitor models.
Wenbo Liu 0005, Yue Zhang 0036, Jianzhe Lin
IEEE Geosci. Remote. Sens. Lett.6
2021 Discriminative feature alignment: Improving transferability of unsupervised domain adaptation by Gaussian-guided latent alignment
Jing Wang 0112, Jiahong Chen, Jianzhe Lin, Leonid Sigal, Clarence W. de Silva
Pattern Recognit.3
2021 Attention-Aware Pseudo-3-D Convolutional Neural Network for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) have been applied for hyperspectral image classification recently. Among this class of deep models, 3-D CNN has been shown to be more effective by learning discriminative features from abundant spectral signatures and spatial contexts in hyperspectral imagery (HSI). However, by simply imposing 3-D CNN to HSI, a large amount of initial information might be lost in this CNN pipeline. The proposed attention-aware pseudo-3-D (AP3D) convolutional network for HSI classification is motivated by two observations. First, each dimension of the 3-D HSI is not equally important, different attention should be paid to different dimensions of the initial HSI image, especially in the first convolution operation. Second, intermediate representations of the 3-D input image at different stages in the 3-D CNN pipeline represent different levels of features and should not be neglected and abandoned. Instead, a 2-D matrix of scores for each feature map should be fed to the final softmax layer. Quantitative and qualitative results demonstrate that the proposed AP3D model outperforms the state-of-the-art HSI classification methods in agricultural and rural/urban data sets: Indian Pines, Pavia University, and Salinas Scene.
Jianzhe Lin, Lichao Mou, Xiao Xiang Zhu 0001, Xiangyang Ji, Z. Jane Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 Unifying Top-Down Views by Task-Specific Domain Adaptation
abstract
In this article, we aim to learn a unified representation of images from satellite/aerial/ground views by exploring their underlying correlations. Inspired by recent advances in domain adaptation (DA), we propose a novel task-specific DA method for this purpose. Different from traditional DA methods, this proposed method not only applies task-specific classifiers1but also introduces domain-specific tasks for different domains during the adaptation process. The experiments are conducted on two newly proposed ground-/satellite-to-aerial scene adaptation (GSSA) data sets. Since the semantic gap between the ground/satellite scenes and the aerial scenes is much larger than that between ground scenes, the DA task between these scenes is more challenging than traditional DA tasks. On GSSA data sets, we not only demonstrate the proposed unsupervised DA method but also explore the few-shot DA in the discussion section. The proposed method is easy to implement, and our method substantially outperforms the state-of-the-art methods on the studied data sets.We hope that the proposed method for the novel GSSA data sets can be a good baseline for future researchers. The related data sets/codes will be available online.
Jianzhe Lin, Tianze Yu, Lichao Mou, Xiao Xiang Zhu 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 Dual Adversarial Network for Unsupervised Ground/Satellite-to-Aerial Scene Adaptation
abstract
Recent domain adaptation work tends to obtain a uniformed representation in an adversarial manner through joint learning of the domain discriminator and feature generator. However, this domain adversarial approach could render sub-optimal performances due to two potential reasons: First, it might fail to consider the task at hand when matching the distributions between the domains. Second, it generally treats the source and target domain data in the same way. In our opinion, the source domain data which serves the feature adaption purpose should be supplementary, whereas the target domain data mainly needs to consider the task-specific classifier. Motivated by this, we propose a dual adversarial network for domain adaptation, where two adversarial learning processes are conducted iteratively, in correspondence with the feature adaptation and the classification task respectively. The efficacy of the proposed method is first demonstrated on Visual Domain Adaptation Challenge (VisDA) 2017 challenge, and then on two newly proposed Ground/Satellite-to-Aerial Scene adaptation tasks. For the proposed tasks, the data for the same scene is collected not only by the traditional camera on the ground, but also by satellite from the out space and unmanned aerial vehicle (UAV) at the high-altitude. Since the semantic gap between the ground/satellite scene and the aerial scene is much larger than that between ground scenes, the newly proposed tasks are more challenging than traditional domain adaptation tasks. The datasets/codes can be found at https://github.com/jianzhelin/DuAN.
Jianzhe Lin, Lichao Mou, Tianze Yu, Xiao Xiang Zhu 0001, Z. Jane Wang 0001
ACM Multimedia1
2020 Xnet: Task-specific attentional domain adaptation for satellite-to-aerial scene
Jianzhe Lin, Kaiwen Yuan, Rabab K. Ward, Z. Jane Wang 0001
Neurocomputing1
2020 DT-LET: Deep transfer learning by exploring where to transfer
Jianzhe Lin, Liang Zhao 0005, Qi Wang 0009, Rabab K. Ward, Z. Jane Wang 0001
Neurocomputing1
2020 An End-to-End Multi-Task Deep Learning Framework for Skin Lesion Analysis
abstract
Automatic skin lesion analysis of dermoscopy images remains a challenging topic. In this paper, we propose an end-to-end multi-task deep learning framework for automatic skin lesion analysis. The proposed framework can perform skin lesion detection, classification, and segmentation tasks simultaneously. To address the class imbalance issue in the dataset (as often observed in medical image datasets) and meanwhile to improve the segmentation performance, a loss function based on the focal loss and the jaccard distance is proposed. During the framework training, we employ a three-phase joint training strategy to ensure the efficiency of feature learning. The proposed framework outperforms state-of-the-art methods on the benchmarks ISBI 2016 challenge dataset towards melanoma classification and ISIC 2017 challenge dataset towards melanoma segmentation, especially for the segmentation task. The proposed framework should be a promising computer-aided tool for melanoma diagnosis.
Lei Song 0003, Jianzhe Lin, Z. Jane Wang 0001, Haoqian Wang
IEEE J. Biomed. Health Informatics2
2019 A Coarse-to-Fine Optimization for Hyperspectral Band Selection
abstract
Hyperspectral band selection is a feature selection method that selects a most representative set of bands to achieve a good performance in several tasks such as classification and anomaly detection. It reduces the burden of storage, transmission, and computation. In this letter, a two-stage band selection algorithm is introduced. It selects bands and refines the result using a linear reconstruction error criterion. Then a coarse-to-fine band selection (CFBS) strategy is applied to the two-stage band selection in order to achieve a better result. CFBS selects bands group by group. Each group is selected based on bands that are not well represented by the previous groups, trying to minimize the linear reconstruction error. Experiments show that the proposed method has a significant advancement compared with other competitors.
Jianzhe Lin, Yanning Zhang 0001
IEEE Geosci. Remote. Sens. Lett.2
2018 Deep Transfer Learning for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) includes a vast quantities of samples, large number of bands, as well as randomly occurring redundancy. Classifying such complex data is challenging, and the classification performance generally is affected significantly by the amount of labeled training samples. Collecting such labeled training samples is labor and time consuming, motivating the idea of borrowing and reusing labeled samples from other preexisting related images. Therefore transfer learning, which can mitigate the semantic gap between existing and new HSI, has recently drawn increasing research attention. However, existing transfer learning methods for HSI which concentrated on how to overcome the divergence among images, may neglect the high level latent features during the transfer learning process. In this paper, we present two novel ideas based on this observation. We propose constructing and connecting higher level features for the source and target HSI data, to further overcome the cross-domain disparity. Different from existing methods, no priori knowledge on the target domain is needed for the proposed classification framework, and the proposed framework works for both homogeneous and heterogenous HSI data. Experimental results on real world hyperspectral images indicate the significance of the proposed method in HSI classification.
Jianzhe Lin, Rabab K. Ward, Z. Jane Wang 0001
MMSP1
2017 Structure Preserving Transfer Learning for Unsupervised Hyperspectral Image Classification
abstract
Recent advances on remote sensing techniques allow easier access to imaging spectrometer data. Manually labeling and processing of such collected hyperspectral images (HSIs) with a vast quantities of samples and a large number of bands is labor and time consuming. To relieve these manual processes, machine learning based HSI processing methods have attracted increasing research attention. A major assumption in many machine learning problems is that the training and testing data are in the same feature space and follow the same distribution. However, this assumption doesn’t always hold true in many real world problems, especially in certain HSI processing problems with extremely insufficient or even without training samples. In this letter, we present a transfer learning framework to address this unsupervised challenge (i.e., without training samples in the target domain), by making the following three main contributions: 1) to the best of our knowledge, this is the first time for transfer learning framework to be used for the classification of totally unknown target HSI data with no training samples; 2) the characteristics of HSI are learned on dual spaces to exploit its structure knowledge to better label HSI samples; and 3) two specific new scenarios suitable for transfer learning are investigated. Experimental results on several real world HSIs support the superiority of the proposed work.
Jianzhe Lin, Chen He 0002, Z. Jane Wang 0001
IEEE Geosci. Remote. Sens. Lett.1
2016 Hyperspectral Image Classification via Multitask Joint Sparse Representation and Stepwise MRF Optimization
abstract
Hyperspectral image (HSI) classification is a crucial issue in remote sensing. Accurate classification benefits a large number of applications such as land use analysis and marine resource utilization. But high data correlation brings difficulty to reliable classification, especially for HSI with abundant spectral information. Furthermore, the traditional methods often fail to well consider the spatial coherency of HSI that also limits the classification performance. To address these inherent obstacles, a novel spectral-spatial classification scheme is proposed in this paper. The proposed method mainly focuses on multitask joint sparse representation (MJSR) and a stepwise Markov random filed framework, which are claimed to be two main contributions in this procedure. First, the MJSR not only reduces the spectral redundancy, but also retains necessary correlation in spectral field during classification. Second, the stepwise optimization further explores the spatial correlation that significantly enhances the classification accuracy and robustness. As far as several universal quality evaluation indexes are concerned, the experimental results on Indian Pines and Pavia University demonstrate the superiority of our method compared with the state-of-the-art competitors.
Yuan Yuan 0001, Jianzhe Lin, Qi Wang 0009
IEEE Trans. Cybern.2
2016 Dual-Clustering-Based Hyperspectral Band Selection by Contextual Analysis
abstract
Hyperspectral image (HSI) involves vast quantities of information that can help with the image analysis. However, this information has sometimes been proved to be redundant, considering specific applications such as HSI classification and anomaly detection. To address this problem, hyperspectral band selection is viewed as an effective dimensionality reduction method that can remove the redundant components of HSI. Various HSI band selection methods have been proposed recently, and the clustering-based method is a traditional one. This agglomerative method has been considered simple and straightforward, while the performance is generally inferior to the state of the art. To tackle the inherent drawbacks of the clustering-based band selection method, a new framework concerning on dual clustering is proposed in this paper. The main contribution can be concluded as follows: 1) a novel descriptor that reveals the context of HSI efficiently; 2) a dual clustering method that includes the contextual information in the clustering process; 3) a new strategy that selects the cluster representatives jointly considering the mutual effects of each cluster. Experimental results on three real-world HSIs verify the noticeable accuracy of the proposed method, with regard to the HSI classification application. The main comparison has been conducted among several recent clustering-based band selection methods and constraint-based band selection methods, demonstrating the superiority of the technique that we present.
Yuan Yuan 0001, Jianzhe Lin, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.2
2016 Salient Band Selection for Hyperspectral Image Classification via Manifold Ranking
abstract
Saliency detection has been a hot topic in recent years, and many efforts have been devoted in this area. Unfortunately, the results of saliency detection can hardly be utilized in general applications. The primary reason, we think, is unspecific definition of salient objects, which makes that the previously published methods cannot extend to practical applications. To solve this problem, we claim that saliency should be defined in a context and the salient band selection in hyperspectral image (HSI) is introduced as an example. Unfortunately, the traditional salient band selection methods suffer from the problem of inappropriate measurement of band difference. To tackle this problem, we propose to eliminate the drawbacks of traditional salient band selection methods by manifold ranking. It puts the band vectors in the more accurate manifold space and treats the saliency problem from a novel ranking perspective, which is considered to be the main contributions of this paper. To justify the effectiveness of the proposed method, experiments are conducted on three HSIs, and our method is compared with the six existing competitors. Results show that the proposed method is very effective and can achieve the best performance among the competitors.
Qi Wang 0009, Jianzhe Lin, Yuan Yuan 0001
IEEE Trans. Neural Networks Learn. Syst.2
2014 In defense of iterated conditional mode for hyperspectral image classification
abstract
Hyperspectral image classification is one of the most significant topics in remote sensing. A large number of methods have been proposed to improve the classification accuracy. However, the improvement often comes at the cost of higher complexity. In this work, we mainly focus on the Markov Random Fields related paradigm, which involves a demanding energy minimization procedure. Traditional methods are prone to employ the advanced optimization techniques. On the contrary, this paper is in defense of a simple yet efficient method for hyperspectral image classification, Iterated Conditional Mode, which has been generally considered inferior to other state-of-the-art methods. Our purpose is successfully achieved by tackling two inherent drawbacks of ICM, sensitive label initialization and local minimum. We apply our method to three real-world hyperspectral images, and compare the results with those of state-of-the-art methods. The comparisons show that the proposed method outperforms its competitors.
Jianzhe Lin, Qi Wang 0009, Yuan Yuan 0001
ICME1