VLDB 2026 Research / reviewers in the wild / expert
Haokui Zhang
dblp:197/5431
· DBLP profile ↗
30ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0002-4336-5558ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 7 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UVLM: Benchmarking Video Language Model for Underwater World UnderstandingabstractRecently, video-language models (VidLMs) have gained widespread attention and adoption. However, existing works primarily focus on terrestrial scenarios, overlooking the highly demanding application needs of underwater observation. To overcome this gap, we introduce UVLM, an under water observation benchmark which is build through a collaborative approach combining human expertise and AI models. To ensure data quality, we have conducted in-depth considerations from multiple perspectives. First, to address the unique challenges of underwater environments, we selected videos that represent typical underwater challenges including light variations, water turbidity, and diverse viewing angles to construct the dataset. Second, to ensure data diversity, the dataset covers a wide range of frame rates, resolutions, 419 classes of marine animals, and various static plants and terrains. Next, for task diversity, we adopted a structured design where observation targets are categorized into two major classes: biological and environmental. Each category includes content observation and change/action observation, totaling 20 subtask types. Finally, we designed several challenging evaluation metrics to enable quantitative comparison and analysis of different methods. Experiments on two representative VidLMs demonstrate that fine-tuning VidLMs on UVLM significantly improves underwater world understanding while also showing potential for slight improvements on existing in-air VidLM benchmarks. Xizhe Xue, Dawei Yan 0001, Lijie Tao, Ying Li 0017, Haokui Zhang, Rong Xiao 0003 |
AAAI | 7 |
| 2026 | Teacher Agent: A Knowledge Distillation-Free Framework for Rehearsal-Based Video Incremental Learning
Shengqin Jiang, Yaoyu Fang, Haokui Zhang, Qingshan Liu 0001, Yuankai Qi, Yang Yang 0002, Peng Wang 0023 |
Int. J. Comput. Vis. | 3 |
| 2025 | TG-LLaVA: Text Guided LLaVA via Learnable Latent EmbeddingsabstractCurrently, inspired by the success of vision-language models (VLMs), an increasing number of researchers are focusing on improving VLMs and have achieved promising results. However, most existing methods concentrate on optimizing the connector and enhancing the language model component, while neglecting improvements to the vision encoder itself. In contrast, we propose Text Guided LLaVA (TG-LLaVA) in this paper, which optimizes VLMs by guiding the vision encoder with text, offering a new and orthogonal optimization direction. Specifically, inspired by the purpose-driven logic inherent in human behavior, we use learnable latent embeddings as a bridge to analyze textual instruction and add the analysis results to the vision encoder as guidance, refining it. Subsequently, another set of latent embeddings extracts additional detailed text-guided information from high-resolution local patches as auxiliary information. Finally, with the guidance of text, the vision encoder can extract text-related features, similar to how humans focus on the most relevant parts of an image when considering a question. This results in generating better answers. Experiments on various datasets validate the effectiveness of the proposed method. Remarkably, without the need for additional training data, our proposed method can bring more benefits to the baseline (LLaVA-1.5) compared with other concurrent methods. Furthermore, the proposed method consistently brings improvement in different settings. Dawei Yan 0001, Hao Chen 0041, Weihua Luo, Wei Dong 0010, Qingsen Yan, Haokui Zhang, Chunhua Shen |
AAAI | 9 |
| 2025 | NN-Former: Rethinking Graph Structure in Neural Architecture RepresentationabstractThe growing use of deep learning necessitates efficient network design and deployment, making neural predictors vital for estimating attributes such as accuracy and latency. Recently, Graph Neural Networks (GNNs) and transformers have shown promising performance in representing neural architectures. However, each of both methods has its disadvantages. GNNs lack the capabilities to represent complicated features, while transformers face poor generalization when the depth of architecture grows. To mitigate the above issues, we rethink neural architecture topology and show that sibling nodes are pivotal while overlooked in previous research. We thus propose a novel predictor leveraging the strengths of GNNs and transformers to learn the enhanced topology. We introduce a novel token mixer that considers siblings, and a new channel mixer named bidirectional graph isomorphism feed-forward network. Our approach consistently achieves promising performance in both accuracy and latency prediction, providing valuable insights for learning Directed Acyclic Graph (DAG) topology. The code is available at https://github.com/XuRuihan/NNFormer. Ruihan Xu 0002, Haokui Zhang, Yaowei Wang 0001, Wei Zeng 0006, Shiliang Zhang |
CVPR | 2 |
| 2025 | Efficient Adaptation of Pre-Trained Vision Transformer Underpinned by Approximately Orthogonal Fine-Tuning StrategyabstractA prevalent approach in Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViT) involves freezing the majority of the backbone parameters and solely learning low-rank adaptation weight matrices to accommodate downstream tasks. These low-rank matrices are commonly derived through the multiplication structure of down-projection and up-projection matrices, exemplified by methods such as LoRA and Adapter. In this work, we observe an approximate orthogonality among any two row or column vectors within any weight matrix of the backbone parameters; however, this property is absent in the vectors of the down/up-projection matrices. Approximate orthogonality implies a reduction in the upper bound of the model's generalization error, signifying that the model possesses enhanced generalization capability. If the fine-tuned down/up-projection matrices were to exhibit this same property as the pre-trained backbone matrices, could the generalization capability of fine-tuned ViTs be further augmented? To address this question, we propose an Approximately Orthogonal Fine-Tuning (AOFT) strategy for representing the low-rank weight matrices. This strategy employs a single learnable vector to generate a set of approximately orthogonal vectors, which form the down/up-projection matrices, thereby aligning the properties of these matrices with those of the backbone. Extensive experimental results demonstrate that our method achieves competitive performance across a range of downstream image classification tasks, confirming the efficacy of the enhanced generalization capability embedded in the down/up-projection matrices. Yiting Yang, Qingsen Yan, Haokui Zhang, Wei Dong 0010, Guoqing Wang 0001, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ICCV | 5 |
| 2025 | Text-Visual Semantic Constrained AI-Generated Image Quality Assessment
Qingsen Yan, Haojian Huang, Peng Wu 0015, Haokui Zhang, Yanning Zhang 0001 |
ACM Multimedia | 5 |
| 2025 | Cell-Wise Self-Optimization: Making Pretrained Model Better in Remote Sensing CountingabstractRecently, remote sensing counting has drawn a lot of attention due to its wide application requirements. However, most existing approaches tend to focus on optimizing the network backend or designing new loss functions, and overlook a foundational component, i.e., the feature extractor, which is usually based on a pre-trained model. This oversight limits the potential for further performance improvement. In this paper, we propose a cell-wise self-optimization method to enhance the feature extractor. By leveraging the powerful representation capabilities of a pre-trained model, our method further refines them for counting tasks, notably improving network performance on limited remote sensing data. Specifically, we design a lightweight cell-wise architecture optimization based on a network architecture search algorithm. It sequentially builds lightweight cells in parallel with the blocks in a pre-trained model, leveraging a newly proposed random path selection strategy for training the optimization framework. Additionally, we propose extracting high-frequency information from the blocks in a pre-trained model as self-guidance optimization to facilitate the learning of the searched cells. Experimental results on several representative datasets demonstrate that our proposed method significantly improves counting performance while substantially reducing parameters. For instance, compared with our baseline, it improves performance by 22.2% in MAE and 17.0% in MSE on the Building dataset, while reducing parameters by 61.2%. Shengqin Jiang, Qian Jie, Fengna Cheng, Haokui Zhang, Yu Liu 0029, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Spatial-Temporal Interleaved Network for Efficient Action RecognitionabstractThe decomposition of 3D convolution will considerably reduce the computing complexity of 3D convolutional neural networks, yet simple stacking restricts the performance of neural networks. To this end, we propose a spatial-temporal interleaved network for efficient action recognition. By deeply analyzing this task, it revisits the structure of 3D neural networks in action recognition from the following perspectives. To enhance the learning of robust spatial-temporal features, we initially propose an interleaved feature interaction module to comprehensively explore cross-layer features and capture the most discriminative information among them. With regards to being lightweight, a boosted parallel pseudo-3D module is introduced with the goal of circumventing a substantial number of computations from the lower to middle levels while enhancing temporal and spatial features in parallel at high levels. Furthermore, we exploit a spatial-temporal differential attention mechanism to suppress redundant features in different dimensions while reaping the benefits of nearly negligible parameters. Lastly, extensive experiments on four action recognition benchmarks are given to show the advantages and efficiency of our proposed method. Specifically, our method attains a 15.2% improvement in Top-1 accuracy compared to our baseline, a stack of full 3D convolutional layers, on the Something-Something V1 dataset while utilizing only 18.2% of the parameters. Shengqin Jiang, Haokui Zhang, Yuankai Qi, Qingshan Liu 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Efficient Adaptation of Pre-trained Vision Transformer via Householder TransformationabstractA common strategy for Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViTs) involves adapting the model to downstream tasks by learning a low-rank adaptation matrix. This matrix is decomposed into a product of down-projection and up-projection matrices, with the bottleneck dimensionality being crucial for reducing the number of learnable parameters, as exemplified by prevalent methods like LoRA and Adapter. However, these low-rank strategies typically employ a fixed bottleneck dimensionality, which limits their flexibility in handling layer-wise variations. To address this limitation, we propose a novel PEFT approach inspired by Singular Value Decomposition (SVD) for representing the adaptation matrix. SVD decomposes a matrix into the product of a left unitary matrix, a diagonal matrix of scaling values, and a right unitary matrix. We utilize Householder transformations to construct orthogonal matrices that efficiently mimic the unitary matrices, requiring only a vector. The diagonal values are learned in a layer-wise manner, allowing them to flexibly capture the unique properties of each layer. This approach enables the generation of adaptation matrices with varying ranks across different layers, providing greater flexibility in adapting pre-trained models. Experiments on standard downstream vision tasks demonstrate that our method achieves promising fine-tuning performance. Wei Dong 0010, Yiting Yang, Zhijun Lin, Qingsen Yan, Haokui Zhang, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
NeurIPS | 7 |
| 2024 | Tripartite-structure transformer for hyperspectral image classificationabstractAbstract Hyperspectral images contain rich spatial and spectral information, which provides a strong basis for distinguishing different land‐cover objects. Therefore, hyperspectral image (HSI) classification has been a hot research topic. With the advent of deep learning, convolutional neural networks (CNNs) have become a popular method for hyperspectral image classification. However, convolutional neural network (CNN) has strong local feature extraction ability but cannot deal with long‐distance dependence well. Vision Transformer (ViT) is a recent development that can address this limitation, but it is not effective in extracting local features and has low computational efficiency. To overcome these drawbacks, we propose a hybrid classification network that combines the strengths of both CNN and ViT, names Spatial‐Spectral Former(SSF). The shallow layer employs 3D convolution to extract local features and reduce data dimensions. The deep layer employs a spectral‐spatial transformer module for global feature extraction and information enhancement in spectral and spatial dimensions. Our proposed model achieves promising results on widely used public HSI datasets compared to other deep learning methods, including CNN, ViT, and hybrid models. Liuwei Wan, Meili Zhou, Shengqin Jiang, Zongwen Bai, Haokui Zhang |
Comput. Intell. | 5 |
| 2024 | Compare and Focus: Multi-Scale View Aggregation for Crowd CountingabstractRecently, some state-of-the-art (SOTA) methods have designed dedicated context extractors to capture the global information that serves as a key clue for describing crowd density. A promising alternative is the transformer-based model which inherently captures long-range context dependencies. Recent related studies have made impressive progress, yet the following issues remain: (1) The size of the heads in the image is large near and small far away. The existing models fail to cope well with these variations. (2) There is an imbalance in the distribution of samples across different densities in the dataset, which leads to poor network performance on density distributions with a small number of samples. To address these issues, we propose to aggregate multi-scale views through Compare and Focus strategies. In terms of the first strategy, we mine differential hints from multi-scale view features to capture heads of varying sizes. This can effectively reduce the influence of redundant information while perceiving the subtleties of various view inputs, making it simpler to establish discriminative representations. As for the second strategy, we introduce a new activation function to formulate the Region of Interest (ROI) extraction module that enables the network to focus on relevant regions effectively. It can alleviate the extreme distribution imbalance of samples with different densities. Finally, several experiments show that our method achieves SOTA performance on four challenging datasets. Shengqin Jiang, Jialu Cai, Haokui Zhang, Yu Liu 0029, Qingshan Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | NAR-Former: Neural Architecture Representation Learning Towards Holistic Attributes PredictionabstractWith the wide and deep adoption of deep learning models in real applications, there is an increasing need to model and learn the representations of the neural networks themselves. These models can be used to estimate attributes of different neural network architectures such as the accuracy and latency, without running the actual training or inference tasks. In this paper, we propose a neural architecture representation model that can be used to estimate these attributes holistically. Specifically, we first propose a simple and effective tokenizer to encode both the operation and topology information of a neural network into a single sequence. Then, we design a multi-stage fusion transformer to build a compact vector representation from the converted sequence. For efficient model training, we further propose an information flow consistency augmentation and correspondingly design an architecture consistency loss, which brings more benefits with less augmentation samples compared with previous random augmentation strategies. Experiment results on NAS-Bench-101, NAS-Bench-201, DARTS search space and NNLQP show that our proposed framework can be used to predict the aforementioned latency and accuracy attributes of both cell architectures and whole deep neural networks, and achieves promising performance. Code is available at https://github.com/yuny220/NAR-Former. Yun Yi, Haokui Zhang, Wenze Hu, Nannan Wang 0001, Xiaoyu Wang 0002 |
CVPR | 2 |
| 2023 | ParCNetV2: Oversized Kernel with Enhanced Attention*abstractTransformers have shown great potential in various computer vision tasks. By borrowing design concepts from transformers, many studies revolutionized CNNs and showed remarkable results. This paper falls in this line of studies. Specifically, we propose a new convolutional neural network, ParCNetV2, that extends the research line of ParCNetV1 by bridging the gap between CNN and ViT. It introduces two key designs: 1) Oversized Convolution (OC) with twice the size of the input, and 2) Bifurcate Gate Unit (BGU) to ensure that the model is input adaptive. Fusing OC and BGU in a unified CNN, ParCNetV2 is capable of flexibly extracting global features like ViT, while maintaining lower latency and better accuracy. Extensive experiments demonstrate the superiority of our method over other convolutional neural networks and hybrid models that combine CNNs and transformers. The code are publicly available at https://github.com/XuRuihan/ParCNetV2. Ruihan Xu 0002, Haokui Zhang, Wenze Hu, Shiliang Zhang, Xiaoyu Wang 0002 |
ICCV | 2 |
| 2023 | Fcaformer: Forward Cross Attention in Hybrid Vision TransformerabstractCurrently, one main research line in designing a more efficient vision transformer is reducing the computational cost of self attention modules by adopting sparse attention or using local attention windows. In contrast, we propose a different approach that aims to improve the performance of transformer-based architectures by densifying the attention pattern. Specifically, we proposed forward cross attention for hybrid vision transformer (FcaFormer), where tokens from previous blocks in the same stage are secondary used. To achieve this, the FcaFormer leverages two innovative components: learnable scale factors (LSFs) and a token merge and enhancement module (TME). The LSFs enable efficient processing of cross tokens, while the TME generates representative cross tokens. By integrating these components, the proposed FcaFormer enhances the interactions of tokens across blocks with potentially different semantics, and encourages more information flows to the lower levels. Based on the forward cross attention (Fca), we have designed a series of FcaFormer models that achieve the best trade-off between model size, computational cost, memory cost, and accuracy. For example, without the need for knowledge distillation to strengthen training, our FcaFormer achieves 83.1% top-1 accuracy on Imagenet with only 16.3 million parameters and about 3.6 billion MACs. This saves almost half of the parameters and a few computational costs while achieving 0.7% higher accuracy compared to distilled EfficientFormer. Code is available at https://github.com/hkzhang-git/FcaFormer Haokui Zhang, Wenze Hu, Xiaoyu Wang 0002 |
ICCV | 1 |
| 2023 | NAR-Former V2: Rethinking Transformer for Universal Neural Network Representation LearningabstractAs more deep learning models are being applied in real-world applications, there is a growing need for modeling and learning the representations of neural networks themselves. An effective representation can be used to predict target attributes of networks without the need for actual training and deployment procedures, facilitating efficient network design and deployment. Recently, inspired by the success of Transformer, some Transformer-based representation learning frameworks have been proposed and achieved promising performance in handling cell-structured models. However, graph neural network (GNN) based approaches still dominate the field of learning representation for the entire network. In this paper, we revisit the Transformer and compare it with GNN to analyze their different architectural characteristics. We then propose a modified Transformer-based universal neural network representation learning model NAR-Former V2. It can learn efficient representations from both cell-structured networks and entire networks. Specifically, we first take the network as a graph and design a straightforward tokenizer to encode the network into a sequence. Then, we incorporate the inductive representation learning capability of GNN into Transformer, enabling Transformer to generalize better when encountering unseen architecture. Additionally, we introduce a series of simple yet effective modifications to enhance the ability of the Transformer in learning representation from graph structures. In encoding entire networks and then predicting the latency, our proposed method surpasses the GNN-based method NNLP by a significant margin on the NNLQP dataset. Furthermore, regarding accuracy prediction on the cell-structured NASBench101 and NASBench201 datasets, our method achieves highly comparable performance to other state-of-the-art methods. The code is available at https://github.com/yuny220/NAR-Former-V2. Yun Yi, Haokui Zhang, Rong Xiao 0003, Nannan Wang 0001, Xiaoyu Wang 0002 |
NeurIPS | 2 |
| 2023 | Adapt-Infomap: Face clustering with adaptive graph refinement in infomapabstractFace clustering is a critical task in computer vision due to the increasing number of applications such as augmented reality or photo album management. The primary challenge in this task arises from the imperfections in image feature representations. Given image features extracted from an existing pre-trained representation model, it remains an unresolved problem that how to leverage the inherent characteristics of similarities among unlabelled images to improve the clustering performance. In order to solve face clustering in an unsupervised manner , we develop an effective and robust framework named as Adapt-Infomap. First, we reformulate face clustering as a process of non-overlapping community detection. Specially, Adapt-Infomap achieves face clustering by minimizing the entropy of information flows (also known as the map equation) on an affinity graph of images. Since the affinity graph of images might contain noisy edges, we develop an outlier detection strategy in Adapt-Infomap to adaptively refine the affinity graph. Experiments with ablation studies demonstrate that Adapt-Infomap significantly outperforms existing methods and achieves new state-of-the-arts on three popular large-scale datasets for face clustering, e.g. , an absolute improvement of more than 10 % and 3 % comparing with prior unsupervised and supervised methods respectively in terms of average of Pairwise F-score. Xiaotian Yu, Aibo Wang, Haokui Zhang, Hanling Yi, Guangming Lu 0002, Xiaoyu Wang 0002 |
Pattern Recognit. | 5 |
| 2022 | ParC-Net: Position Aware Circular Convolution with Merits from ConvNets and Transformer
Haokui Zhang, Wenze Hu, Xiaoyu Wang 0002 |
ECCV (26) | 1 |
| 2022 | Connecting Compression Spaces with Transformer for Approximate Nearest Neighbor Search
Haokui Zhang, Buzhou Tang, Wenze Hu, Xiaoyu Wang 0002 |
ECCV (14) | 1 |
| 2022 | Memory-Efficient Hierarchical Neural Architecture Search for Image Restoration
Haokui Zhang, Ying Li 0017, Hao Chen 0041, Chengrong Gong, Zongwen Bai, Chunhua Shen |
Int. J. Comput. Vis. | 1 |
| 2022 | Pseudo-LiDAR-Based Road DetectionabstractRoad detection is a critically important task for self-driving cars. By employing LiDAR data, recent works have significantly improved the accuracy of road detection. However, relying on LiDAR sensors limits the application of those methods when only cameras are available. In this paper, we propose a novel road detection approach with RGB images being the only input. Specifically, we exploit pseudo-LiDAR using depth estimation and propose a feature fusion network in which RGB images and learned depth information are fused for improved road detection. To optimize the network architecture and improve the efficiency of our network, we propose a method to search for the information propagation paths. Finally, to reduce the computational cost, we design a modality distillation strategy to avoid using depth estimation networks during inference. The resulting model eliminates the reliance on LiDAR sensors and achieves state-of-the-art performance on two challenging benchmarks, KITTI and R2D. Libo Sun 0002, Haokui Zhang, Wei Yin 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Grafting Transformer on Automatically Designed Convolutional Neural Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification has been a hot topic for decides, as hyperspectral images have rich spatial and spectral information and provide strong basis for distinguishing different land-cover objects. Benefiting from the development of deep learning technologies, deep learning based HSI classification methods have achieved promising performance. Recently, several neural architecture search (NAS) algorithms have been proposed for HSI classification, which further improve the accuracy of HSI classification to a new level. In this paper, NAS and Transformer are combined for handling HSI classification task for the first time. Compared with previous work, the proposed method has two main differences. First, we revisit the search spaces designed in previous HSI classification NAS methods and propose a novel hybrid search space, consisting of the space dominated cell and the spectrum dominated cell. Compared with search spaces proposed in previous works, the proposed hybrid search space is more aligned with the characteristic of HSI data, that is, HSIs have a relatively low spatial resolution and an extremely high spectral resolution. Second, to further improve the classification accuracy, we attempt to graft the emerging transformer module on the automatically designed convolutional neural network (CNN) to add global information to local region focused features learned by CNN. Experimental results on three public HSI datasets show that the proposed method achieves much better performance than comparison approaches, including manually designed network and NAS based HSI classification methods. Especially on the most recently captured dataset Houston University, overall accuracy is improved by nearly 6 percentage points. Code is available at: https://github.com/Cecilia-xue/HyT-NAS. Xizhe Xue, Haokui Zhang, Bei Fang, Zongwen Bai, Ying Li 0017 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | 3-D-ANAS: 3-D Asymmetric Neural Architecture Search for Fast Hyperspectral Image ClassificationabstractHyperspectral images (HSIs) provide abundant spectral and spatial information, playing an irreplaceable role in land-cover classification. Recently, based on deep learning (DL) technologies, an increasing number of HSI classification approaches have been proposed, which demonstrate promising performance. However, previous studies suffer from two major drawbacks: 1) the architecture of most DL models is manually designed, relies on specialized knowledge, and is relatively tedious. Moreover, in HSI classifications, datasets captured by different sensors have different physical properties. Correspondingly, different models need to be designed for different datasets, which further increases the workload of designing architectures and 2) the mainstream framework is a patch-to-pixel framework. The overlap regions of patches of adjacent pixels are calculated repeatedly, which increases computational cost and time cost. In addition, the classification accuracy is sensitive to the patch size, which is artificially set based on extensive investigation experiments. To overcome the issues mentioned above, we first propose a 3-D asymmetric neural network search algorithm and leverage it to automatically search for efficient architectures for HSI classifications. By analyzing the characteristics of HSIs, we specifically build a 3-D asymmetric decomposition search space, where spectral and spatial information is processed with different decomposition convolutions. Furthermore, we propose a new fast classification framework, i.e., pixel-to-pixel classification framework, which has no repetitive operations and reduces the overall cost. Experiments on three public HSI datasets captured by different sensors demonstrate the networks designed by our 3-D asymmetric neural architecture search (3-D-ANAS) achieve competitive performance compared to several state-of-the-art methods, while having a much faster inference speed. Code is available at:https://github.com/hkzhang91/3D-ANAS. Haokui Zhang, Chengrong Gong, Yunpeng Bai, Zongwen Bai, Ying Li 0017 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Learning Structure Affinity for Video Depth EstimationabstractDepth estimation is a structure learning problem. The affinity among neighbouring pixels plays an important role in inferring depth values. In this paper, we propose to learn structure affinity in both spatial and temporal domain for accurate depth estimation from monocular videos. Specifically, we first propose a convolutional spatial temporal propagation network (CSTPN) that learns affinity among neighbouring video frames. Secondly, we employ a structure knowledge distillation scheme that transfers the spatial temporal affinity learned by cumbersome network to compact network. By calculating pixel-wise similarities between neighboring frames and neighbouring sequences, our knowledge distillation scheme efficiently captures both short-term and long-term spatial temporal affinity. Finally, we apply a warping loss based on optical flow between video frames to further enforce the temporal affinity. Experiment results show that our proposed depth estimation approach outperform the state-of-the-art methods on both indoor and outdoor benchmark datasets. Yuanzhouhan Cao, Yidong Li, Haokui Zhang, Chao Ren 0002, Yifan Liu 0001 |
ACM Multimedia | 3 |
| 2021 | Hyperspectral Image Classification With Spatial Consistence Using Fully Convolutional Spatial Propagation NetworkabstractIn recent years, deep convolutional neural networks (CNNs) have demonstrated impressive ability to represent hyperspectral images (HSIs) and achieved encouraging results in HSI classification. However, the existing CNN-based models operate at the patch level, in which a pixel is separately classified into classes using a patch of images around it. This patch-level classification will lead to a large number of repeated calculations, and it is hard to identify the appropriate patch size that is beneficial to classification accuracy. In addition, the conventional CNN models operate convolutions with local receptive fields, which cause the failure of contextual spatial information modeling. To overcome these aforementioned limitations, we propose a novel end-to-end, pixel-to-pixel, fully convolutional spatial propagation network (FCSPN) for HSI classification. Our FCSPN consists of a 3-D fully convolution network (3D-FCN) and a convolutional spatial propagation network (CSPN). Specifically, the 3D-FCN is first introduced for reliable preliminary classification, in which a novel dual separable residual (DSR) unit is proposed to effectively capture spectral and spatial information simultaneously with fewer parameters. Moreover, the channel-wise attention mechanism is adapted in the 3D-FCN to grasp the most informative channels from redundant channel information. Finally, the CSPN is introduced to capture the spatial correlations of HSIs via learning a local linear spatial propagation, which allows maintaining the HSI spatial consistency and further refining the classification results. Experimental results on three HSI benchmark data sets demonstrate that the proposed FCSPN achieves state-of-the-art performance on HSI classification. Yenan Jiang, Ying Li 0017, Shanrong Zou, Haokui Zhang, Yunpeng Bai |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | D3D: Dual 3-D Convolutional Network for Real-Time Action RecognitionabstractThree-dimensional convolutional neural networks (3D CNNs) have been explored to learn spatio-temporal information for video-based human action recognition. Expensive computational cost and memory demand resulted from standard 3D CNNs, however, hinder their application in practical scenarios. In this article, we address the aforementioned limitations by proposing a novel dual 3-D convolutional network (D3DNet) with two complementary lightweight branches. A coarse branch maintains large temporal receptive field by a fast temporal downsampling strategy and simulates the expensive 3-D convolutions using a combination of more efficient spatial convolutions and temporal convolutions. Meanwhile, a fine branch progressively downsamples the video in the temporal domain and adopts 3-D convolutional units with reduced channel capacities to capture multiresolution spatio-temporal information. Instead of learning these two branches independently, a shallow spatiotemporal downsampling module is shared for these two branches for efficient low-level feature learning. Besides, lateral connections are learned to effectively fuse the information from the two branches at multiple stages. The proposed network makes good balance between inference speed and action recognition performance. Based on RGB information only, it achieves competing performance on five popular video-based action recognition datasets, with inference speed of 3200 FPS on a single NVIDIA GTX 2080Ti card. Shengqin Jiang, Yuankai Qi, Haokui Zhang, Zongwen Bai, Xiaobo Lu, Peng Wang 0023 |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Memory-Efficient Hierarchical Neural Architecture Search for Image DenoisingabstractRecently, neural architecture search (NAS) methods have attracted much attention and outperformed manually designed architectures on a few high-level vision tasks. In this paper, we propose HiNAS (Hierarchical NAS), an effort towards employing NAS to automatically design effective neural network architectures for image denoising. HiNAS adopts gradient based search strategies and employs operations with adaptive receptive field to build an flexible hierarchical search space. During the search stage, HiNAS shares cells across different feature levels to save memory and employ an early stopping strategy to avoid the collapse issue in NAS, and considerably accelerate the search speed. The proposed HiNAS is both memory and computation efficient, which takes only about 4.5 hours for searching using a single GPU. We evaluate the effectiveness of our proposed HiNAS on two different datasets, namely an additive white Gaussian noise dataset BSD500, and a realistic noise dataset SIM1800. Experimental results show that the architecture found by HiNAS has fewer parameters and enjoys a faster inference speed, while achieving highly competitive performance compared with state-of-the-art methods. We also present analysis on the architectures found by NAS. HiNAS also shows good performance on experiments for image de-raining. Haokui Zhang, Ying Li 0017, Hao Chen 0041, Chunhua Shen |
CVPR | 1 |
| 2020 | Meta Learning with Differentiable Closed-form Solver for Fast Video Object SegmentationabstractVideo object segmentation plays a vital role to many robotic tasks, beyond the satisfied accuracy, quickly adapt to the new scenario with very limited annotations and conduct a quick inference are also important. In this paper, we are specifically concerned with the task of fast segmenting all pixels of a target object in all frames, given the annotation mask in the first frame. Even when such annotation is available, this remains a challenging problem because of the changing appearance and shape of the object over time. In this paper, we tackle this task by formulating it as a meta-learning problem, where the base learner grasping the semantic scene understanding for a general type of objects, and the meta learner quickly adapting the appearance of the target object with a few examples. Our proposed meta-learning method uses a closed form optimizer, the so-called "ridge regression", which has been shown to be conducive for fast and better training convergence. Moreover, we propose a mechanism, named "block splitting", to further speed up the training process as well as to reduce the number of learning parameters. In comparison with the state-of-the art methods, our proposed framework achieves significant boost up in processing speed, while having highly comparable performance compared to the best performing methods on the widely used datasets. Video demo can be found here1. Yu Liu 0029, Lingqiao Liu, Haokui Zhang, Seyed Hamid Rezatofighi, Qingsen Yan, Ian D. Reid 0001 |
IROS | 3 |
| 2019 | Exploiting Temporal Consistency for Real-Time Video Depth EstimationabstractAccuracy of depth estimation from static images has been significantly improved recently, by exploiting hierarchical features from deep convolutional neural networks (CNNs). Compared with static images, vast information exists among video frames and can be exploited to improve the depth estimation performance. In this work, we focus on exploring temporal information from monocular videos for depth estimation. Specifically, we take the advantage of convolutional long short-term memory (CLSTM) and propose a novel spatial-temporal CSLTM (ST-CLSTM) structure. Our ST-CLSTM structure can capture not only the spatial features but also the temporal correlations/consistency among consecutive video frames with negligible increase in computational cost. Additionally, in order to maintain the temporal consistency among the estimated depth frames, we apply the generative adversarial learning scheme and design a temporal consistency loss. The temporal consistency loss is combined with the spatial loss to update the model in an end-to-end fashion. By taking advantage of the temporal information, we build a video depth estimation framework that runs in real-time and generates visually pleasant results. Moreover, our approach is flexible and can be generalized to most existing depth estimation frameworks. Code is available at: https://tinyurl.com/STCLSTM Haokui Zhang, Ying Li 0017, Yuanzhouhan Cao, Yu Liu 0029, Chunhua Shen, Youliang Yan |
ICCV | 1 |
| 2019 | Hyperspectral Image Classification Based on 3-D Separable ResNet and Transfer LearningabstractDeep learning (DL) has proven to be a promising technique for hyperspectral image (HSI) classification. However, due to complex network structure and massive parameters, it is challenging to achieve satisfying classification accuracy with only a small number of training samples. In this letter, we propose a novel HSI classification method by collaborating the 3-D separable ResNet (3-D-SRNet) with cross-sensor transfer learning. The 3-D-SRNet replaces 3-D convolutions with spatial and spectral separable 3-D convolutions, thus showing much less parameters than models that use standard 3-D convolutions. First, we pretrain a classification model with the proposed 3-D-SRNet on the source HSI data set with sufficient training samples compared with the target HSI data set. Then, the pretrained model is transferred to the target HSI data set for fine-tuning to finish the classification task. It is worth noting that the source data for pretraining can be captured by the different sensor with the target data. Compared with the conventional 3-D-ResNet, the proposed 3-D-SRNet has less parameters involving lower computation cost while achieving better classification performance. Experimental results on three benchmark data sets show that our method outperforms several state-of-the-art methods in HSI classification with small training samples. Yenan Jiang, Ying Li 0017, Haokui Zhang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | Hyperspectral Classification Based on Lightweight 3-D-CNN With Transfer LearningabstractRecently, hyperspectral image (HSI) classification approaches based on deep learning (DL) models have been proposed and shown promising performance. However, because of very limited available training samples and massive model parameters, DL methods may suffer from overfitting. In this paper, we propose an end-to-end 3-D lightweight convolutional neural network (CNN) (abbreviated as 3-D-LWNet) for limited samples-based HSI classification. Compared with conventional 3-D-CNN models, the proposed 3-D-LWNet has a deeper network structure, less parameters, and lower computation cost, resulting in better classification performance. To further alleviate the small sample problem, we also propose two transfer learning strategies: 1) cross-sensor strategy, in which we pretrain a 3-D model in the source HSI data sets containing a greater number of labeled samples and then transfer it to the target HSI data sets and 2) cross-modal strategy, in which we pretrain a 3-D model in the 2-D RGB image data sets containing a large number of samples and then transfer it to the target HSI data sets. In contrast to previous approaches, we do not impose restrictions over the source data sets, in which they do not have to be collected by the same sensors as the target data sets. Experiments on three public HSI data sets captured by different sensors demonstrate that our model achieves competitive performance for HSI classification compared to several state-of-the-art methods. Haokui Zhang, Ying Li 0017, Yenan Jiang, Peng Wang 0015, Qiang Shen 0001, Chunhua Shen |
IEEE Trans. Geosci. Remote. Sens. | 1 |