EDBT 2026 Demo / reviewers in the wild / expert
Junjie Zhang 0002
dblp:99/6243-2
· DBLP profile ↗
47ranked-venue papers
13as first author
41since 2021 · last 2026
0000-0002-0033-0494ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 9 first-author · 26 since 2021Artificial intelligence and machine learning · 15 · 6 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FTTrack: RGB-T Tracking With Frequency-Adaptive Fusion and Temporal Enhancement
Yutong Gu, Zichun Zhou, Hongwen Yu, Junjie Zhang 0002 |
IEEE Signal Process. Lett. | 5 |
| 2026 | Learning Compact Representations With an Information Bottleneck for Camouflaged Object DetectionabstractFrequency domain-based methods have demonstrated promising performance in Camouflaged Object Detection (COD) tasks because of their enhanced power for distinguishing between objects and the background in the frequency domain. However, these methods often overlook the interference caused by task-irrelevant cues such as background textures. These extraneous factors are learned alongside task-relevant features by the employed network, increasing the number of false positives. Therefore, we propose a camouflaged object detection method based on the Information Bottleneck (IB) theory. The aim is to obtain a robust representation that retains the essential features needed for prediction while minimizing the redundant information derived from both the RGB and frequency domains. Specifically, we propose a Feature Selection Information Bottleneck Module (FSIBM). By explicit supervision, this module minimizes the mutual information between the fused feature from two domains and the predictive features, thereby weakening task-irrelated information. Simultaneously, the FSIBM maximizes the mutual information between the predictive features and the ground truth (i.e., emphasizing task-related elements). Additionally, we introduce a Cross-Domain Awareness Interaction Module (CDAIM), which establishes self-reinforcement for the object attributes within each domain and facilitates cross-domain complementarity. This enables the capture of sufficient discriminative features from both domains. To verify the generalization ability of the proposed method, we applied it to three benchmark datasets, on which our method outperformed the corresponding state-of-the-art methods. Our code is released athttps://github.com/KwunYat/CODIB. Guanyi Li, Junjie Zhang 0002, Wubang Yuan, Gloria Jin, Dan Zeng 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Multi-Modal Multi-Platform Person Re-Identification: Benchmark and Method
Ruiyang Ha, Songyi Jiang, Bikang Pan, Yihang Zhu, Junjie Zhang 0002, Xiatian Zhu, Shaogang Gong, Jingya Wang 0001 |
ICCV | 6 |
| 2025 | CostDiff: Residual Diffusion-Based Cost Map Refinement for Open-Vocabulary Semantic Segmentation
Yutao Rao, Fangyu Wu 0001, Junjie Zhang 0002 |
PRCV (12) | 4 |
| 2025 | Enhancing origin-destination flow prediction via bi-directional spatio-temporal inference and interconnected feature evolution
Piao Yu, Xu Zhang 0039, Yongshun Gong, Jian Zhang 0002, Haoliang Sun, Junjie Zhang 0002, Xinxin Zhang 0004, Yilong Yin |
Expert Syst. Appl. | 6 |
| 2025 | Fine-grained visual tracking via distribution-aware mask modeling and temporal propagation
Junjie Zhang 0002, Hongwen Yu, Fangyu Wu 0001, Xiaoshui Huang, Jian Zhang 0002 |
Knowl. Based Syst. | 2 |
| 2025 | Establishing Nuanced Multimodal Attention for Weakly Supervised Semantic Segmentation of Remote Sensing ScenesabstractWeakly Supervised Semantic Segmentation (WSSS) with image-level labels reduces reliance on pixel-level annotations for remote sensing (RS) imagery. However, in natural scenes, WSSS frequently faces challenges such as imprecise localization, extraneous activations, and class ambiguity. These challenges are particularly pronounced in RS images, characterized by complex backgrounds, substantial scale variations, and dense small-object distributions, complicating the distinction between intra-class variations and inter-class similarities. To tackle these challenges, we introduce a class-constrained multi-modal attention framework aimed at enhancing the localization accuracy of class activation maps (CAMs). Specifically, we design class-specific tokens to capture the visual characteristics of each target class. As these tokens initially lack explicit constraints, we integrate the textual branch of the RemoteCLIP model to leverage class-related linguistic priors, which collaborate with visual features to encode the specific semantics of diverse objects. Furthermore, the multi-modal collaborative optimization module dynamically establishes tailored attention mechanisms for both global and regional features, thereby improving class discriminability among targets to mitigate challenges like inter-class similarity and dense small-object distributions. By refining class-specific attention, textual semantic attention, and patch-level pairwise affinity weights, the quality of generated pseudo-masks is markedly enhanced. Concurrently, to ensure domain-invariant feature learning, we align the backbone features with the CLIP visual embedding by minimizing the distribution disparity between the two in the latent space, semantic consistency is therefore preserved. The experimental results validate the effectiveness and robustness of our proposed method, achieving significant performance improvements on two representative RS WSSS datasets. Junjie Zhang 0002, Huaxi Huang, Fangyu Wu 0001, Hongwen Yu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | 3DBench: A scalable benchmark for object and scene-level instruction-tuning of 3D large language models
Tianci Hu, Junjie Zhang 0002, Yutao Rao, Dan Zeng 0001, Hongwen Yu, Xiaoshui Huang |
Neural Networks | 2 |
| 2025 | DSENet++: A Coarse-to-Fine Framework for Enhanced Sub-Region Detection in Aerial Images
Xiangjie Wang, Liang Chen 0004, Junjie Zhang 0002, Jian Zhang 0002, Shiming Ge, Dan Zeng 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Precision in pursuit: a multi-consistency joint approach for infrared anti-UAV tracking
Junjie Zhang 0002, Pangrong Shi, Xiaoqiang Zhu, Dan Zeng 0001 |
Vis. Comput. | 1 |
| 2024 | Fine-Grained Urban Flow Inference with Dynamic Multi-scale Representation Learning
Shilu Yuan, Wei Liu 0007, Xinxin Zhang 0004, Meng Chen 0003, Junjie Zhang 0002, Yongshun Gong |
DASFAA (2) | 6 |
| 2024 | DSENet: An Object-Wise Density-Informed Coarse-to-Fine Object Detector for Aerial ImageabstractObject detection in aerial images remains formidable due to substantial object scale variations, and uneven object distributions. Previous methods widely adopt the coarse-to-fine methodology where detectors focus on large-scale objects coarsely. Sub-regions that contain densely distributed small ones are captured and detected finely. However, two pivotal assessment factors of sub-regions, positional precision, and detection difficulty, deserve further consideration. In this paper, we propose an object-wise density-informed DSENet including consecutive stages termed "Discernment, Selection, Elevation ". Specifically, the sophisticated object-wise density map that considers both object scales and angles, helps discern more positional-precise sub-regions. Then sub-regions with high detection difficulty are selected based on density intensities and coarse detections collaboratively. Finally, the fine detector head instead of the full detector, fine-tuned with selected sub-regions efficiently, elevates what and where coarse detections are mediocre. Extensive experiments show that DSENet achieves state-of-the-art performance on two popular aerial image datasets, VisDrone and DOTA-V1.5. Xiangjie Wang, Junjie Zhang 0002, Jian Zhang 0002, Dan Zeng 0001 |
ICME | 3 |
| 2024 | Densely Connected Transformer with Frequency Awareness and Sam Guidance for Semi-Supervised Hyperspectral Image ClassificationabstractAdvancements in Hyperspectral Image (HSI) spatial resolution pose challenges in pixel-wise classification. Semi-supervised self-training shows potential by using pseudo-labels from unlabeled samples. However, the Hughes phenomenon and environmental factors often lead to spectral variability and undermine pseudo-label credibility. To address above issues, we propose a densely connected Transformer leveraging Discrete Wavelet Transform for extracting nuanced spatial-spectral features and redundancy removal, and we design a filtering strategy guided by the Segment Anything Model (SAM) to retain reliable pseudo labeled samples given the spatial and semantic consistency of HSI regions. Experiments show promising performance of proposed model on high-resolution HSIs compared to trending methods under limited supervision. Yutao Rao, Liwei Sun, Junjie Zhang 0002, Jian Zhang 0002, Dan Zeng 0001 |
ICME | 3 |
| 2024 | 3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset
Junjie Zhang 0002, Tianci Hu, Xiaoshui Huang, Yongshun Gong, Dan Zeng 0001 |
IJCAI | 1 |
| 2024 | A streamlined framework for BEV-based 3D object detection with prior masking
Qinglin Tong, Junjie Zhang 0002, Chenggang Yan 0001, Dan Zeng 0001 |
Image Vis. Comput. | 2 |
| 2024 | TMSDNet: Transformer with multi-scale dense network for single and multi-view 3D reconstructionabstractAbstract 3D reconstruction is a long‐standing problem. Recently, a number of studies have emerged that utilize transformers for 3D reconstruction, and these approaches have demonstrated strong performance. However, transformer‐based 3D reconstruction methods tend to establish the transformation relationship between the 2D image and the 3D voxel space directly using transformers or rely solely on the powerful feature extraction capabilities of transformers. They ignore the crucial role played by deep multi‐scale representation of the object in the voxel feature domain, which can provide extensive global shape and local detail information about the object in a multi‐scale manner. In this article, we propose a novel framework TMSDNet (transformer with multi‐scale dense network) for single‐view and multi‐view 3D reconstruction with transformer to solve this problem. Based on our well‐designed combined‐transformer Block, which is canonical encoder–decoder architecture, voxel features with spatial order can be extracted from the input image, which are used to further extract multi‐scale global features in parallel using a multi‐scale residual attention module. Furthermore, a residual dense attention block is introduced for deep local features extraction and adaptive fusion. Finally, the reconstructed objects are produced with the voxel reconstruction block. Experiment results on the benchmarks such as ShapeNet and Pix3D datasets demonstrate that TMSDNet outperforms the existing state‐of‐the‐art reconstruction methods substantially. Xiaoqiang Zhu, Xinsheng Yao, Junjie Zhang 0002, Lihua You, Xiaosong Yang, Jian J. Zhang 0001, Dan Zeng 0001 |
Comput. Animat. Virtual Worlds | 3 |
| 2024 | DCTracker: Rethinking MOT in soccer events under dual views via cascade association
Long Hu, Junjie Zhang 0002, Weiyi Lv, Yongshun Gong, Jingya Wang 0001, Jian Zhang 0002, Dan Zeng 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Leveraging Frequency-Guided Mixer and Target-Aware Attention for Ground-Based Cloud DetectionabstractCompared to satellite imagery, ground-based cameras capture cloud data (ground-to-sky data) with higher temporal and spatial resolutions, providing more detailed cloud information. However, the spectral information available in ground-to-sky data is limited. Therefore, extracting features with strong discrimination from optical remote sensing images (ORSIs) is challenging. Currently, deep learning-based cloud detection methods face two main challenges. Firstly, although Convolutional Neural Networks (CNNs) effectively extract high-frequency (HF) components from images through convolutions, they struggle to capture low-frequency (LF) components, which are capable of representing global features and target structures. Secondly, in ORSIs, the spectral characteristics of thin clouds and the sky are similar, making it difficult to distinguish cloud regions from the background. To address these challenges, we propose a network consisting of two main modules: the Mixer Module (MM) and the Cloud Aware Attention Module (CAAM). The MM comprises a HF and a LF components extraction branch. The HF branch extracts local textures through max-pooling and parallel convolution operations. The LF branch captures long-range dependency by decomposing a large kernel convolution. It leverages the advantages of both convolution and self-attention to effectively capture global features. In addition, we introduce the CAAM, which quantifies images into histograms to separate clouds from the background and enhances the perception of clouds using attention mechanism. We conducted experiments using both daytime and nighttime cloud image data from the SWINySeg dataset with mIoU reaching 88.93% and OA reaching 93.97%. The results demonstrate that our proposed method achieves promising performance compared to state-of-the-art cloud detection methods. Chenyu Dong, Guanyi Li, Yixiao Gu, Junjie Zhang 0002, Dan Zeng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Multi-Level Information Fusion Network With Edge Information Injection for Single-Band Cloud DetectionabstractCurrent cloud detection methods have demonstrated effectiveness by utilizing the rich spectral features of multi-spectral images. Compared to multispectral images, single-band infrared images offer higher efficiency in terms of sampling and processing speed. However, single-band cloud detection methods have not been fully developed, and existing methods based on multispectral cloud detection have some limitations when applied directly to single-band images: Firstly, they often blend shallow features containing spatial details with deep features providing high-level semantic information, yet struggle to disentangle features with strong discrimination representing cloud edges and bodies from limited information. Additionally, the correlation between features at different aspects is not fully reasoned, resulting in blurred boundary segmentation. To address these issues, we introduce a Multi-level Information Fusion Network (MIFNet) with an integrated edge information injection strategy. Our method effectively decouples clouds into their fundamental components: body and edge (Low-Frequency (LF) and High-Frequency (HF) components), enabling the comprehensive acquisition of strong discriminative features. Specifically, we propose an Edge Feature Extraction Module (EFEM) that isolates the cloud body through low-pass filtering, while the cloud’s edge is extracted by subtracting lower-level features from LF components. Furthermore, we employ a Feature Refinement Module (FRM) to locate the cloud body’s position precisely. Building upon this foundation, we devise a Graph Reasoning Module (GRM) to facilitate the full inference of feature correlations at different levels and to model the global interdependence between edges and semantics. Through comprehensive evaluations on benchmark datasets comprising infrared band images from Landsat 8 and MODIS satellites, we demonstrate that our proposed MIFNet outperforms state-of-the-art methods, yielding promising results in cloud detection accuracy. Our code is publicly available at https://github.com/KwunYat/MIFNet. Guanyi Li, Junjie Zhang 0002, Enquan Yang, Dan Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | One-Shot Multiple Object Tracking With Robust ID PreservationabstractMaintaining identity consistency and avoiding ID-switch during tracking is one of the primary focuses of multiple object tracking (MOT). One-shot MOT methods which jointly learn the detection and tracking models in one single network (hence namely, one-shot) have achieved promising results in tracking accuracy and speed. However, their capabilities of maintaining ID consistency are somehow weakened. The reason for this weakened ID consistency is two-fold: (1) the ID features learned by one-shot methods are not discriminative enough due to their heatmap-based single-location representation. (2) severe occlusion in the MOT scene leads to feature ambiguity and high ID-switch. In this paper, we propose a one-shot MOT system with strong ID consistency called PID-MOT (Preserved ID MOT). Specifically, we devise a visibility branch to predict the object occlusion level, and a predicted visibility map will be used in both Feature Refinement Model (FRM) and a visibility-guided two-stage association strategy (VGTAS). FRM is designed to strengthen the location-based features and enrich the identity information. VGTAS is proposed for tackling objects with high and low visibility separately. In addition, we initialize the parameters of our model by training on the recently emerged abundant synthetic MOTSynth dataset from scratch rather than the commonly used COCO dataset for full training. Finally, we carry out our method on the commonly used MOT datasets and the experimental results demonstrate that the proposed PID-MOT achieves especially good performances in ID F1 score (IDF1) and ID-Switch (IDS) compared with other state-of-the-art one-shot trackers, with comparable overall HOTA/MOTA performance. The code is available at https://github.com/Kroery/PIDMOT. Weiyi Lv, Ning Zhang 0023, Junjie Zhang 0002, Dan Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Weakly Supervised Semantic Segmentation With Consistency-Constrained Multiclass Attention for Remote Sensing ScenesabstractObtaining image-level class labels for Remote Sensing (RS) images is a relatively straightforward process, sparking significant interest in Weakly Supervised Semantic Segmentation (WSSS). However, RS images present challenges beyond those encountered in generic WSSS, including complex backgrounds, densely distributed small objects, and considerable scale variations. To address above issues, we introduce a COnsistency-COnstrained Multi-Class Attention model, noted asCocoaNet. Specifically, CocoaNet endeavors to capture both semantic correlation and class distinctiveness using a Global-Local Adaptive Attention mechanism, which integrates the self-attention to model global correlation, complemented by a Local Perception branch that intensifies focus on local regions. The resulting class-specific attention weights and patch-level pairwise affinity weights are employed to optimize the initial CAMs. This mechanism proves highly effective in mitigating inter-class interference and managing the distribution of densely clustered small objects. Moreover, we invoke a Consistency Constraint to rectify activation inaccuracy. By utilizing a Siamese structure for the mutual supervision of features extracted from images at different scales, we address substantial scale variations in RS scenes. Simultaneously, a Class Contrast Loss is adopted to enhance the discriminativeness of class-specific features. Departing from the conventional CAM optimization, which is rather complex and time-consuming, we harness the prior knowledge from generic Segment Anything model to design a joint optimization strategy that refines target boundaries and further promotes discriminative visual features. We validate the effectiveness of our proposed approach on three benchmark datasets in multi-class RS scenarios, experimental results demonstrate that our model yield promising advancements compared to state-of-the-art methods. Junjie Zhang 0002, Yongshun Gong, Jian Zhang 0002, Liang Chen 0004, Dan Zeng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Frequency-Aware Multi-Modal Fine-Tuning for Few-Shot Open-Set Remote Sensing Scene ClassificationabstractFew-shot open-set recognition, as a new paradigm, leveraging a limited amount of supervised data to identify specific Remote Sensing (RS) scene categories and generalize to novel ones. However, the data bias induced by the small sample size not only causes severe overfitting within base classes, but also impairs the capacity for inference to identify RS scenes in hitherto unobserved categories. Furthermore, owing to environmental influences, RS images frequently manifest notable intra-class disparities and comparatively low inter-class distinctions, intensifying the challenge in obtaining suitable classifiers. To address above issues, we investigate the utilization of a Multi-modal Foundational Model (MFM) infused with essential domain knowledge to mitigate the generalization limitations encountered in few-shot scenarios. Recognizing that existing MFMs with a visual-text dual-branch structure are primarily tailored for natural scenes, we propose a custom Frequency Distribution-based Multi-modal Fine-Tuning strategy (FreqDiMFT) in a parameter-efficient manner. More specifically, within the vision branch, we address the high inter-class similarity and intra-class diversity in RS images by embedding the local-global frequency distribution information to facilitate the recognition of RS scenes. To further amplify the model's generalization ability post transfer, we introduce an adaptive feature refinement module designed for Transformers, proficient in filtering redundant features resulting from domain disparities. To mitigate the domain drift on the textual branch, we adopt an input format that combines basic templates with domain expertise from RS end to generate more discriminative class prototypes. To fully verify the effectiveness of our FreqDiMFT in a more practical setting, we collect a Large-Scale hybrid dataset (LSRS). Extensive experiments demonstrate that, even with a scant number of training samples, our strategy yields advanced performances compared to state-of-the-art models. Junjie Zhang 0002, Yutao Rao, Xiaoshui Huang, Guanyi Li, Dan Zeng 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | STAT: Multi-Object Tracking Based on Spatio-Temporal Topological ConstraintsabstractThe mainstream tracking-by-detection paradigm for multi-object tracking generally conducts detection first, followed by Re-IDentification (Re-ID) and motion estimation. The associations between the predicted boxes and existing tracks are then performed via visual and motion association. However, challenges such as irregular motion patterns, similar appearances, and frequent occlusions often arise, making object tracking a nontrivial task. In this article, we propose a multi-object tracker based on Spatio-TemporAl Topological (STAT) constraints to address the above issues. More specifically, we design the Feature Adaptive Association Module (FAAM) to establish the association between motion and appearance regionally, completing a complementary combination of appearance and motion features. Among these, the Appearance Feature Update Module (AFUM) is proposed to manage the appearance updates of tracked objects by imposing constraints based on the spatial locations and the degree of object occlusion, while temporal consistency is adopted to smooth the appearance states of tracks to mitigate the accumulation of appearance noise. Moreover, the Robust Motion Tracking Module (RMTM) is established to reduce the impact of irregular motions and certain unreliable detection results. The proposed module includes a higher weighted momentum term to accommodate the excessive motion amplitude and considers low-confidence boxes accompanied by the stage-wise association strategy for high-confidence boxes. Extensive experiments on DanceTrack and benchmark MOT datasets verify the effectiveness of our STAT tracker, especially the state-of-the-art results on DanceTrack, which is characterized by irregular motion and indistinguishable appearance attributes. Junjie Zhang 0002, Xinyu Zhang 0015, Chenggang Yan 0001, Dan Zeng 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | GLCSA-Net: global-local constraints-based spectral adaptive network for hyperspectral image inpainting
Jia Li 0032, Junjie Zhang 0002, Chenggang Yan 0001, Dan Zeng 0001 |
Vis. Comput. | 3 |
| 2024 | Region-guided network with visual cues correction for infrared small target detection
Junjie Zhang 0002, Dan Zeng 0001 |
Vis. Comput. | 1 |
| 2023 | A Pyramid Attention Network With Edge Information Injection for Remote-Sensing Object DetectionabstractRemote sensing images (RSIs) are often characterized by the high spatial resolution, strong object scale effects, and complex scenes, which poses great challenges to the object detection. Although mainstream neural network-based methods work well in detecting common objects, they often fail to fully exploit the detailed structural information in the spatial domain, leading to the poor performance for objects with diverse scales and distributions under complicated backgrounds. To address the above issue, we propose a pyramid attention network with edge information injection for remote sensing object detection. Considering each object is composed of the inner body and outer profile parts that corresponding to the low and high frequency components of image respectively, the difference between the original image and its low frequency component is beneficial for obtaining the high frequency counterpart. We design the Edge Information Extraction Module (EIEM) to mine the detailed edge features at multiple scales, and subsequently inject them into features at corresponding scales in the backbone network. As for promoting the performance in complex scenes, we introduce a Pyramid Feature Fusion (PFF) module, which leverages both local and global attention for establishing the long-range channel dependency, thereby highlighting objects that need to be concentrated on. To verify the effectiveness of our proposed method, we conduct extensive experiments on DIOR and RSOD datasets with mean Average Precision (mAP) reaching 74.93% and 96.44% respectively, demonstrating that our model achieved SOTA performance compared to mainstream methods. Junjie Zhang 0002, Anqi Ding, Guanyi Li, Liangang Zhang, Dan Zeng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | Progressive Recurrent Neural Network for Multispectral Remote Sensing Image DestripingabstractAn unstable imaging system often introduces additional stripe noise in multispectral remote sensing images during the data acquisition process given a variety of factors. The complicated stripe distributions lead to the residual stripe in the results of existing methods, thus increasing the difficulty of destriping in practice. Mainstream deep learning-based methods show the encouraging destriping performance on multispectral remote sensing images. However, they often require the model to handle the varying degrees of stripe noise in a single shot for each image, which results in the poor destriping performance when facing practical cases with diverse stripe distributions. To address the above issue, we propose a Progressive Recurrent Neural Network (PRNet) to remove the stripe noise for each degraded image in an iterative manner. More specifically, a progressive destriping strategy is designed to gradually restore the clean image, in which the Main Recurrent Module (MRM) is introduced to iteratively process the stripe removal results generated from previous timesteps until the clean image is obtained. Furthermore, since the uniformity of the entire image is supposed to be significantly enhanced after the destriping, it is necessary to take the local spatial correlation into account during the destriping. Therefore, we present the Patch-based Sequence Module (PSM) to leverage the local spatial correlation by splitting the image into multi-scale patch sequences and capturing the relationship among different patches. Extensive experimental results on different datasets demonstrate that the proposed model yields superior destriping performance compared to other methods, especially for removing the stripe noise with complex distributions. Jia Li 0032, Junjie Zhang 0002, Jungong Han, Chenggang Yan 0001, Dan Zeng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Genre-Conditioned Long-Term 3D Dance Generation Driven by MusicabstractDancing to music is an artistic behavior of humans, however, letting machines generate dances from music is still challenging. Most existing works have been made progress in tackling the problem of motion prediction conditioned by music, yet they rarely consider the importance of the musical genre. In this paper, we focus on generating long-term 3D dance from music with a specific genre. Specifically, we construct a pure transformer-based architecture to correlate motion features and music features. To utilize the genre information, we propose to embed the genre categories into the transformer decoder so that it can guide every frame. Moreover, different from previous inference schemes, we introduce the motion queries to output the dance sequence in parallel that significantly improves the efficiency. Extensive experiments on AIST++[1] dataset show that our model outperforms state-of-the-art methods with a much faster inference speed. Yuhang Huang 0006, Junjie Zhang 0002, Qian Bao, Dan Zeng 0001, Zhineng Chen, Wu Liu 0005 |
ICASSP | 2 |
| 2022 | Remote Sensing Image Denoising Based on Multi-Scale Feature Fusion and Regional Contextual InformationabstractThe various types of noise in Remote Sensing (RS) images resulting from environmental factors and the imaging system, often significantly degrade the imaging quality and impair high-level visual tasks. Therefore, denoising plays an essential role in the applications of RS images. Traditional methods mainly focus on dealing with a single type of noise, while the denoising performance is rather limited with the complex noise in practice. Given the advanced representation learning ability of deep neural networks, investigations have been made to apply them to the RS image denoising. However, existing methods tend to pay more attention to global features, the detailed local information are often overlooked. Therefore, in this paper, we propose a hyperspectral RS image denoising model by leveraging both multi-scale feature fusion and regional contextual information. More specifically, the proposed model includes two branches, i.e., the global branch based on Multi-scale Feature Fusion Module (MFFM) to aggregate global features from multiple scales and the local branch based on Transformer Attention Module (TAM) to explore the regional context. The denoised RS image is then obtained by combining feature maps from two branches. Extensive experimental results demonstrate that our model performs favorably on both simulated and real noisy RS images. The proposed model is also evaluated on the high-level visual tasks including object detection and clustering, which further illustrates the potential of our model for facilitating downstream tasks. Anqi Ding, Zhouyin Cai, Jia Li 0032, Junjie Zhang 0002 |
MMSP | 4 |
| 2022 | Lightweight Remote Sensing Image Denoising via Knowledge DistillationabstractSince multispectral remote sensing images (RSIs) contain the abundant information of the surface environment, they have been widely applied in diverse research areas, such as earth observation, agricultural monitoring, and geological exploration. However, RSIs generally suffer from the interference of random noise during the recording and transmission, which significantly impacts the accuracy and reliability of subsequent tasks. Existing neural network-based denoising methods often rely on the heavy parameters to model the generation process of the clean image from its noisy counterpart. However, it is time-consuming deploying such model in real application scenarios. To address the above issue, we present a lightweight denoising model based on the proposed Simplified Residual Spatial-Spectral Module (SRSSM) to effectively extract the spatial and spectral features of RSIs, while maintaining a low computation complexity. A training strategy based on the knowledge distillation is designed to constrain the feature distribution by referring to a larger teacher model. Extensive experiments demonstrate the effectiveness and efficiency of the proposed method compared against existing models. Moreover, the downstream tasks including object detection and clustering are performed on the denoised images to further validate the necessity of denoising process. Zhouyin Cai, Jia Li 0032, Junjie Zhang 0002 |
MMSP | 4 |
| 2022 | Privacy-Preserving Student Learning with Differentially Private Data-Free DistillationabstractDeep learning models can achieve high inference accuracy by extracting rich knowledge from massive well-annotated data, but may pose the risk of data privacy leakage in practical deployment. In this paper, we present an effective teacher-student learning approach to train privacy-preserving deep learning models via differentially private data-free distillation. The main idea is generating synthetic data to learn a student that can mimic the ability of a teacher well-trained on private data. In the approach, a generator is first pretrained in a data-free manner by incorporating the teacher as a fixed discriminator. With the generator, massive synthetic data can be generated for model training without exposing data privacy. Then, the synthetic data is fed into the teacher to generate private labels. Towards this end, we propose a label differential privacy algorithm termed selective randomized response to protect the label information. Finally, a student is trained on the synthetic data with the supervision of private labels. In this way, both data privacy and label privacy are well protected in a unified framework, leading to privacy-preserving models. Extensive experiments and analysis clearly demonstrate the effectiveness of our approach. Bochao Liu, Jianghu Lu, Junjie Zhang 0002, Dan Zeng 0001, Zhenxing Qian, Shiming Ge |
MMSP | 4 |
| 2022 | Multi-Object Tracking with Adaptive Cost MatrixabstractMulti-object tracking (MOT) aims at detecting and assigning identities for objects in videos. Complicated scenes, severe occlusions, irregular motions, and ambiguous appearances of objects hinder the further advance, which occurs frequently in pedestrian tracking. To tackle these challenges, we present a simple yet effective MOT framework focusing on two-fold: a more robust motion feature and a proper association paradigm. The Hybrid Motion Feature (HMF) integrates Intersection over Union (IoU), Euclidean distance metric, and area ratio information, and the latter two resort to the historical average variations to improve the cost matrix construction. Moreover, the new association paradigm, namely the Adaptive Calculation Method (ACM) performs better by avoiding the manual weighting of the motion and appearance-based cost matrix. In addition, the new correction method, i.e., reinitializing the state of the Kalman Filter when a severe mismatch occurs between the ground truth and the predicted trajectory, mitigates the effect of irregular motions. We achieved 80.7 MOTA, 78.5 IDF1, and 64.0 HOTA, outperforming the state-of-the-art model ByteTrack on the public MOT17 benchmark. Bozheng Lit, Junjie Zhang 0002 |
MMSP | 4 |
| 2022 | Controllable blending of line and polygon skeleton-based convolution surfaces with finite support kernels
Xiaoqiang Zhu, Sihu Liu, Chenjie Fan, Chenze Song, Junjie Zhang 0002, Dan Zeng 0001, Xiaogang Jin 0001 |
Comput. Graph. | 6 |
| 2022 | Expression-tailored talking face generation with adaptive cross-modal weighting
Dan Zeng 0001, Shuaitao Zhao, Junjie Zhang 0002 |
Neurocomputing | 3 |
| 2022 | Hyperspectral Anomaly Detection via Low-Rank Decomposition and Morphological FilteringabstractTo effectively detect anomalies and eliminate the influence of noise on anomaly detection (AD), we propose a hyperspectral AD method based on low-rank decomposition and morphological filtering (LRDMF). For one thing, given the different ways in which anomalies and noise occur in the spectral bands, a low-rank decomposition model is proposed to decompose the original hyperspectral image (HSI) into the background, anomaly, and noise components, where a superpixel segmentation method and the sparse representation (SR) model are used to construct a robust background dictionary. For another thing, considering that the anomalies in HSI possess small area characteristics, a morphological filtering method is applied to preserve the small connected components. Finally, anomalies are detected by jointly considering the LRDMF results. The experimental results conducted on two real hyperspectral datasets demonstrate that the proposed method outperforms some of the state-of-the-art methods. Yating Xu, Junjie Zhang 0002, Dan Zeng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | TOAN: Target-Oriented Alignment Network for Fine-Grained Image Categorization With Few Labeled SamplesabstractIn this paper, we study the fine-grained categorization problem under the few-shot setting, i.e., each fine-grained class only contains a few labeled examples, termed Fine-Grained Few-Shot classification (FGFS). The core predicament in FGFS is the high intra-class variance yet low inter-class fluctuations in the dataset. In traditional fine-grained classification, the high intra-class variance can be somewhat relieved by conducting the supervised training on the abundant labeled samples. However, with few labeled examples, it is hard for the FGFS model to learn a robust class representation with the significantly higher intra-class variance. Moreover, the inter- and intra-class variance are closely related. The significant intra-class variance in FGFS often aggravates the low inter-class variance issue. To address the above challenges, we propose a Target-Oriented Alignment Network (TOAN) to tackle the FGFS problem from both intra- and inter-class perspective. To reduce the intra-class variance, we propose a target-oriented matching mechanism to reformulate the spatial features of each support image to match the query ones in the embedding space. To enhance the inter-class discrimination, we devise discriminative fine-grained features by integrating local compositional concept representations with the global second-order pooling. We conducted extensive experiments on four public datasets for fine-grained categorization, and the results show the proposed TOAN obtains the state-of-the-art. Huaxi Huang, Junjie Zhang 0002, Litao Yu, Jian Zhang 0002, Qiang Wu 0001, Chang Xu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Adaptive Material Matching for Hyperspectral Imagery DestripingabstractDue to instrument instability, slit contamination, and light interference, hyperspectral images often suffer from striping artifacts, which greatly impairs the data quality. Real hyperspectral data are usually characterized by a small amount of historical data, complex material distribution, insignificant periodicity of noise, and so on, which brings significant challenges for the destriping task. However, the assumptions made by traditional destriping methods are often inconsistent with these characteristics. To this end, we propose a novel destriping method based on adaptive material matching (MAM) without making explicit assumptions of hyperspectral data. Specifically, to identify pixels that belong to the same material, we propose a principal material analysis (PMA) to adaptively generate thresholds within each superpixel. The pixels are matched by thresholding their vertical gradients and leveraging both inner stripe gradient feature (ISGF) and neighbor-stripe geometry feature (NSGF). Correction pixels selected from the same material can then be used to calculate the offsets and gains of pixels to adjust adjacent columns. To further improve the stability of the destriping process, we generate a set of correction candidates for each column and select the optimal candidate by considering the prior distribution and destriping nonuniformity. The stripe noise within the whole image is finally removed by iteratively performing the correction between adjacent columns. We compare the proposed model against traditional and deep learning methods on both synthetic and real hyperspectral images. The promising results indicate that MAM can effectively remove the image stripes, retain original image information, and improve the nonuniformity. Jia Li 0032, Junjie Zhang 0002, Kai Zhao 0012, Dan Zeng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | PTN: A Poisson Transfer Network for Semi-supervised Few-shot LearningabstractThe predicament in semi-supervised few-shot learning (SSFSL) is to maximize the value of the extra unlabeled data to boost the few-shot learner. In this paper, we propose a Poisson Transfer Network (PTN) to mine the unlabeled information for SSFSL from two aspects. First, the Poisson Merriman–Bence–Osher (MBO) model builds a bridge for the communications between labeled and unlabeled examples. This model serves as a more stable and informative classifier than traditional graph-based SSFSL methods in the message-passing process of the labels. Second, the extra unlabeled samples are employed to transfer the knowledge from base classes to novel classes through contrastive learning. Specifically, we force the augmented positive pairs close while push the negative ones distant. Our contrastive transfer scheme implicitly learns the novel-class embeddings to alleviate the over-fitting problem on the few labeled data. Thus, we can mitigate the degeneration of embedding generality in novel classes. Extensive experiments indicate that PTN outperforms the state-of-the-art few-shot and SSFSL models on miniImageNet and tieredImageNet benchmark datasets. Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Qiang Wu 0001, Chang Xu 0002 |
AAAI | 2 |
| 2021 | Neural Architecture Search for Joint Human Parsing and Pose EstimationabstractHuman parsing and pose estimation are crucial for the understanding of human behaviors. Since these tasks are closely related, employing one unified model to perform two tasks simultaneously allows them to benefit from each other. However, since human parsing is a pixel-wise classification process while pose estimation is usually a regression task, it is non-trivial to extract discriminative features for both tasks while modeling their correlation in the joint learning fashion. Recent studies have shown that Neural Architecture Search (NAS) has the ability to allocate efficient feature connections for specific tasks automatically. With the spirit of NAS, we propose to search for an efficient network architecture (NPPNet) to tackle two tasks at the same time. On the one hand, to extract task-specific features for the two tasks and lay the foundation for the further searching of feature interaction, we propose to search their encoder-decoder architectures, respectively. On the other hand, to ensure two tasks fully communicate with each other, we propose to embed NAS units in both multi-scale feature interaction and high-level feature fusion to establish optimal connections between two tasks. Experimental results on both parsing and pose estimation benchmark datasets have demonstrated that the searched model achieves state-of-the-art performances on both tasks.1 Dan Zeng 0001, Yuhang Huang 0006, Qian Bao, Junjie Zhang 0002, Chi Su, Wu Liu 0005 |
ICCV | 4 |
| 2021 | Exploring the auxiliary learning for long-tailed visual recognition
Junjie Zhang 0002, Lingqiao Liu, Peng Wang 0023, Jian Zhang 0002 |
Neurocomputing | 1 |
| 2021 | Low-Rank Pairwise Alignment Bilinear Network For Few-Shot Fine-Grained Image ClassificationabstractDeep neural networks have demonstrated advanced abilities on various visual classification tasks, which heavily rely on the large-scale training samples with annotated ground-truth. However, it is unrealistic always to require such annotation in real-world applications. Recently, Few-Shot learning (FS), as an attempt to address the shortage of training samples, has made significant progress in generic classification tasks. Nonetheless, it is still challenging for current FS models to distinguish the subtle differences between fine-grained categories given limited training data. To filling the classification gap, in this paper, we address the Few-Shot Fine-Grained (FSFG) classification problem, which focuses on tackling the fine-grained classification under the challenging few-shot learning setting. A novel low-rank pairwise bilinear pooling operation is proposed to capture the nuanced differences between the support and query images for learning an effective distance metric. Moreover, a feature alignment layer is designed to match the support image features with query ones before the comparison. We name the proposed model Low-Rank Pairwise Alignment Bilinear Network (LRPABN), which is trained in an end-to-end fashion. Comprehensive experimental results on four widely used fine-grained classification data sets demonstrate that our LRPABN model achieves the superior performances compared to state-of-the-art methods. Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Jingsong Xu, Qiang Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2019 | Mind Your Neighbours: Image Annotation With Metadata Neighbourhood Graph Co-Attention NetworksabstractAs the visual reflections of our daily lives, images are frequently shared on the social network, which generates the abundant 'metadata' that records user interactions with images. Due to the diverse contents and complex styles, some images can be challenging to recognise when neglecting the context. Images with the similar metadata, such as 'relevant topics and textual descriptions', 'common friends of users' and 'nearby locations', form a neighbourhood for each image, which can be used to assist the annotation. In this paper, we propose a Metadata Neighbourhood Graph Co-Attention Network (MangoNet) to model the correlations between each target image and its neighbours. To accurately capture the visual clues from the neighbourhood, a co-attention mechanism is introduced to embed the target image and its neighbours as graph nodes, while the graph edges capture the node pair correlations. By reasoning on the neighbourhood graph, we obtain the graph representation to help annotate the target image. Experimental results on three benchmark datasets indicate that our proposed model achieves the best performance compared to the state-of-the-art methods. Junjie Zhang 0002, Qi Wu 0001, Jian Zhang 0002, Chunhua Shen, Jianfeng Lu 0003 |
CVPR | 1 |
| 2019 | Compare More Nuanced: Pairwise Alignment Bilinear Network for Few-Shot Fine-Grained LearningabstractThe recognition ability of human beings is developed in a progressive way. Usually, children learn to discriminate various objects from coarse to fine-grained with limited supervision. Inspired by this learning process, we propose a simple yet effective model for the Few-Shot Fine-Grained (FSFG) recognition, which tries to tackle the challenging fine-grained recognition task using meta-learning. The proposed method, named Pairwise Alignment Bilinear Network (PABN), is an end-to-end deep neural network. Unlike traditional deep bilinear networks for fine-grained classification, which adopt the self-bilinear pooling to capture the subtle features of images, the proposed model uses a novel pairwise bilinear pooling to compare the nuanced differences between base images and query images for learning a deep distance metric. In order to match base image features with query image features, we design feature alignment losses before the proposed pairwise bilinear pooling. Experiment results on four fine-grained classification datasets and one generic few-shot dataset demonstrate that the proposed model outperforms both the state-of-the-art few-shot fine-grained and general few-shot methods. Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Qiang Wu 0001, Jingsong Xu |
ICME | 2 |
| 2019 | Heritage image annotation via collective knowledge
Junjie Zhang 0002, Qi Wu 0001, Jian Zhang 0002, Chunhua Shen, Jianfeng Lu 0003, Qiang Wu 0001 |
Pattern Recognit. | 1 |
| 2018 | Kill Two Birds With One Stone: Weakly-Supervised Neural Network for Image Annotation and Tag RefinementabstractThe number of social images has exploded by the wide adoption of social networks, and people like to share their comments about them. These comments can be a description of the image, or some objects, attributes, scenes in it, which are normally used as the user-provided tags. However, it is well-known that user-provided tags are incomplete and imprecise to some extent. Directly using them can damage the performance of related applications, such as the image annotation and retrieval. In this paper, we propose to learn an image annotation model and refine the user-provided tags simultaneously in a weakly-supervised manner. The deep neural network is utilized as the image feature learning and backbone annotation model, while visual consistency, semantic dependency, and user-error sparsity are introduced as the constraints at the batch level to alleviate the tag noise. Therefore, our model is highly flexible and stable to handle large-scale image sets. Experimental results on two benchmark datasets indicate that our proposed model achieves the best performance compared to the state-of-the-art methods. Junjie Zhang 0002, Qi Wu 0001, Jian Zhang 0002, Chunhua Shen, Jianfeng Lu 0003 |
AAAI | 1 |
| 2018 | Goal-Oriented Visual Question Generation via Intermediate Rewards
Junjie Zhang 0002, Qi Wu 0001, Chunhua Shen, Jian Zhang 0002, Jianfeng Lu 0003, Anton van den Hengel |
ECCV (5) | 1 |
| 2018 | Multilabel Image Classification With Regional Latent Semantic DependenciesabstractDeep convolution neural networks (CNNs) have demonstrated advanced performance on single-label image classification, and various progress also has been made to apply CNN methods on multilabel image classification, which requires annotating objects, attributes, scene categories, etc., in a single shot. Recent state-of-the-art approaches to the multilabel image classification exploit the label dependencies in an image, at the global level, largely improving the labeling capacity. However, predicting small objects and visual concepts is still challenging due to the limited discrimination of the global visual features. In this paper, we propose a regional latent semantic dependencies model (RLSD) to address this problem. The utilized model includes a fully convolutional localization architecture to localize the regions that may contain multiple highly dependent labels. The localized regions are further sent to the recurrent neural networks to characterize the latent semantic dependencies at the regional level. Experimental results on several benchmark datasets show that our proposed model achieves the best performance compared to the state-of-the-art models, especially for predicting small objects occurring in the images. Also, we set up an upper bound model (RLSD+ft-RPN) using bounding-box coordinates during training, and the experimental results also show that our RLSD can approach the upper bound without using the bounding-box annotations, which is more realistic in the real world. Junjie Zhang 0002, Qi Wu 0001, Chunhua Shen, Jian Zhang 0002, Jianfeng Lu 0003 |
IEEE Trans. Multim. | 1 |