EDBT 2026 Demo / reviewers in the wild / expert
Jiayuan Fan 0001
dblp:76/10698-1
· DBLP profile ↗
53ranked-venue papers
7as first author
39since 2021 · last 2026
0000-0001-7494-0255ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 5 first-author · 26 since 2021Artificial intelligence and machine learning · 16 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SRNet: Self-supervised structure regularization for stereo matching
Jun Cheng 0003, Zaiwang Gu, Weide Liu, Jiayuan Fan 0001, Zhengguo Li, Chuan-Sheng Foo |
Neurocomputing | 4 |
| 2025 | DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
Zhende Song, Jiamu Sheng, Chi Zhang 0007, Shengji Tang, Jiayuan Fan 0001, Tao Chen 0003 |
ACM Multimedia | 6 |
| 2025 | Scene123: One Prompt to 3D Scene Generation via Video-Assisted and Consistency-Enhanced MAEabstractAs Artificial Intelligence Generated Content (AIGC) advances, a variety of methods have been developed to generate text, images, videos, and 3D shapes from single or multimodal inputs, contributing efforts to emulate human-like cognitive content creation. However, generating realistic large-scale scenes from a single input presents a challenge due to the complexities involved in ensuring consistency across extrapolated views generated by models. Benefiting from recent video generation models and implicit neural representations, we propose Scene123, a 3D scene generation model, which combines a video generation framework to ensure realism and diversity with implicit neural fields integrated with Masked Autoencoders (MAE) to effectively ensure the consistency of unseen areas across views. Specifically, the input image (or a text-generated image) is first warped to simulate adjacent views, with the invisible regions filled using the consistency-enhanced MAE model. Nonetheless, the synthesized images often exhibit inconsistencies in viewpoint alignment, thus we utilize the produced views to optimize a neural radiance field, enhancing geometric consistency. Moreover, to further enhance the details and texture fidelity of generated views, we employ a GAN-based Loss against images derived from the input image through the video generation model. Extensive experiments demonstrate that our method can generate realistic and consistent scenes from a single prompt. Both qualitative and quantitative results indicate that our approach surpasses existing state-of-the-art methods. Fukun Yin, Jiayuan Fan 0001, Wanzhang Li, Xin Chen 0040, Gang Yu 0002 |
ACM Multimedia | 3 |
| 2025 | BridgeNet: Comprehensive and Effective Feature Interactions via Bridge Feature for Multi-Task Dense PredictionsabstractMulti-task dense prediction aims at handling multiple pixel-wise prediction tasks within a unified network simultaneously for visual scene understanding. However, cross-task feature interactions of current methods are still suffering from incomplete levels of representations, less discriminative semantics in feature participants, and inefficient pair-wise task interaction processes. To tackle these under-explored issues, we propose a novel BridgeNet framework, which extracts comprehensive and discriminative intermediate Bridge Features, and conducts interactions based on them. Specifically, a Task Pattern Propagation (TPP) module is first applied to ensure highly semantic task-specific feature participants are prepared for subsequent interactions, and a Bridge Feature Extractor (BFE) is specially designed to selectively integrate both high-level and low-level representations to generate the comprehensive bridge features. Then, instead of conducting heavy pair-wise cross-task interactions, a Task-Feature Refiner (TFR) is developed to efficiently take guidance from bridge features and form final task predictions. To the best of our knowledge, this is the first work considering the completeness and quality of feature participants in cross-task interactions. Extensive experiments are conducted on NYUD-v2, Cityscapes and PASCAL Context benchmarks, and the superior performance shows the proposed architecture is effective and powerful in promoting different dense prediction tasks simultaneously. Jingdong Zhang 0003, Jiayuan Fan 0001, Peng Ye 0006, Bo Zhang 0069, Hancheng Ye, Baopu Li, Yancheng Cai, Tao Chen 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | DualMamba: A Lightweight Spectral-Spatial Mamba-Convolution Network for Hyperspectral Image ClassificationabstractThe effectiveness and efficiency of modeling complex spectral–spatial relations are crucial for hyperspectral image (HSI) classification. Most existing methods based on convolution neural networks (CNNs) and transformers still suffer from heavy computational burdens and have room for improvement in capturing the global–local spectral–spatial feature representation. To this end, we propose a novel lightweight parallel design called a lightweight dual-stream Mamba-convolution network (DualMamba) for HSI classification. Specifically, a parallel lightweight Mamba and CNN block are developed to extract global and local spectral–spatial features. First, the cross-attention spectral–spatial Mamba module (CAS2MM) is proposed to leverage the global modeling of Mamba at linear complexity. In this module, dynamic positional embedding (DPE) is designed to enhance the spatial location information of visual sequences. The lightweight spectral–spatial Mamba blocks comprise an efficient scanning strategy and a lightweight Mamba design to efficiently extract global spectral–spatial features. And the cross-attention spectral–spatial fusion (CAS2F) is designed to learn cross correlation and fuse spectral–spatial features. Second, the lightweight spectral–spatial residual convolution module is proposed with lightweight spectral and spatial branches to extract local spectral–spatial features through residual learning. Finally, the adaptive global–local fusion is proposed to dynamically combine global Mamba features and local convolution features for a global–local spectral–spatial representation. Compared with state-of-the-art HSI classification methods, experimental results demonstrate that DualMamba achieves significant classification accuracy on three public HSI datasets and a superior reduction in model parameters and floating-point operations (FLOPs). Jiamu Sheng, Peng Ye 0006, Jiayuan Fan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | PM-INR: Prior-Rich Multi-Modal Implicit Large-Scale Scene Neural RepresentationabstractRecent advancements in implicit neural representations have contributed to high-fidelity surface reconstruction and photorealistic novel view synthesis. However, with the expansion of the scene scale, such as block or city level, existing methods will encounter challenges because traditional sampling cannot cope with the cubically growing sampling space. To alleviate the dependence on filling the sampling space, we explore using multi-modal priors to assist individual points to obtain more global semantic information and propose a priorrich multi-modal implicit neural representation network, Pm-INR, for the outdoor unbounded large-scale scene. The core of our method is multi-modal prior extraction and crossmodal prior fusion modules. The former encodes codebooks from different modality inputs and extracts valuable priors, while the latter fuses priors to maintain view consistency and preserve unique features among multi-modal priors. Finally, feature-rich cross-modal priors are injected into the sampling regions to allow each region to perceive global information without filling the sampling space. Extensive experiments have demonstrated the effectiveness and robustness of our method for outdoor unbounded large-scale scene novel view synthesis, which outperforms state-of-the-art methods in terms of PSNR, SSIM, and LPIPS. Fukun Yin, Wen Liu 0003, Jiayuan Fan 0001, Xin Chen 0040, Gang Yu 0002, Tao Chen 0003 |
AAAI | 4 |
| 2024 | LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and PlanningabstractRecent progress in Large Multimodal Models (LMM) has opened up great possibilities for various applications in the field of human-machine interactions. However, developing LMMs that can comprehend, reason, and plan in complex and diverse 3D environments remains a challenging topic, especially considering the demand for understanding permutation-invariant point cloud representations of the 3D scene. Existing works seek help from multi-view images by projecting 2D features to 3D space, which inevitably leads to huge computational overhead and performance degradation. In this paper, we present LL3DA, a Large Language 3D Assistant that takes point cloud as the direct input and responds to both text instructions and visual interactions. The additional visual interaction enables LMMs to better comprehend human interactions with the 3D environment and further remove the ambiguities within plain texts. Experiments show that LL3DA achieves remarkable results and surpasses various 3D vision-language models on both 3D Dense Captioning and 3D Question Answering. Sijin Chen, Xin Chen 0040, Chi Zhang 0007, Mingsheng Li, Gang Yu 0002, Hao Fei 0001, Hongyuan Zhu 0002, Jiayuan Fan 0001, Tao Chen 0003 |
CVPR | 8 |
| 2024 | MotionChain: Conversational Motion Controllers via Multimodal Prompts
Biao Jiang, Xin Chen 0040, Chi Zhang 0007, Fukun Yin, Zhuoyuan Li 0006, Gang Yu 0002, Jiayuan Fan 0001 |
ECCV (26) | 7 |
| 2024 | Through the Real World Haze Scenes: Navigating the Synthetic-to-Real Gap in Challenging Image DehazingabstractDehazing real-world hazy images is challenging due to the complexity of natural haze, varying haze conditions, details preservation, and the risk of overexposure. Existing methods excel in synthetic hazy scenarios but struggle in the real world because they don’t use all available features. Classical dehazing techniques primarily focus on low-level dehazing enhancements, whereas deep learning-based methods extract more intricate weather-related features. However, both of these approaches exhibit limitations in effectively addressing the real-world dehazing. To address these challenges, we introduce an innovative approach that combines the strengths of both modalities to dehaze and enhance the visibility of real-world hazy scenes. Firstly, we extract both low-level and deep features and then employ a pre-trained vector quantization GAN to create well-detailed data patches. A decoder, with a normalized module, effectively utilizes these high-quality features. Additionally, we introduce a controllable operation to improve feature matching. To further enhance dehazing and generalizability, the decoder’s output undergoes a sequence of gamma-correction operations and generates a series of multi-exposure images that are combined to create a haze-free and higher-quality image. Our method effectively reduces haziness, enhances sharpness, preserves natural colors, and minimizes artifacts in challenging real-world scenarios. The approach surpasses five SOTA methods in both qualitative and quantitative evaluations across three key metrics, utilizing three synthetic and two real-world hazy datasets. Notably, it achieves a substantial improvement in real-world datasets over the second-best method, with 0.5702 and 0.129 in FADE metrics for the RTTS and Fattal datasets, respectively. Mohammad Mahdizadeh, Chong Yu 0001, Jiayuan Fan 0001, Tao Chen 0003 |
ICRA | 4 |
| 2024 | Spear: Evaluate the Adversarial Robustness of Compressed Neural Models
Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001, Jiayuan Fan 0001 |
IJCAI | 4 |
| 2024 | G-Former: A Grouping Transformer for Weakly Supervised Point Cloud SegmentationabstractRecent advancements in weakly supervised point cloud semantic segmentation have diminished the reliance on extensive annotations, thereby enhancing the efficacy of understanding the real-world environment. However, existing approaches, such as PSD [25], SQN [5] and OTOC [10], often overlook the valuable global class-related prior knowledge present in point clouds beyond the scope of labels. To fully leverage this prior knowledge, which suggests that points of the same class should be close in feature space and each class should have a representative feature, we propose G-Former, a grouping transformer model. G-Former incorporates the idea of clustering into the overall model by defining clusters aligned to classes and assigning learnable tensors as cluster centers. Points are then grouped into these clusters based on the similarity of their features to the cluster centers. The core components of G-Former include a Hierarchy Cluster Structure (HCS) and a Grouping Module (GM). The former consists of two sets of clusters, one for classes while the other serves as a middle layer to help class clusters handle large-scale point features. The latter facilitates grouping the point cloud into different clusters. With the help of the grouping transformer model, G-Former further proposes a series of cluster center constraints to augment inter-class distances and diminish intra-class distances to enhance the discriminability of points. Experimental results on ScanNet v2 and S3DIS datasets demonstrate that G-Former outperforms previous methods with limited labels (0.1% or 1%) by a significant margin and is even comparable to fully supervised methods. Zehan Huang, Fukun Yin, Jiayuan Fan 0001, Xin Chen 0040, Hongyuan Zhu 0002, Bin Wang 0008, Tao Chen 0003 |
IJCNN | 3 |
| 2024 | DeNKD: Decoupled Non-Target Knowledge Distillation for Complementing Transformer-Based Unsupervised Domain AdaptationabstractThere is a growing need to explore the potential of transformers in Unsupervised Domain Adaptation (UDA) due to their increasing success in various vision tasks. However, the application of transformers in UDA has yet to be thoroughly investigated and requires further research. In this study, our primary focus is to design a novel pipeline specifically tailored for transformer-based UDA, to address a crucial challenge: the overemphasis on the transfer of target-oriented information, mainly caused by the self-attention blocks in transformers and the cross-domain adversarial learning scheme. First, we show that non-target information, including semantic contextual information such as background features and non-target classes, must be addressed in the domain adaptation process. Recognizing the importance of incorporating non-target knowledge, we propose a decoupled non-target knowledge distillation method called DeNKD. DeNKD decouples non-target information across domains at both feature and logit levels. This decoupling is achieved through a bi-directional knowledge distillation approach that facilitates the interaction and exchange of non-target knowledge to facilitate an effective transformer-based cross-domain knowledge transfer. We perform extensive evaluations on several well-established UDA benchmark datasets. The results consistently show that DeNKD outperforms other methods, achieving the best performance across the board. For example, on the Office-Home dataset, DeNKD achieves an accuracy of 85.54%, while on the VisDA-2017 dataset, it achieves an accuracy of 89.95%. These results highlight the effectiveness of DeNKD in transformer-based UDA and its potential for improving cross-domain adaptation performance. Peng Ye 0006, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Few-Shot Cross-Domain Object Detection With Instance-Level Prototype-Based Meta-LearningabstractIn typical unsupervised domain adaptive object detection, it is assumed that extensive unlabeled training data from the target domain can be easily obtained. However, in some access-constrained scenarios, massive target data cannot be guaranteed, but acquiring only a few target samples and annotating them may costs less. Therefore, inspired by the meta-learning success in few-shot tasks, we propose an Instance-level Prototype learning Network (IPNet) for solving the domain adaptive object detection under the supervised few-shot scenario in this work. To compensate for the target domain data deficiency, we fuse cropped instances from labeled images in both domains to learn a representative prototype for each class, by enforcing features of the same class’s instances but from different domains to be as close as possible. These prototypes are further employed to discriminate various features’ salience in an image, and separate foreground and background regions for respective domain alignment. Extensive experiments are conducted on several cross-domain scenarios, and their results show the consistent accuracy gains of the IPNet over state-of-the-art methods, e.g., 10.4% mAP increase on Cityscapes-to-FoggyCityscapes setting and 3.0% mAP increase on Sim10k-to-Cityscapes setting. Lin Zhang 0055, Bo Zhang 0069, Botian Shi, Jiayuan Fan 0001, Tao Chen 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | U²ConvFormer: Marrying and Evolving Nested U-Net and Scale-Aware Transformer for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification plays an important role in the human exploration of the Earth. Recent research of deep learning-based HSI classification has been fast-growing, but still suffers from three obstacles: First, existing deep learning-based HSI works lack of extraction and utilization of multigrained multiscale information and multiscale local-to-global information. Second, most previous works have too fixed-sized receptive fields in their convolutional network parts to handle HSI classification problems, and pay no attention to the existence of asymmetries in the spectral-spatial dimension of the HSI data. Third, most networks for HSI classification are hand-craft. To this end, we propose a novel architecture in this article, which is the first to combine the advantages of nested U-Net and scale-aware Transformer, named U2ConvFormer. Specifically, the nested U-Net structure can fully extract and aggregate multiscale spectral-spatial features at both inter- and inner stage granularity. The scale-aware Transformer takes multiscale local spectral-spatial features from the encoder of nested U-Net and produces multiscale global spectral-spatial features for its decoder. After that, we design a novel plug-and-play searchable operation called asymmetric spectral-spatial convolution (A2SConv), where asymmetric spectral-spatial feature pooling and multiscale feature extraction can be concurrently searched. Furthermore, we develop a customized search strategy to automatically design U2ConvFormer, which uses advanced neural architecture search (NAS) methods to enable the customization of suitable models for different hyperspectral datasets. Experimental results on three benchmark datasets, including Indian Pines, Pavia University and Houston University 2018, validate the superiority of our proposed U2ConvFormer, which achieves new state-of-the-art performance across different benchmark datasets. Lin Zhan, Peng Ye 0006, Jiayuan Fan 0001, Tao Chen 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Exploring Multi-Timestep Multi-Stage Diffusion Features for Hyperspectral Image ClassificationabstractThe effectiveness of spectral-spatial feature learning is crucial for the hyperspectral image (HSI) classification task. Diffusion models, as a new class of groundbreaking generative models, have the ability to learn both contextual semantics and textual details from the distinct timestep dimension, enabling the modeling of complex spectral-spatial relations in HSIs. However, existing diffusion-based HSI classification methods only utilize manually selected single-timestep single-stage features, limiting the full exploration and exploitation of rich contextual semantics and textual information hidden in the diffusion model. To address this issue, we propose a novel diffusion-based feature learning framework that explores Multi-Timestep Multi-Stage Diffusion features for HSI classification for the first time, called MTMSD. Specifically, the diffusion model is first pretrained with unlabeled HSI patches to mine the connotation of unlabeled data, and then is used to extract the multi-timestep multi-stage diffusion features. To effectively and efficiently leverage multi-timestep multi-stage features, two strategies are further developed. One strategy is class & timestep-oriented multi-stage feature purification module with the inter-class and inter-timestep prior for reducing the redundancy of multi-stage features and alleviating memory constraints. The other one is selective timestep feature fusion module with the guidance of global features to adaptively select different timestep features for integrating texture and semantics. Both strategies facilitate the generality and adaptability of the MTMSD framework for diverse patterns of different HSI data. Extensive experiments are conducted on four public HSI datasets, and the results demonstrate that our method outperforms state-of-the-art methods for HSI classification, especially on the challenging Houston 2018 dataset. The codes are available at https://github.com/zjyaccount/MTMSD. Jiamu Sheng, Peng Ye 0006, Jiayuan Fan 0001, Tong He 0001, Bin Wang 0008, Tao Chen 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Lightweight Model Pre-Training via Language Guided Knowledge DistillationabstractThis paper studies the problem of pre-training for small models, which is essential for many mobile devices. Current state-of-the-art methods on this problem transfer the representational knowledge of a large network (as a Teacher) into a smaller model (as a Student) using self-supervised distillation, improving the performance of the small model on downstream tasks. However, existing approaches are insufficient in extracting the crucial knowledge that is useful for discerning categories in downstream tasks during the distillation process. In this paper, for the first time, we introduce language guidance to the distillation process and propose a new method named Language-Guided Distillation (LGD) system, which uses category names of the target downstream task to help refine the knowledge transferred between the teacher and student. To this end, we utilize a pre-trained text encoder to extract semantic embeddings from language and construct a textual semantic space called Textual Semantics Bank (TSB). Furthermore, we design a Language-Guided Knowledge Aggregation (LGKA) module to construct the visual semantic space, also named Visual Semantics Bank (VSB). The task-related knowledge is transferred by driving a student encoder to mimic the similarity score distribution inferred by a teacher over TSB and VSB. Compared with other small models obtained by either ImageNet pre-training or self-supervised distillation, experiment results show that the distilled lightweight model using the proposed LGD method presents state-of-the-art performance and is validated on various downstream tasks, including classification, detection, and segmentation. Mingsheng Li, Lin Zhang 0055, Mingzhen Zhu, Gang Yu 0002, Jiayuan Fan 0001, Tao Chen 0003 |
IEEE Trans. Multim. | 6 |
| 2023 | Boost Vision Transformer with GPU-Friendly Sparsity and QuantizationabstractThe transformer extends its success from the language to the vision domain. Because of the stacked self-attention and cross-attention blocks, the acceleration deployment of vision transformer on GPU hardware is challenging and also rarely studied. This paper thoroughly designs a compression scheme to maximally utilize the GPU-friendly 2:4 fine-grained structured sparsity and quantization. Specially, an original large model with dense weight parameters is first pruned into a sparse one by 2:4 structured pruning, which considers the GPU's acceleration of 2:4 structured sparse pattern with FP16 data type, then the floating-point sparse model is further quantized into a fixed-point one by sparse-distillation-aware quantization aware training, which considers GPU can provide an extra speedup of 2:4 sparse calculation with integer tensors. A mixed-strategy knowledge distillation is used during the pruning and quantization process. The proposed compression scheme is flexible to support supervised and unsupervised learning styles. Experiment results show GPUSQ-ViT scheme achieves state-of-the-art compression by reducing vision transformer models$\mathbf{6.4}-\mathbf{12.7}\times$on model size and$\mathbf{30.3}-\mathbf{62} \times$on FLOPs with negligible accuracy degradation on ImageNet classification, COCO detection and ADE20K segmentation benchmarking tasks. Moreover, GPUSQ-ViT can boost actual deployment performance by$\mathbf{1.39}-\mathbf{1.79}\times$and$\mathbf{3.22}-\mathbf{3.43}\times$of latency and throughput on A100 GPU, and$\mathbf{1.57}-\mathbf{1.69}\times$and$\mathbf{2.11}-\mathbf{2.51}\times$improvement of latency and throughput on AGX Orin. Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001, Jiayuan Fan 0001 |
CVPR | 4 |
| 2023 | JNDMix: Jnd-Based Data Augmentation for No-Reference Image Quality AssessmentabstractDespite substantial progress in no-reference image quality assessment (NR-IQA), previous training models often suffer from over-fitting due to the limited scale of used datasets, resulting in model performance bottlenecks. To tackle this challenge, we explore the potential of leveraging data augmentation to improve data efficiency and enhance model robustness. However, most existing data augmentation methods incur a serious issue, namely that it alters the image quality and leads to training images mismatching with their original labels. Additionally, although only a few data augmentation methods are available for NR-IQA task, their ability to enrich dataset diversity is still insufficient. To address these issues, we propose a effective and general data augmentation based on just noticeable difference (JND) noise mixing for NR-IQA task, named JNDMix. In detail, we randomly inject the JND noise, imperceptible to the human visual system (HVS), into the training image without any adjustment to its label. Extensive experiments demonstrate that JNDMix significantly improves the performance and data efficiency of various state-of-the-art NR-IQA models and the commonly used baseline models, as well as the generalization ability. More importantly, JNDMix facilitates MANIQA to achieve the state-of-the-art performance on LIVEC and KonIQ-10k. Jiamu Sheng, Jiayuan Fan 0001, Peng Ye 0006, Jianjian Cao |
ICASSP | 2 |
| 2023 | A2S-NAS: Asymmetric Spectral-Spatial Neural Architecture Search for Hyperspectral Image ClassificationabstractExisting deep learning-based hyperspectral image (HSI) classification works still suffer from the limitation of the fixed-sized receptive field, leading to difficulties in distinctive spectral-spatial features for ground objects with various sizes and arbitrary shapes. Meanwhile, plenty of previous works ignore asymmetric spectral-spatial dimensions in HSI. To address the above issues, we propose a multi-stage search architecture in order to overcome asymmetric spectral-spatial dimensions and capture significant features. First, the asymmetric pooling on the spectral-spatial dimension maximally retains the essential features of HSI. Then, the 3D convolution with a selectable range of receptive fields overcomes the constraints of fixed-sized convolution kernels. Finally, we extend these two searchable operations to different layers of each stage to build the final architecture. Extensive experiments are conducted on two challenging HSI benchmarks including Indian Pines and Houston University, and results demonstrate the effectiveness of the proposed method with superior performance compared with the related works. Lin Zhan, Jiayuan Fan 0001, Peng Ye 0006, Jianjian Cao |
ICASSP | 2 |
| 2023 | A Large-Scale Outdoor Multi-modal Dataset and Benchmark for Novel View Synthesis and Implicit Scene ReconstructionabstractNeural Radiance Fields (NeRF) [24] has achieved impressive results in single object scene reconstruction and novel view synthesis, as demonstrated on many single modality and single object focused indoor scene datasets like DTU [14], BMVS [42], and NeRF Synthetic [24]. However, the study of NeRF on large-scale outdoor scene reconstruction is still limited, as there is no unified outdoor scene dataset for large-scale NeRF evaluation due to expensive data acquisition and calibration costs. In this work, we propose a large-scale outdoor multi-modal dataset, OMMO dataset, containing complex objects and scenes with calibrated images, point clouds and prompt annotations. A new benchmark for several outdoor NeRF-based tasks is established, such as novel view synthesis, diverse 3D representation, and multi-modal NeRF. To create the dataset, we capture and collect a large number of real fly-view videos and select high-quality and high-resolution clips from them. Then we design a quality review module to refine images, remove low-quality frames and fail-to-calibrate scenes through a learning-based automatic evaluation plus manual review. Finally, volunteers are employed to label and review the prompt annotation for each scene and keyframe. Compared with existing NeRF datasets, our dataset contains abundant real-world urban and natural scenes with various scales, camera trajectories, and lighting conditions. Experiments show that our dataset can benchmark most state-of-the-art NeRF methods on different tasks. The dataset can be found at the following link: https://ommo.luchongshan.com/. Chongshan Lu, Fukun Yin, Xin Chen 0040, Wen Liu 0003, Tao Chen 0003, Gang Yu 0002, Jiayuan Fan 0001 |
ICCV | 7 |
| 2023 | Rethinking Pseudo-Label-Based Unsupervised Person Re-ID with Hierarchical Prototype-based GraphabstractUnsupervised person re-identification (Re-ID) aims to match individuals without manual annotations. However, existing methods often struggle with intra-class variations due to differences in person poses and camera styles such as resolution and environment information. Additionally, clustering may produce incorrect pseudo-labels, compounding the issue. To address these challenges, we propose a novel hierarchical prototype-based graph network (HPG-Net) for unsupervised person Re-ID. Our approach uses a hierarchical prototype-based graph structure to describe person images by attributes of poses and camera styles, with each graph node representing the average of image features as a prototype. We then apply a hierarchical contrastive learning module to enhance the feature learning at each level, reducing the impact of intra-class differences caused by extraneous attributes. We also calculate the similarity between samples and each level of prototypes, maintaining prototype-based graph consistency with the mean-teacher network to mitigate the accumulation errors caused by pseudo-labels. Experimental results on three benchmarks show that our method outperforms state-of-the-art (SOTA) works. Moreover, we achieve promising performance on an occluded dataset. Ben Sha, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001 |
ACM Multimedia | 4 |
| 2023 | PDF: Point Diffusion Implicit Function for Large-scale Scene Neural RepresentationabstractRecent advances in implicit neural representations have achieved impressive results by sampling and fusing individual points along sampling rays in the sampling space. However, due to the explosively growing sampling space, finely representing and synthesizing detailed textures remains a challenge for unbounded large-scale outdoor scenes. To alleviate the dilemma of using individual points to perceive the entire colossal space, we explore learning the surface distribution of the scene to provide structural priors and reduce the samplable space and propose a Point Diffusion implicit Function, PDF, for large-scale scene neural representation. The core of our method is a large-scale point cloud super-resolution diffusion module that enhances the sparse point cloud reconstructed from several training images into a dense point cloud as an explicit prior. Then in the rendering stage, only sampling points with prior points within the sampling radius are retained. That is, the sampling space is reduced from the unbounded space to the scene surface. Meanwhile, to fill in the background of the scene that cannot be provided by point clouds, the region sampling based on Mip-NeRF 360 is employed to model the background representation. Expensive experiments have demonstrated the effectiveness of our method for large-scale scene novel view synthesis, which outperforms relevant state-of-the-art baselines. Yuhan Ding, Fukun Yin, Jiayuan Fan 0001, Xin Chen 0040, Wen Liu 0003, Chongshan Lu, Gang Yu 0002, Tao Chen 0003 |
NeurIPS | 3 |
| 2023 | Performance-Aware Approximation of Global Channel Pruning for Multitask CNNsabstractGlobal channel pruning (GCP) aims to remove a subset of channels (filters) across different layers from a deep model without hurting the performance. Previous works focus on either single task model pruning or simply adapting it to multitask scenario, and still face the following problems when handling multitask pruning: 1) Due to the task mismatch, a well-pruned backbone for classification task focuses on preserving filters that can extract category-sensitive information, causing filters that may be useful for other tasks to be pruned during the backbone pruning stage; 2) For multitask predictions, different filters within or between layers are more closely related and interacted than that for single task prediction, making multitask pruning more difficult. Therefore, aiming at multitask model compression, we propose a Performance-Aware Global Channel Pruning (PAGCP) framework. We first theoretically present the objective for achieving superior GCP, by considering the joint saliency of filters from intra- and inter-layers. Then a sequentially greedy pruning strategy is proposed to optimize the objective, where a performance-aware oracle criterion is developed to evaluate sensitivity of filters to each task and preserve the globally most task-related filters. Experiments on several multitask datasets show that the proposed PAGCP can reduce the FLOPs and parameters by over 60% with minor performance drop, and achieves 1.2x ∼ 3.3x acceleration on both cloud and mobile platforms. Our code is available at http://www.github.com/HankYe/PAGCP.git. Hancheng Ye, Bo Zhang 0069, Tao Chen 0003, Jiayuan Fan 0001, Bin Wang 0008 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Hyperspectral Image Classification Using Spectral-Spatial Token Enhanced Transformer With Hash-Based Positional EmbeddingabstractHyperspectral image (HSI) classification aims to distinguish the category of a land coverage object for each pixel. In an effective way, the transformer architecture has been successfully introduced for the HSI classification task with promising performance. However, existing transformer-based HSI classification methods still suffer from the inability to fully explore both spectral information and spatial information in HSIs. To this end, we propose a Spectral-Spatial Token Enhanced Transformer (SSTE-Former) method with the hash-based positional embedding, which is the first to exploit multiscale spectral-spatial information for transformer-based HSI classification in-depth. Specifically, SSTE-Former accepts multiscale HSI cubes centered on the target pixel, that are preprocessed by PCA. Then, a designed multiscale CNN architecture is utilized to extract short-range spectral-spatial features and generate token embeddings. In parallel, a novel hash-based spatially enhanced positional embedding tailored for HSI cubes is developed to model the correlations within and across multiscale token embeddings. Finally, multiscale token embeddings and hash-based positional embeddings are concatenated and flattened into the transformer encoder for long-range spectral-spatial feature fusion. We conduct extensive experiments on four benchmark HSI datasets and achieve superior performance compared with the state-of-the-art HSI classification methods. Jiayuan Fan 0001, Peng Ye 0006, Mingzhen Zhu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Exploring Kernel-Based Texture Transfer for Pose-Guided Person Image GenerationabstractPose-guided person image generation that aims to transfer the pose of a given person to a target pose has recently received lots of research attention. Due to the spatial misalignment and occlusions of different local body parts by pose variations, this task is still challenging especially in maintaining high-fidelity textures and body structures in generated images. Besides, most works also suffer from the limited number of texture styles in the given person datasets, restricting the diversity of generated persons' appearances. To solve these problems, we design a Kernel-based Texture-Fusion Joint Refinement Network (TFJR-Net) to jointly refine the structure and texture information of generated images. First, we leverage a bone-map representation to guide the generation of human parsing maps, which has more structure priors and richer context information than traditional key-point maps, thus reduce the uncertainty of generated body structures. Next, a Texture-Kernel Injection Normalization module (TKIN) is proposed to inject the per-region texture-kernel into the corresponding semantic region from the human parsing map, which decouples the texture and shape information, and also preserves fine-grained features for complex textures. Furthermore, we are the first to introduce external texture patterns outside of the dataset in human semantic regions such as the upper clothes. We fuse the two texture domains in a shared texture space through our designed texture-fusion TKIN modules. Extensive experiments are conducted on the Deepfashion dataset, with the DTD dataset as an external texture source. The experimental results demonstrate the superiority of our proposed method in generating persons of better textures and structures than state-of-the-art works, and also show the generalization ability of our proposed method to absorb diversified external textures for generating person images. The source codes are available athttps://github.com/pilgrim00/TKIN. Jiaxiang Chen, Jiayuan Fan 0001, Hancheng Ye, Jie Li 0040, Yongbin Liao, Tao Chen 0003 |
IEEE Trans. Multim. | 2 |
| 2022 | β-DARTS: Beta-Decay Regularization for Differentiable Architecture SearchabstractNeural Architecture Search (NAS) has attracted increasingly more attention in recent years because of its capability to design deep neural network automatically. Among them, differential NAS approaches such as DARTS, have gained popularity for the search efficiency. However, they suffer from two main issues, the weak robustness to the performance collapse and the poor generalization ability of the searched architectures. To solve these two problems, a simple-but-efficient regularization method, termed as Beta-Decay, is proposed to regularize the DARTS-based NAS searching process. Specifically, Beta-Decay regularization can impose constraints to keep the value and variance of activated architecture parameters from too large. Furthermore, we provide in-depth theoretical analysis on how it works and why it works. Experimental results on NAS-Bench-201 show that our proposed method can help to stabilize the searching process and makes the searched network more transferable across different datasets. In addition, our search scheme shows an outstanding property of being less dependent on training time and data. Comprehensive experiments on a variety of search spaces and datasets validate the effectiveness of the proposed method. The code is available at https://github.com/Sunshine-Ye/Beta-DARTS. Peng Ye 0006, Baopu Li, Yikang Li 0002, Tao Chen 0003, Jiayuan Fan 0001, Wanli Ouyang |
CVPR | 5 |
| 2022 | Learning Cross-Image Object Semantic Relation in Transformer for Few-Shot Fine-Grained Image ClassificationabstractFew-shot fine-grained learning aims to classify a query image into one of a set of support categories with fine-grained differences. Although learning different objects' local differences via Deep Neural Networks has achieved success, how to exploit the query-support cross-image object semantic relations in Transformer-based architecture remains under-explored in the few-shot fine-grained scenario. In this work, we propose a Transformer-based double-helix model, namely HelixFormer, to achieve the cross-image object semantic relation mining in a bidirectional and symmetrical manner. The HelixFormer consists of two steps: 1) Relation Mining Process (RMP) across different branches, and 2) Representation Enhancement Process (REP) within each individual branch. By the designed RMP, each branch can extract fine-grained object-level Cross-image Semantic Relation Maps (CSRMs) using information from the other branch, ensuring better cross-image interaction in semantically related local object regions. Further, with the aid of CSRMs, the developed REP can strengthen the extracted features for those discovered semantically-related local regions in each branch, boosting the model's ability to distinguish subtle feature differences of fine-grained objects. Extensive experiments conducted on five public fine-grained benchmarks demonstrate that HelixFormer can effectively enhance the cross-image object semantic relation matching for recognizing fine-grained objects, achieving much better performance over most state-of-the-art methods under 1-shot and 5-shot scenarios. Bo Zhang 0069, Jiakang Yuan, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001, Botian Shi |
ACM Multimedia | 5 |
| 2022 | What Makes for Effective Few-shot Point Cloud Classification?abstractDue to the emergence of powerful computing resources and large-scale annotated datasets, deep learning has seen wide applications in our daily life. However, most current methods require extensive data collection and retraining when dealing with novel classes never seen before. On the other hand, we humans can quickly recognize new classes by looking at a few samples, which motivates the recent popularity of few-shot learning (FSL) in machine learning communities. Most current FSL approaches work on 2D image domain, however, its implication in 3D perception is relatively under-explored. Not only needs to recognize the unseen examples as in 2D domain, 3D few-shot learning is more challenging with unordered structures, high intra-class variances and subtle inter-class differences. Moreover, different architectures and learning algorithms make it difficult to study the effectiveness of existing 2D methods when migrating to the 3D domain.In this work, for the first time, we perform systematic and extensive studies of recent 2D FSL and 3D backbone networks for benchmarking few-shot point cloud classification, and we suggest a strong baseline and learning architectures for 3D FSL. Then, we propose a novel plug-and-play component called Cross-Instance Adaptation (CIA) module, to address the high intra-class variances and subtle inter-class differences issues, which can be easily inserted into current baselines with significant performance improvement. Extensive experiments on two newly introduced benchmark datasets, ModelNet40-FS and ShapeNet70-FS, demonstrate the superiority of our proposed network for 3D FSL. Chuangguan Ye, Hongyuan Zhu 0002, Yongbin Liao, Yanggang Zhang, Tao Chen 0003, Jiayuan Fan 0001 |
WACV | 6 |
| 2022 | Efficient Joint-Dimensional Search with Solution Space Regularization for Real-Time Semantic Segmentation
Peng Ye 0006, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001, Chen Lin 0003, Chongyan Zuo, Qinghua Chi, Wanli Ouyang |
Int. J. Comput. Vis. | 4 |
| 2022 | Point Cloud Instance Segmentation With Semi-Supervised Bounding-Box MiningabstractPoint cloud instance segmentation has achieved huge progress with the emergence of deep learning. However, these methods are usually data-hungry with expensive and time-consuming dense point cloud annotations. To alleviate the annotation cost, unlabeled or weakly labeled data is still less explored in the task. In this paper, we introduce the first semi-supervised point cloud instance segmentation framework (SPIB) using both labeled and unlabelled bounding boxes as supervision. To be specific, our SPIB architecture involves a two-stage learning procedure. For stage one, a bounding box proposal generation network is trained under a semi-supervised setting with perturbation consistency regularization (SPCR). The regularization works by enforcing an invariance of the bounding box predictions over different perturbations applied to the input point clouds, to provide self-supervision for network learning. For stage two, the bounding box proposals with SPCR are grouped into some subsets, and the instance masks are mined inside each subset with a novel semantic propagation module and a property consistency graph module. Moreover, we introduce a novel occupancy ratio guided refinement module to refine the instance masks. Extensive experiments on the challenging ScanNet v2 dataset demonstrate our method can achieve competitive performance compared with the recent fully-supervised methods. Yongbin Liao, Hongyuan Zhu 0002, Yanggang Zhang, Chuangguan Ye, Tao Chen 0003, Jiayuan Fan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Densely Semantic Enhancement for Domain Adaptive Region-Free DetectorsabstractUnsupervised domain adaptive object detection aims to adapt a well-trained detector from its original source domain with rich labeled data to a new target domain with unlabeled data. Previous works focus on improving the domain adaptability of region-based detectors,e.g., Faster-RCNN, through matching cross-domain instance-level features that are explicitly extracted from a region proposal network (RPN). However, this is unsuitable for region-free detectors such as single shot detector (SSD), which perform a dense prediction from all possible locations in an image and do not have the RPN to encode such instance-level features. As a result, they fail to align important image regions and crucial instance-level features between the domains of region-free detectors. In this work, we propose an adversarial module, namely, densely semantic enhancement module (DSEM), to strengthen the cross-domain matching of instance-level features for region-free detectors. Firstly, to emphasize the important regions of image, the DSEM learns to predict a transferable foreground enhancement mask that can be utilized to suppress the background disturbance in an image. Secondly, considering that region-free detectors recognize objects of different scales using multi-layer feature maps, the DSEM encodes multi-scale representations across different domains. Finally, the DSEM is pluggable into different region-free detectors, ultimately achieving the densely semantic feature matching via adversarial learning. Extensive experiments have been conducted on PASCAL VOC, Clipart, Comic, W atercolor, and FoggyCityscape benchmarks, and their results well demonstrate that the proposed approach not only improves the domain adaptability of region-free detectors but also outperforms existing domain adaptive region-based detectors under various domain shift settings. Bo Zhang 0069, Tao Chen 0003, Bin Wang 0008, Xiaofeng Wu 0003, Liming Zhang 0001, Jiayuan Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | SC-EADNet: A Self-Supervised Contrastive Efficient Asymmetric Dilated Network for Hyperspectral Image ClassificationabstractUnsupervised and semisupervised feature learning has recently emerged as an effective way to reduce the reliance on expensive data collection and annotation for hyperspectral image (HSI) analysis. Existing unsupervised and semisupervised convolutional neural network (CNN)-based HSI classification works still face two challenges: underutilization of pixel-wise multiscale contextual information for feature learning and expensive computational cost, for example, large floating-point operations per seconds (FLOPs), due to the lack of lightweight design. To utilize the unlabeled pixels in the HSIs more efficiently, we propose a self-supervised contrastive efficient asymmetric dilated network (SC-EADNet) for HSI classification. There are two novelties in the SC-EADNet. First, a self-supervised multiscale pixel-wise contextual feature learning model is proposed, which generates multiple patches around each hyperspectral pixel and develops a contrastive learning framework to learn from these patches for HSI classification. Second, a lightweight feature extraction network EADNet, composed of multiple plug-and-play efficient asymmetric dilated convolution (EADC) blocks, is designed and inserted into the contrastive learning framework. The EADC block adopts different dilation rates to capture the spatial information of objects with varying shapes and sizes. Compared with other unsupervised, semisupervised, and supervised learning methods, our SC-EADNet provides competitive classification performance on four hyperspectral datasets, including Indian Pines, Pavia University, Salinas, and Houston 2013, but few FLOPs and fast computational speed. Mingzhen Zhu, Jiayuan Fan 0001, Qihang Yang 0002, Tao Chen 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Sample-Centric Feature Generation for Semi-Supervised Few-Shot LearningabstractSemi-supervised few-shot learning aims to improve the model generalization ability by means of both limited labeled data and widely-available unlabeled data. Previous works attempt to model the relations between the few-shot labeled data and extra unlabeled data, by performing a label propagation or pseudo-labeling process using an episodic training strategy. However, the feature distribution represented by the pseudo-labeled data itself is coarse-grained, meaning that there might be a large distribution gap between the pseudo-labeled data and the real query data. To this end, we propose a sample-centric feature generation (SFG) approach for semi-supervised few-shot image classification. Specifically, the few-shot labeled samples from different classes are initially trained to predict pseudo-labels for the potential unlabeled samples. Next, a semi-supervised meta-generator is utilized to produce derivative features centering around each pseudo-labeled sample, enriching the intra-class feature diversity. Meanwhile, the sample-centric generation constrains the generated features to be compact and close to the pseudo-labeled sample, ensuring the inter-class feature discriminability. Further, a reliability assessment (RA) metric is developed to weaken the influence of generated outliers on model learning. Extensive experiments validate the effectiveness of the proposed feature generation approach on challenging one- and few-shot image classification benchmarks. Bo Zhang 0069, Hancheng Ye, Gang Yu 0002, Bin Wang 0008, Yike Wu 0001, Jiayuan Fan 0001, Tao Chen 0003 |
IEEE Trans. Image Process. | 6 |
| 2021 | EADNet: Efficient Asymmetric Dilated Network For Semantic SegmentationabstractDue to real-time image semantic segmentation needs on power constrained edge devices, there has been an increasing desire to design lightweight semantic segmentation neural network, to simultaneously reduce computational cost and increase inference speed. In this paper, we propose an efficient asymmetric dilated semantic segmentation network, named EADNet, which consists of multiple developed asymmetric convolution branches with different dilation rates to capture the variable shapes and scales information of an image. Specially, a multi-scale multi-shape receptive field convolution (MMRFC) block with only a few parameters is designed to capture such information. Experimental results on the Cityscapes dataset demonstrate that our proposed EADNet achieves segmentation mIoU of 67.1% with smallest number of parameters (only 0.35M) among mainstream lightweight semantic segmentation networks. Qihang Yang 0002, Tao Chen 0003, Jiayuan Fan 0001, Chongyan Zuo, Qinghua Chi |
ICASSP | 3 |
| 2021 | HSEGAN: Hair Synthesis and Editing Using Structure-Adaptive Normalization on Generative Adversarial NetworkabstractHuman hair is a kind of special material with complex and varied high-frequency details. It is a challenging task to synthesize and edit realistic and fine-grained hair using deep learning methods. In this paper, we propose HSEGAN, a novel framework consisting of two condition modules encoding foreground hair and background respectively, followed by a hair synthesis generator that synthesizes the final result based on the encoded input. For the purpose of efficient and effective hair generation, we propose hair structure-adaptive normalization (HSAN) and use several HSAN residual blocks to build the hair synthesis generator. HSEGAN allows for explicit manipulation of hair at three different levels, including color, structure and shape. Extensive experiments on FFHQ dataset demonstrate our method can generate higher-quality hair images than state-of-the-art methods, yet consume less time in the inference stage. Wanling Fan, Jiayuan Fan 0001, Gang Yu 0002, Tao Chen 0003 |
ICIP | 2 |
| 2021 | Spcr: semi-supervised point cloud instance segmentation with perturbation consistency regularizationabstractPoint cloud instance segmentation is steadily improving with the development of deep learning. However, current progress is hindered by the expensive cost of collecting dense point cloud labels. To this end, we propose the first semi-supervised point cloud instance segmentation architecture, which is called semi-supervised point cloud instance segmentation with perturbation consistency regularization (SPCR). It is capable to alleviate the data-hungry bottleneck of existing strongly supervised methods. Specifically, SPCR enforces an invariance of the predictions over different perturbations applied to the input point clouds. We firstly introduce various perturbation schemes on inputs to force the network to be robust and easily generalized to the unseen and unlabeled data. Further, perturbation consistency regularization is then conducted on predicted instance masks from various transformed inputs to provide self-supervision for network learning. Extensive experiments on the challenging ScanNet v2 dataset demonstrate our method can achieve competitive performance compared with the state-of-the-art of fully supervised methods. Yongbin Liao, Hongyuan Zhu 0002, Tao Chen 0003, Jiayuan Fan 0001 |
ICIP | 4 |
| 2021 | Global-to-Local Dynamic Feature Aggregation for Unsupervised Person Re-IdentificationabstractMost of existing unsupervised person re-identification algorithms use single-scale structure to extract global features. However, the features extracted from different scales and spatial locations are crucial for person re-identification task mainly using body parts to identify. In this paper, we propose a novel global-local architecture called dynamic feature aggregation network (DFANet) to learn discriminative patch features at various semantic levels on unlabelled datasets and aggregate them with a dynamic aggregation mechanism. Specifically, DFANet starts with a global feature learning stage to learn global features at various semantic levels. Then a local stage formed by attention patch generation network and dynamic aggregation module is deployed to extract distinct and notable local features on unlabelled datasets. Extensive experimental results on Market1501 and DukeMTMC show that our proposed method outperforms state-of-the-art works. Wei Li 0132, Jiayuan Fan 0001, Yanwei Fu 0001 |
ICME | 2 |
| 2021 | Object-aware Long-short-range Spatial Alignment for Few-Shot Fine-Grained Image ClassificationabstractThe goal of few-shot fine-grained image classification is to recognize rarely seen fine-grained objects in the query set, given only a few samples of this class in the support set. Previous works focus on learning discriminative image features from a limited number of training samples for distinguishing various fine-grained classes, but ignore one important fact that spatial alignment of the discriminative semantic features between the query image with arbitrary changes and the support image, is also critical for computing the semantic similarity between each support-query pair. In this work, we propose an object-aware long-short-range spatial alignment approach, which is composed of a foreground object feature enhancement (FOE) module, a long-range semantic correspondence (LSC) module and a short-range spatial manipulation (SSM) module. The FOE is developed to weaken background disturbance and encourage higher foreground object response. To address the problem of long-range object feature misalignment between support-query image pairs, the LSC is proposed to learn the transferable long-range semantic correspondence by a designed feature similarity metric. Further, the SSM module is developed to refine the transformed support feature after the long-range step to align short-range misaligned features (or local details) with the query features. Extensive experiments have been conducted on four benchmark datasets, and the results show superior performance over most state-of-the-art methods under both 1-shot and 5-shot classification scenarios. Yike Wu 0001, Bo Zhang 0069, Gang Yu 0002, Weixi Zhang, Bin Wang 0008, Tao Chen 0003, Jiayuan Fan 0001 |
ACM Multimedia | 7 |
| 2021 | Coarse-to-Fine Gaze Redirection with Numerical and Pictorial GuidanceabstractGaze redirection aims at manipulating the gaze of a given face image with respect to a desired direction (i.e., a reference angle) and it can be applied to many real life scenarios, such as video-conferencing or taking group photos. However, previous work on this topic mainly suffers of two limitations: (1) Low-quality image generation and (2) Low redirection precision. In this paper, we propose to alleviate these problems by means of a novel gaze redirection framework which exploits both a numerical and a pictorial direction guidance, jointly with a coarse-to-fine learning strategy. Specifically, the coarse branch learns the spatial transformation which warps input image according to desired gaze. On the other hand, the fine-grained branch consists of a generator network with conditional residual image learning and a multi-task discriminator. This second branch reduces the gap between the previously warped image and the ground-truth image and recovers finer texture details. Moreover, we propose a numerical and pictorial guidance module (NPG) which uses a pictorial gazemap description and numerical angles as an extra guide to further improve the precision of gaze redirection. Extensive experiments on a benchmark dataset show that the proposed method outperforms the state-of-the-art approaches in terms of both image quality and redirection precision. The code is available at https://github.com/jingjingchen777/CFGR Jichao Zhang, Enver Sangineto, Tao Chen 0003, Jiayuan Fan 0001, Nicu Sebe |
WACV | 5 |
| 2020 | PIDNet: An Efficient Network for Dynamic Pedestrian Intrusion DetectionabstractVision-based dynamic pedestrian intrusion detection (PID), judging whether pedestrians intrude an area-of-interest (AoI) by a moving camera, is an important task in mobile surveillance. The dynamically changing AoIs and a number of pedestrians in video frames increase the difficulty and computational complexity of determining whether pedestrians intrude the AoI, which makes previous algorithms incapable of this task. In this paper, we propose a novel and efficient multi-task deep neural network, PIDNet, to solve this problem. PIDNet is mainly designed by considering two factors: accurately segmenting the dynamically changing AoIs from a video frame captured by the moving camera and quickly detecting pedestrians from the generated AoI-contained areas. Three efficient network designs are proposed and incorporated into PIDNet to reduce the computational complexity: 1) a special PID task backbone for feature sharing, 2) a feature cropping module for feature cropping, and 3) a lighter detection branch network for feature compression. In addition, considering there are no public datasets and benchmarks in this field, we establish a benchmark dataset to evaluate the proposed network and give the corresponding evaluation metrics for the first time. Experimental results show that PIDNet can achieve 67.1% PID accuracy and 9.6 fps inference speed on the proposed dataset, which serves as a good baseline for the future vision-based dynamic PID study. Jingchen Sun, Jiming Chen 0001, Tao Chen 0003, Jiayuan Fan 0001, Shibo He |
ACM Multimedia | 4 |
| 2020 | BURSTS: A bottom-up approach for robust spotting of texts in scenes
Jiayuan Fan 0001, Tao Chen 0003, Feng Zhou 0003 |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | SS-HCNN: Semi-Supervised Hierarchical Convolutional Neural Network for Image ClassificationabstractThe availability of large-scale annotated data and uneven separability of different data categories become two major impediments of deep learning for image classification. In this paper, we present a Semi-Supervised Hierarchical Convolutional Neural Network (SS-HCNN) to address these two challenges. A large-scale unsupervised maximum margin clustering technique is designed, which splits images into a number of hierarchical clusters iteratively to learn cluster-level CNNs at parent nodes and category-level CNNs at leaf nodes. The splitting uses the similarity of CNN features to group visually similar images into the same cluster, which relieves the uneven data separability constraint. With the hierarchical cluster-level CNNs capturing certain high-level image category information, the category-level CNNs can be trained with a small amount of labelled images, and this relieves the data annotation constraint. A novel cluster splitting criterion is also designed which automatically terminates the image clustering in the tree hierarchy. The proposed SS-HCNN has been evaluated on the CIFAR-100 and ImageNet classification datasets. Experiments show that the SS-HCNN trained using a portion of labelled training images can achieve comparable performance with other fully trained CNNs using all labelled images. Additionally, the SS-HCNN trained using all labelled images clearly outperforms other fully trained CNNs. Tao Chen 0003, Shijian Lu, Jiayuan Fan 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | S-CNN: Subcategory-Aware Convolutional Networks for Object DetectionabstractThe marriage between the deep convolutional neural network (CNN) and region proposals has made breakthroughs for object detection in recent years. While the discriminative object features are learned via a deep CNN for classification, the large intra-class variation and deformation still limit the performance of the CNN based object detection. We propose a subcategory-aware CNN (S-CNN) to solve the object intra-class variation problem. In the proposed technique, the training samples are first grouped into multiple subcategories automatically through a novel instance sharing maximum margin clustering process. A multi-component Aggregated Channel Feature (ACF) detector is then trained to produce more latent training samples, where each ACF component corresponds to one clustered subcategory. The produced latent samples together with their subcategory labels are further fed into a CNN classifier to filter out false proposals for object detection. An iterative learning algorithm is designed for the joint optimization of image subcategorization, multi-component ACF detector, and subcategory-aware CNN classifier. Experiments on INRIA Person dataset, Pascal VOC 2007 dataset and MS COCO dataset show that the proposed technique clearly outperforms the state-of-the-art methods for generic object detection. Tao Chen 0003, Shijian Lu, Jiayuan Fan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Superpixel Guided Deep-Sparse-Representation Learning for Hyperspectral Image ClassificationabstractThis paper presents a new technique for hyperspectral image (HSI) classification by using superpixel guided deep-sparse-representation learning. The proposed technique constructs a hierarchical architecture by exploiting the sparse coding to learn the HSI representation. Specifically, a multiple-layer architecture using different superpixel maps is designed, where each superpixel map is generated by downsampling the superpixels gradually along with enlarged spatial regions for labeled samples. In each layer, sparse representation of pixels within every spatial region is computed to construct a histogram via the sum-pooling with l1normalization. Finally, the representations (features) learned from the multiple-layer network are aggregated and trained by a support vector machine classifier. The proposed technique has been evaluated over three public HSI data sets, including the Indian Pines image set, the Salinas image set, and the University of Pavia image set. Experiments show superior performance compared with the state-of-the-art methods. Jiayuan Fan 0001, Tao Chen 0003, Shijian Lu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Unsupervised Feature Learning for Land-Use Scene RecognitionabstractThis paper proposes a novel unsupervised feature learning algorithm for land-use scene recognition on very high resolution remote sensing imagery. The proposed technique utilizes a multipath sparse coding architecture in order to capture multiple aspects of discriminative structures within complex remote sensing sceneries. Unlike the previous sparse coding and bag-of-visual-words-based techniques that rely on the handcrafted feature descriptors such as scale-invariant feature transform, the proposed technique extracts dense low-level features from the raw data, including the visual (RGB) data and near-infrared (NIR) data, using image patches of varying sizes at different layers. The proposed technique has been evaluated on three data sets, including the 21-category UC Merced landuse RGB data set with a 1-ft spatial resolution, the 9-category ground scene RGB-NIR data set, and the 10-category Singapore land-use RGB-NIR data set with a 0.5-m spatial resolution. The experimental results show that the proposed technique outperforms the state-of-the-art methods. Jiayuan Fan 0001, Tao Chen 0003, Shijian Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Landmark recognition with compact BoW histogram and ensemble ELM
Jiuwen Cao, Tao Chen 0003, Jiayuan Fan 0001 |
Multim. Tools Appl. | 3 |
| 2015 | Reversible watermarking using enhanced local predictionabstractReversible watermarking has drawn extensive attentions in recent years due to its broad applications of digital forensics and data security. This paper proposes a novel reversible watermarking approach, which has the perfect reversibility of the embedded data and the original image. In order to increase the embedding capacity, the proposed approach utilizes the enhanced local prediction to reduce the prediction error of every pixel value. Different from traditional reversible watermarking algorithms considering all the image pixels equivalently before prediction, an enhanced image is first computed by multiplying the original image with its saliency map. By linearly formulating each pixel value in an original image as a weighted sum of the enhanced pixel values in its local neighborhood, the correlation coefficients as learned weights are then solved as a least squares solution. Based on two state-of-the-art datasets, the experimental results show that the proposed approach achieves large embedding capacity with relatively low visual distortion. Jiayuan Fan 0001, Tao Chen 0003 |
ICIP | 1 |
| 2015 | Vegetation coverage detection from very high resolution satellite imageryabstractAutomatic vegetation coverage detection plays a key role for monitoring and management of land usage, environmental variation, and urban planning. This paper presents a novel vegetation coverage detection technique for very high resolution multi-spectral satellite imagery. The proposed technique consists of two stages including a supervised patch-level scoring stage and an unsupervised pixel-level classification stage. In the first stage, a support vector regression (SVR) technique is developed which scores each image patch and generates a coarse patch-level vegetation map. In the second stage, an unsupervised pixel-level vegetation classification technique is developed, which produces a more detailed vegetation map by re-scoring those uncertain pixels based on the computed SVR scores. Experiments on very high resolution multi-spectral satellite images show that the proposed technique outperforms the state-of-the-art methods in both patch-level and pixel-level vegetation detection. Jiayuan Fan 0001, Tao Chen 0003, Shijian Lu |
VCIP | 1 |
| 2015 | Context-aware vocabulary tree for mobile landmark recognition
Tao Chen 0003, Shijian Lu, Jiayuan Fan 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Context-aware codebook learning for mobile landmark recognitionabstractThis paper presents a codebook learning based mobile landmark recognition technique based on context information that is acquired from mobile devices. Previous codebook learning methods are mainly developed on nonmobile platforms such as desktop PC, hence underutilize context features such as location and direction information as provided by the mobile devices. The proposed technique employs both the direction and location information to learn the codebook for mobile landmark recognition. A set of direction-aware leaf codewords are first generated by using direction data to decompose the leaf nodes of the original SVT. A visual word significance learning algorithm is then developed by considering location information to generate a compact codebook for image encoding. Experiments on the NTU50Landmark database show that the proposed method can achieve good recognition performance in mobile landmark recognition. Tao Chen 0003, Jiayuan Fan 0001, Shijian Lu |
ICIP | 2 |
| 2013 | A novel approach for partial blur detection and segmentationabstractThis paper proposes a novel approach for partial blur detection and segmentation. The local blur kernels of image blocks are firstly estimated and then a reblurring technique is used to measure relative blur degrees of the local blur kernels. The output of reblurring is a metric to classify blurred and non-blurred image blocks. Furthermore, block-based and pixel-based techniques are incorporated for a fine segmentation of blurred and non-blurred regions. Our approach is evaluated for out-of-focus and motion blurred images. The experimental results show that the proposed approach detects and segments the blurred and non-blurred regions in partial blurred images with 88% accuracy for natural out-of-focus blur, 86% accuracy for artificial out-of-focus blur and 83% accuracy for artificial motion blur, which outperforms the state-of-the-art approaches of partial blur detection and segmentation. Khosro Bahrami, Alex Chichung Kot, Jiayuan Fan 0001 |
ICME | 3 |
| 2013 | Estimating EXIF Parameters Based on Noise Features for Image Manipulation DetectionabstractWe propose in this paper a novel technique to correlate statistical image noise features with three EXchangeable Image File format (EXIF) header features for manipulation detection. By formulating each EXIF feature as a weighted sum of selected statistical image noise features using sequential floating forward selection, the weights are then solved as a least squares solution for modeling the correlation between the intact image and the corresponding EXIF header. Image manipulations like brightness and contrast adjustment can affect these noise features and lead to enlarged numerical difference between each actual and its estimated EXIF feature from the noise features. By using the numerical difference as a manipulation indicator, we achieve excellent performance in detecting common brightness and contrast adjustment. Based on cameras of different brands, our manipulation detection is also demonstrated to work well in a blind mode, where the camera brand/model source is unavailable. Several detection examples suggest that our model can be applied in detecting real-world forgeries. Jiayuan Fan 0001, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2011 | Modeling the EXIF-Image correlation for image manipulation detectionabstractEXchangeable Image File format (EXIF) is a metadata header containing shot-related camera settings such as aperture, exposure time, ISO speed etc. These settings can affect the photo content in many ways. In this paper, we investigate the underlying EXIF-Image correlation and propose a novel model, which correlates image statistical noise features with several commonly used EXIF features. By formulating each EXIF feature as a weighted combination of different image statistical noise features, we first select a compact image statistical noise feature set using sequential floating forward selection. The underlying correlation as a set of regression weights is then solved using a least squares solution. When applying our learned correlation to detect image manipulation, we achieve average test accuracies of 94.6%, 94.1% and 94.9% in three different cameras to detect the presence of common image brightness and contrast adjustment. Jiayuan Fan 0001, Alex Chichung Kot, Farook Sattar |
ICIP | 1 |