Dejun Zhang

dblp:116/6802 · DBLP profile ↗
← Back
43ranked-venue papers
10as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Temporal and spatial context aware voxel transformer for semantic scene completion
Yiqi Wu, Changliang Li, Jiale He, Cuilian Lei, Yilin Chen 0001, Dejun Zhang
Neural Networks7
2026 Weather-aware multi-granularity representation learning for cross-view geo-localization under adverse conditions
Shifeng Xu, Xujie Long, Dejun Zhang
Pattern Recognit.5
2026 SAM-Zero3D: Extending Segment Anything to Zero Shot 3D Scene Segmentation via Iterative Global-Local Interaction
abstract
Lifting multi-view 2D masks generated by the Segment Anything Model (SAM) into 3D space offers a promising direction for zero-shot 3D scene segmentation, but view-dependent occlusions and limited fields of view often cause incomplete observations and cross-view inconsistencies, resulting in fragmented semantics and geometric misalignment. To address this, we propose SAM-Zero3D, which extends SAM to the 3D domain through a structured fusion pipeline with two complementary branches. The global anchor point-guided branch projects 3D anchors into multi-view masks to construct a cross-view affinity graph, identifies consistent mask groups via connected component analysis, and assigns 3D masks via majority voting and nearest-neighbor propagation. The local geometry-driven branch partitions the point cloud into fine-grained regions, estimates region-level semantic similarity from aggregated mask distributions, and progressively merges similar regions through a multi-stage merging strategy. An iterative global–local interaction further refines both branches by aligning global semantic priors with local geometric cues. Extensive experiments on ShapeNetPart, ScanNetV2, and ScanNet200 show that SAM-Zero3D significantly outperforms existing zero-shot baselines, achieving accurate and structure-aware segmentation without any 3D training or supervision.
Dejun Zhang, Shifeng Xu, Yanzi Bai, Yiqi Wu, Jun Liu 0036
IEEE Trans. Circuits Syst. Video Technol.1
2026 A Fast Injection and Low-Overhead Compensation Method for IR-Drop in RRAM-Based CNN In-Memory Computing
Debao Wei, Jingyuan Qu, Yanlong Zeng, Dejun Zhang, Liyan Qiao
IEEE Trans. Very Large Scale Integr. Syst.4
2025 Audio Aesthetics Prediction System QAM16k Based on Pre-trained Audio Encoder
abstract
Meta Audiobox introduces a groundbreaking framework for audio aesthetics assessment, effectively addressing the core limitations of traditional MOS-like evaluation systems, namely ambiguous scoring objectives and imprecise quality deficit attribution. In the AudioMOS Challenge 20251, Track 2 focuses on Audiobox-aesthetics-style prediction tasks. This study presents the T04 team's system QAM16k for Track 2: leveraging a pre-trained 16 kHz Qwen2-Audio Encoder for audio feature extraction, and mapping the representations to four-dimensional scores via four customized multi-layer perceptrons (MLPs). To explore the potential of full-band information, we investigated a band-split feature fusion architecture. Although this approach did not outperform the 16 kHz system in empirical tests, its design principles provide valuable insights for future research. Experimental results on the AES-Natural dataset demonstrate that QAM16k achieved superior performance across multiple metrics compared to the open-source Meta Audiobox baseline. In Track 2 of AudioMOS Challenge 2025, the T04 system ranked second in 18 out of 32 evaluation metrics2, validating the effectiveness of the proposed system.A1udioMOS Challenge 2025: https://sites.google.com/view/voicemos-challenge/audiomos-challenge-20252AudioMOS Challenge 2025 Track2 Results: https://docs.google.com/spreadsheets/d/17s9hKRwbvDlcGDgJUm5tN6UHX7rqBk2lyv8xFwUsUPg/edit?gid=0#gid=0
Linping Xu, Ziqian Wu, Dejun Zhang
ASRU3
2025 RLRFusion: RCS-based LiDAR-Radar Fusion for 3D Object Detection
abstract
In the field of autonomous driving and intelligent transportation systems, 3D object detection plays a critical role in ensuring safe and efficient driving. Achieving reliable object detection relies on robust sensing technologies, where LiDAR provides accurate spatial perception, and radar offers extended detection range and speed information because of its longer wave-lengths. To capitalize on the strengths of both sensors, a novel RCS-based LiDAR-Radar fusion network, named RLRFusion, is proposed for 3D object detection in this paper. The network takes LiDAR and radar point clouds as inputs and processes them through dual bird's-eye view (BEV) feature extraction streams, followed by the BEV fusion and detection module to produce the detection results. In input-level fusion, a cross-modal pillar encoder is introduced to address the sparse radar data and its lack of height information. In feature-level fusion, an RCS-aware fusion encoder leverages the Radar Cross Section (RCS) distribution by mapping pillar features to their surroundings, enhancing object size estimation and mitigating the challenges faced by LiDAR in adverse weather conditions. Experimental results show that RLRFusion achieves competitive performance on the nuScenes dataset, with strong detection results even in rainy conditions. The source code of our method is available at: https://github.com/djZzgroupIRLRFusion.
Yiqi Wu, Jiale He, Xiantao Cai, Dejun Zhang, Changliang Li, Yilin Chen 0001
CSCWD4
2025 Visual Keyword Spotting with Multi-Encoder for MAVSR 2025
abstract
This paper systematically describes the Fosafer system designed for the Mandarin Audio-Visual Speech Recognition (MAVSR) Challenge 2025 Track 2. The purpose of Track 2 is to evaluate the performance of visual speech analysis systems in identifying whether a speaker has pronounced specific keywords in silent video sequences, also known as visual keyword spotting (VKWS). In this paper, we propose to improve VKWS with a video augmentation method and visual multi-encoder. Specifically, we first propose a simple but effective video augmentation method to fully utilize the scarce video data in VKWS. Secondly, visual multi-encoders like Transformer and Conformer for lip motion feature extraction are implied. Finally, we adopt the strategy of model fusion to further enhance model performance. Ablation experiments show that each module of our proposed system has a different level of enhancement to the visual keyword spotting results. The proposed method ranked first in Track 2 of the MAVSR 2025 Challenge, achieving a Mean Average Precision (mAP) of $36.226 \%$ on the test dataset.
Dejun Zhang, Xupeng Jia
FG1
2025 Mitigating Category Imbalance: Fosafer System for the Multimodal Emotion and Intent Joint Understanding Challenge
abstract
This paper presents Fosafer’s approach to the Track 2 Mandarin in the Multimodal Emotion and Intent Joint Understanding (MEIJU) challenge, which focuses on achieving joint recognition of emotion and intent in Mandarin, despite the issue of category imbalance. To alleviate this issue, we use a variety of data augmentation techniques across text, video, and audio modalities. Additionally, we introduce the Sample-Weighted Focal Contrastive (SWFC) loss, designed to address the challenges of recognizing minority class samples and those that are semantically similar but difficult to distinguish. Moreover, we fine-tune the Hubert model to adapt the emotion and intent joint recognition. To mitigate modal competition, we introduce a modal dropout strategy. For the final predictions, a plurality voting approach is used to determine the results. The experimental results demonstrate the effectiveness of our method, which achieves the second-best performance in the Track 2 Mandarin challenge.
Honghong Wang, Dejun Zhang
ICASSP3
2025 Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge
Shangkun Huang, Dejun Zhang, Xupeng Jia, Jintao Kang
INTERSPEECH4
2025 Multimodal 3D Few-Shot Classification via Gaussian Mixture Discriminant Analysis
abstract
Abstract While pre‐trained 3D vision‐language models are becoming increasingly available, there remains a lack of frameworks that can effectively harness their capabilities for few‐shot classification. In this work, we propose PointGMDA, a training‐free framework that combines Gaussian Mixture Models (GMMs) with Gaussian Discriminant Analysis (GDA) to perform robust classification using only a few labeled point cloud samples. Our method estimates GMM parameters per class from support data and computes mixture‐weighted prototypes, which are then used in GDA with a shared covariance matrix to construct decision boundaries. This formulation allows us to model intra‐class variability more expressively than traditional single‐prototype approaches, while maintaining analytical tractability. To incorporate semantic priors, we integrate CLIP‐style textual prompts and fuse predictions from geometric and textual modalities through a hybrid scoring strategy. We further introduce PointGMDA‐T, a lightweight attention‐guided refinement module that learns residuals for fast feature adaptation, improving robustness under distribution shift. Extensive experiments on ModelNet40 and ScanObjectNN demonstrate that PointGMDA outperforms strong baselines across a variety of few‐shot settings, with consistent gains under both training‐free and fine‐tuned conditions. These results highlight the effectiveness and generality of our probabilistic modeling and multimodal adaptation framework. Our code is publicly available at https://github.com/djzgroup/PointGMDA .
Yiqi Wu, HuaChao Wu, Ronglei Hu, Yilin Chen 0001, Dejun Zhang
Comput. Graph. Forum5
2025 FlowST-Net: Tackling non-uniform spatial and temporal distributions for scene flow estimation in point clouds
Xiaohu Yan, Xuefeng Tan, Yiqi Wu, Dejun Zhang
Neurocomputing5
2025 Iteration and SDA-Driven LDPC Decoding Latency Reduction for 3-D TLC NAND Flash Memory
abstract
To enhance the reliability of 3-D TLC NAND flash memory, low-density parity-check (LDPC) codes have become widely adopted. However, as the number of read, program or erasures increases, the raw bit error rate (RBER) of read data in flash memory chips also rises, leading to the challenge of increased LDPC decoding latency. To address this, an idea of utilizing the decoding correct probability of LDPC to reduce latency is proposed. First, by analyzing the encoding method of TLC NAND flash memory, a single direction characterization error model based on the correlation between individual pages is constructed. Next, through experimental validation, the feasibility of using the number of iterations as an indicator of decoding correct probability is demonstrated, leading to the proposal of a low-overhead and high-performance Iterative Alternative Correct Probability Optimization (IACPO) scheme. Leveraging the fact that LDPC decoding exhibits a high success probability within a specific range of read data, the existence of successful decoding area (SDA) is confirmed through a large number of real experiments, and the distribution characteristics of SDA in TLC NAND flash memory are analyzed. Finally, the Utilizing LDPC Decoding Correct-Probability Optimization (ULDCO) scheme to further optimize latency is proposed. This scheme uses SDA to improve the probability of correct data reading, particularly in the middle and later stages of flash memory life, in combination with the IACPO scheme. Experimental results show that the proposed IACPO and ULDCO schemes are not only generally applicable but also achieve significantly iterative latency reduction by 73.43% and 81.51%, respectively, compared to traditional schemes, with negligible storage and computation overhead. These results clearly demonstrate the superior performance of the proposed schemes in reducing iteration latency.
Debao Wei, Yongchao Wang 0001, Dejun Zhang, Huqi Xiang, Liyan Qiao
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 IDWA: A Importance-Driven Weight Allocation Algorithm for Low Write-Verify Ratio RRAM-Based In-Memory Computing
abstract
Resistive random access memory (RRAM)-based in-memory computing (IMC) architectures are currently receiving widespread attention. Since this computing approach relies on the analog characteristics of the devices, the write variation of RRAM can affect the computational accuracy to varying degrees. Conventional write–verify (W&V) procedures are performed on all weight parameters, resulting in significant time overhead. To address this issue, we propose a training algorithm that can recover the offline IMC accuracy impacted by write variation with a lower cost of W&V overhead. We introduce a importance-driven weight allocation (IDWA) algorithm during the training process of the neural network. This algorithm constrains the values of less important weights to suppress the diffusion of variation interference on this part of the weights, thus reducing unnecessary accuracy degradation. Additionally, we employ a layer-wise optimization algorithm to identify important weights in the neural network for W&V operations. Extensive testing across various deep neural networks (DNNs) architectures and datasets demonstrates that our proposed selective W&V methodology consistently outperforms current state-of-the-art selective W&V techniques in both accuracy preservation and computational efficiency. At same accuracy levels, it delivers a speed improvement of$6\times \sim 32\times $compared to other advanced methods.
Jingyuan Qu, Debao Wei, Dejun Zhang, Yanlong Zeng, Zhelong Piao, Liyan Qiao
IEEE Trans. Very Large Scale Integr. Syst.3
2024 3D Contour Generation based on Diffusion Probabilistic Models
abstract
The contours of objects effectively represent essential 3D information such as shapes and boundaries. Existing contour detection methods mostly rely on thresholding or neural networks to classify points as edge or non-edge points. However, these methods often lack generalization ability on datasets with different shapes, leading to issues such as missing contours and discontinuous distribution. To address this, we propose a 3D point cloud contour generation method based on the denoising diffusion probabilistic model (DDPM). Our method treats the complete point cloud as an explicit condition to guide the generation of contour from noise. Specifically, our method trains the DDPM to be a conditional generative network customized for contour generation tasks. In the network, we design the Conditional Feature Extraction (CFE) module that obtains multi-scale feature information, and the Conditional Feature Fusion (CFF) module embeds this information in the generation process to guide contour generation. The experimental results demonstrate the effectiveness of our method. The source code of our method is available at: https://github.com/djzgroup/ContourGeneration.
Yiqi Wu, Kelin Song, Fazhi He, Dejun Zhang
CSCWD5
2024 Mitigating Intra-Class Variance in Few-Shot Point Cloud Classification
abstract
Due to the significant intra-class variance of 3D point clouds, it becomes challenging to characterize prototype features with a small number of instances in few-shot classification. The significant feature discrepancies among instances also hinder category determination. In this paper, we propose a few-shot point cloud classification network based on prototype learning. We mitigate intra-class variance and enhance classification performance from three aspects of the network. Firstly, we enrich point cloud features through a multi-scale grouping and pooling strategy. Subsequently, we engage in learning compensatory information from support features to update preliminary prototype features. Finally, we enhance both prototype and query features through instance feature fusion. We conducted few-shot point cloud classification experiments on benchmark datasets, and the results indicate that our approach achieves state-of-the-art performance. The source code of our method is available at https://github.com/djzgroup/FewshotClassification.
Yiqi Wu, Kelin Song, Dejun Zhang
ICASSP4
2024 Active Learning with Core-Set Sampling and Scale-Sensitive Loss for 3D Object Detection
abstract
Deep learning-based 3D object detectors often require large-scale labeled 3D datasets, which can be expensive to annotate. To tackle this issue, we introduce a core-set sampling strategy within an active learning framework, selecting highly informative data from a data pool to reduce reliance on such datasets. Additionally, we introduce a scale-sensitive loss function into the 3D object detector to mitigate the disparities in the influence of small and large objects on model learning, thereby enhancing the accuracy of small object detection. Our experiments on the KITTI dataset demonstrate that our method achieves comparable results to the baseline using only 20% of the data. Notably, our approach outperforms the baseline in small object detection, with an 8% accuracy improvement for pedestrians and a 4% improvement for cyclists. The source code of the proposed method is available at https://github.com/djzgroup/al-cs-ssl.
Dejun Zhang, Xiaowei Lin, Benxin Yi, Yiqi Wu
ICASSP1
2024 Enhanced ASR FOR Stuttering Speech: Combining Adversarial and Signal-Based Data Augmentation
abstract
This paper presents our submission to the SLT2024 StutteringSpeech Challenge, focusing on augmenting stuttering data using straightforward and effective techniques. We combined adversarial and signal-based data augmentation methods, including modifying speech rate and rhythm, inserting silence segments, repeating speech segments, and applying Generative Adversarial Network-based (GAN-based) perturbation. These techniques enabled us to generate stuttering speech from fluent speech, which we used to train our automatic speech recognition (ASR) model, enhancing its robustness for individuals with stuttering. Our system achieved a character error rate (CER) of 12.30% in the StutteringSpeech Challenge Track 2, demonstrating a relative improvement of 35.87% over the official baseline and securing first place in the competition.
Shangkun Huang, Dejun Zhang
SLT2
2024 Integrating Self-Supervised Pre-Training With Adversarial Learning for Synthesized Song Detection
abstract
Existing spoofing detection systems often perform poorly when applied to highly realistic synthetic song datasets. To address it, we propose a method that integrates self-supervised pre-training with adversarial learning. Initially, we utilize wav2vec 2.0 to extract audio representations, which are subsequently fed into a back-end classifier. A ResBlock-based network is then employed to capture fine-grained audio features. Additionally, we enhance a RawNet-based model by incorporating a gradient reversal layer and applying adversarial training to improve generalization to unknown algorithms. Finally, the outputs of various models are combined at the score level. Experimental results demonstrate that our approach achieves Equal Error Rates (EER) of 1.57% on the test sets of the Controlled Singing Voice Deepfake Detection (CtrSVDD) track, providing relative reductions of 84.89% compared to the baseline B02. Our system achieves the first place during the CtrSVDD track of SVDD challenge 2024.
Dejun Zhang
SLT3
2024 Two-stage video anomaly detection based on dual-stream networks and multi-instance learning
abstract
Abstract To promptly detect abnormal events in surveillance videos, this article designs a video anomaly detection method based on multiple instance learning. Generally, abnormal events occur less frequently compared to normal events. Traditional video surveillance relies on manual operation to monitor scenes and detect abnormal events by watching surveillance videos. However, watching surveillance footage is a labor‐intensive task, and prolonged observation can lead to visual fatigue and lack of concentration, which in turn results in missed detections and false positives [1]. Therefore, it is crucial to develop intelligent algorithms for video anomaly detection. The method can detect whether segments of a video contain abnormal events. First, the I3D network is used as a feature extractor to capture spatiotemporal features from the input video. Then, the spatiotemporal information is processed and input into a segment‐level anomaly detector based on multiple instance learning for detection. The authors treat abnormal videos as positive bags and normal videos as negative bags, and automatically learn a deep anomaly ranking model that can predict abnormal segments. Finally, the results of the training were tested and analyzed, demonstrating that the model is capable of detecting abnormal traffic segments.
Dejun Zhang, Wenbo Fang, Zirong Lyu, Chen Xiong
IET Image Process.1
2024 Unsupervised non-rigid point cloud registration based on point-wise displacement learning
Yiqi Wu, Dejun Zhang, Yilin Chen 0001
Multim. Tools Appl.3
2024 Unsupervised distribution-aware keypoints generation from 3D point clouds
Yiqi Wu, Xingye Chen, Kelin Song, Dejun Zhang
Neural Networks5
2024 Bridging the Domain Gap in Scene Flow Estimation via Hierarchical Smoothness Refinement
abstract
This article introduces SmoothFlowNet3D, an innovative encoder-decoder architecture specifically designed for bridging the domain gap in scene flow estimation. To achieve this goal, SmoothFlowNet3D divides the scene flow estimation task into two stages: initial scene flow estimation and smoothness refinement. Specifically, SmoothFlowNet3D comprises a hierarchical encoder that extracts multi-scale point cloud features from two consecutive frames, along with a hierarchical decoder responsible for predicting the initial scene flow and further refining it to achieve smoother estimation. To generate the initial scene flow, a cross-frame nearest-neighbor search operation is performed between the features extracted from two consecutive frames, resulting in forward and backward flow embeddings. These embeddings are then combined to form the bidirectional flow embedding, serving as input for predicting the initial scene flow. Additionally, a flow smoothing module based on the self-attention mechanism is proposed to predict the smoothing error and facilitate the refinement of the initial scene flow for more accurate and smoother estimation results. Extensive experiments demonstrate that the proposed SmoothFlowNet3D approach achieves state-of-the-art performance on both synthetic datasets and real LiDAR point clouds, confirming its effectiveness in enhancing scene flow smoothness.
Dejun Zhang, Xuefeng Tan, Jun Liu 0036
ACM Trans. Multim. Comput. Commun. Appl.1
2023 An Intra-BRNN and GB-RVQ Based END-TO-END Neural Audio Codec
Linping Xu, Dejun Zhang, Xianjun Xia, Yijian Xiao, Piao Ding, Shenyi Song, Sixing Yin, Ferdous Sohel
INTERSPEECH3
2022 Coarse-to-fine pipeline for 3D wireframe reconstruction from point cloud
Xuefeng Tan, Dejun Zhang, Yiqi Wu, Yilin Chen 0001
Comput. Graph.2
2022 Investigating Impacts of Ambient Air Pollution on the Terrestrial Gross Primary Productivity (GPP) From Remote Sensing
abstract
In contrast to the threats to urban human health, impacts of air pollutants on the ecosystem photosynthesis seem to be less concerned. The existence of aerosols could promote photosynthesis by increasing the ratio of diffuse to direct solar radiation; on the contrary, ozone (O3) could inhibit photosynthesis, as it is detrimental to leaf stomata. However, it is unknown whether these two opposite impacts worldwide cancel each other out. In the current mainstream methods, earth system models may show conflicts within situexperimental results due to their relatively coarse resolution. In virtue of satellite remote sensing and a global eddy covariance (EC) network, we studied ten years of data to explore the impacts of aerosol and O3on photosynthesis by fitting an explainable machine learning model. The impacts of aerosol on gross primary productivity (GPP) were positive in many cases, yet very weak. By means of the nitrogen dioxide (NO2) to formaldehyde (HCHO) ratio, O3was seen with positive impacts on photosynthesis under the NOx-sensitive regime, but the apparent positive impacts correlated with the plant phenology. Under the volatile organic compound (VOC)-sensitive regime, the impacts of O3on GPP were not obvious, which was likely due to the prioritized depletion of O3by NO2and VOCs. The impacts of air pollutants depended on many factors and results varied case by case, but the overall net impacts were negative.
Songyan Zhu, Jian Xu 0008, Jingya Zeng, Qiaolin Zeng, Dejun Zhang
IEEE Geosci. Remote. Sens. Lett.7
2022 Satellite Remote Sensing of Daily Surface Ozone in a Mountainous Area
abstract
High-levels of surface ozone (O3) pollution threaten human and environmental health. Chongqing, a mountainous municipality located in southwest China, is exposed to serious O3 pollution and requires more studies. Due to its complex terrain and always foggy weather, it is difficult to maintain many in-situ sites in Chongqing, and Chemical Transportation Model (CTM) simulations are also challenged. The recently launched (in 2017) Sentinel-5p satellite provides O3 columns with advanced spatiotemporal resolution. Without the dependence on CTMs, we linked O3 columns and surface monitoring data from 2019 to 2021 in virtue of a deep forest machine-learning model. Compared with another widely used machine-learning model and previous studies, our results showed great advantages in estimating surface O3 on a daily scale. Validated against in-situ sites in Chongqing, averaged R2 of cross-validations reached 0.9 while the root mean squared error (RMSE) and mean bias error (MBE) were 13.57 and 0.37 μg/m3. We found out that the model performance is associated with relative height difference between training sites and the test site. The model performed stably when the height difference was lower than 200 m, but obvious performance degradation was seen when the height difference exceeding 400 m.
Songyan Zhu, Jian Xu 0008, Qiaolin Zeng, Dejun Zhang
IEEE Geosci. Remote. Sens. Lett.5
2021 Three-Module Modeling For End-to-End Spoken Language Understanding Using Pre-Trained DNN-HMM-Based Acoustic-Phonetic Model
abstract
In spoken language understanding (SLU), what the user says is converted to his/her intent.Recent work on end-to-end SLU has shown that accuracy can be improved via pre-training approaches.We revisit ideas presented by Lugosch et al. using speech pre-training and three-module modeling; however, to ease construction of the end-to-end SLU model, we use as our phoneme module an open-source acoustic-phonetic model from a DNN-HMM hybrid automatic speech recognition (ASR) system instead of training one from scratch.Hence we fine-tune on speech only for the word module, and we apply multi-target learning (MTL) on the word and intent modules to jointly optimize SLU performance.MTL yields a relative reduction of 40% in intent-classification error rates (from 1.0% to 0.6%).Note that our three-module model is a streaming method.The final outcome of the proposed three-module modeling approach yields an intent accuracy of 99.4% on FluentSpeech, an intent error rate reduction of 50% compared to that of Lugosch et al.Although we focus on real-time streaming methods, we also list non-streaming methods for comparison.
Nick J. C. Wang, Yandan Sun, Haimei Kang, Dejun Zhang
Interspeech5
2020 Weight asynchronous update: Improving the diversity of filters in a deep convolutional network
abstract
Deep convolutional networks have obtained remarkable achievements on various visual tasks due to their strong ability to learn a variety of features. A well-trained deep convolutional network can be compressed to 20%–40% of its original size by removing filters that make little contribution, as many overlapping features are generated by redundant filters. Model compression can reduce the number of unnecessary filters but does not take advantage of redundant filters since the training phase is not affected. Modern networks with residual, dense connections and inception blocks are considered to be able to mitigate the overlap in convolutional filters, but do not necessarily overcome the issue. To do so, we propose a new training strategy, weight asynchronous update, which helps to significantly increase the diversity of filters and enhance the representation ability of the network. The proposed method can be widely applied to different convolutional networks without changing the network topology. Our experiments show that the stochastic subset of filters updated in different iterations can significantly reduce filter overlap in convolutional networks. Extensive experiments show that our method yields noteworthy improvements in neural network performance.
Dejun Zhang, Linchao He, Mengting Luo, Zhanya Xu, Fazhi He
Comput. Vis. Media1
2020 Multimodal image registration using histogram of oriented gradient distance and data-driven grey wolf optimizer
Xiaohu Yan, Yongjun Zhang 0002, Dejun Zhang, Neng Hou
Neurocomputing3
2020 Registration of Multimodal Remote Sensing Images Using Transfer Optimization
abstract
Multimodal image registration is critical yet challenging for remote sensing image processing. Due to the large nonlinear intensity differences between the multimodal images, conventional search algorithms tend to get trapped into local optima when optimizing the transformation parameters by maximizing mutual information (MI). To address this problem, inspired by transfer learning, we propose a novel search algorithm named transfer optimization (TO), which can be applied to any optimizer. In TO, an optimizer transfers its better individuals to the other optimizer in each iteration. Thus, TO can share information between two optimizers and take advantage of their search mechanisms, which is helpful to avoid the local optima. Then, the registration of the multimodal remote sensing images using TO is presented. We compare the proposed algorithm with several state-of-the-art algorithms on real and simulated image pairs. Experimental results demonstrate the superiority of our algorithm in terms of registration accuracy.
Xiaohu Yan, Yongjun Zhang 0002, Dejun Zhang, Neng Hou, Bin Zhang 0046
IEEE Geosci. Remote. Sens. Lett.3
2020 Learning motion representation for real-time spatio-temporal action localization
Dejun Zhang, Linchao He, Zhigang Tu 0001, Shifu Zhang, Boxiong Yang
Pattern Recognit.1
2020 Unsupervised Learning of Optical Flow With CNN-Based Non-Local Filtering
abstract
Estimating optical flow from successive video frames is one of the fundamental problems in computer vision and image processing. In the era of deep learning, many methods have been proposed to use convolutional neural networks (CNNs) for optical flow estimation in an unsupervised manner. However, the performance of unsupervised optical flow approaches is still unsatisfactory and often lagging far behind their supervised counterparts, primarily due to over-smoothing across motion boundaries and occlusion. To address these issues, in this paper, we propose a novel method with a new post-processing term and an effective loss function to estimate optical flow in an unsupervised, end-to-end learning manner. Specifically, we first exploit a CNN-based non-local term to refine the estimated optical flow by removing noise and decreasing blur around motion boundaries. This is implemented via automatically learning weights of dependencies over a large spatial neighborhood. Because of its learning ability, the method is effective for various complicated image sequences. Secondly, to reduce the influence of occlusion, a symmetrical energy formulation is introduced to detect the occlusion map from refined bi-directional optical flows. Then the occlusion map is integrated to the loss function. Extensive experiments are conducted on challenging datasets, i.e. FlyingChairs, MPI-Sintel and KITTI to evaluate the performance of the proposed method. The state-of-the-art results demonstrate the effectiveness of our proposed method.
Zhigang Tu 0001, Dejun Zhang, Jun Liu 0036, Baoxin Li, Junsong Yuan 0001
IEEE Trans. Image Process.3
2020 Part-based visual tracking with spatially regularized correlation filters
Dejun Zhang, Lu Zou, Zhuyang Xie, Fazhi He, Yiqi Wu, Zhigang Tu 0001
Vis. Comput.1
2019 SO-HandNet: Self-Organizing Network for 3D Hand Pose Estimation With Semi-Supervised Learning
abstract
3D hand pose estimation has made significant progress recently, where Convolutional Neural Networks (CNNs) play a critical role. However, most of the existing CNN-based hand pose estimation methods depend much on the training set, while labeling 3D hand pose on training data is laborious and time-consuming. Inspired by the point cloud autoencoder presented in self-organizing network (SO-Net), our proposed SO-HandNet aims at making use of the unannotated data to obtain accurate 3D hand pose estimation in a semi-supervised manner. We exploit hand feature encoder (HFE) to extract multi-level features from hand point cloud and then fuse them to regress 3D hand pose by a hand pose estimator (HPE). We design a hand feature decoder (HFD) to recover the input point cloud from the encoded feature. Since the HFE and the HFD can be trained without 3D hand pose annotation, the proposed method is able to make the best of unannotated data during the training phase. Experiments on four challenging benchmark datasets validate that our proposed SO-HandNet can achieve superior performance for 3D hand pose estimation via semi-supervised learning.
Yujin Chen, Zhigang Tu 0001, Liuhao Ge, Dejun Zhang, Ruizhi Chen, Junsong Yuan 0001
ICCV4
2019 Reconstructed similarity for faster GANs-based word translation to mitigate hubness
Dejun Zhang, Mengting Luo, Fazhi He
Neurocomputing1
2019 A survey of variational and CNN-based optical flow techniques
Zhigang Tu 0001, Wei Xie 0008, Dejun Zhang, Ronald Poppe, Remco C. Veltkamp, Baoxin Li, Junsong Yuan 0001
Signal Process. Image Commun.3
2019 Action-Stage Emphasized Spatiotemporal VLAD for Video Action Recognition
abstract
Despite outstanding performance in image recognition, convolutional neural networks (CNNs) do not yet achieve the same impressive results on action recognition in videos. This is partially due to the inability of CNN for modeling long-range temporal structures especially those involving individual action stages that are critical to human action recognition. In this paper, we propose a novel action-stage (ActionS) emphasized spatiotemporal Vector of Locally Aggregated Descriptors (ActionS-STVLAD) method to aggregate informative deep features across the entire video according to adaptive video feature segmentation and adaptive segment feature sampling (AVFS-ASFS). In our ActionSST- VLAD encoding approach, by using AVFS-ASFS, the key frame features are chosen and the corresponding deep features are automatically split into segments with the features in each segment belonging to a temporally coherent ActionS. Then, based on the extracted key frame feature in each segment, a flow-guided warping technique is introduced to detect and discard redundant feature maps, while the informative ones are aggregated by using our exploited similarity weight. Furthermore, we exploit an RGBF modality to capture motion salient regions in the RGB images corresponding to action activity. Extensive experiments are conducted on four public benchmarks - HMDB51, UCF101, Kinetics and ActivityNet for evaluation. Results show that our method is able to effectively pool useful deep features spatiotemporally, leading to state-of-the-art performance for videobased action recognition.
Zhigang Tu 0001, Hongyan Li 0003, Dejun Zhang, Justin Dauwels, Baoxin Li, Junsong Yuan 0001
IEEE Trans. Image Process.3
2018 Service-Oriented Feature-Based Data Exchange for Cloud-Based Design and Manufacturing
abstract
With the rapid development of service-oriented computing (SOC)/service-oriented architecture (SOA), cloud computing and web services, cloud-based design and manufacture (CBDM) is emerging as state-of-the-art technologies and methodologies to enable collaborative product development (CPD). CBDM-enabled CPD can provide cost-effective, flexible and scalable solutions to collaborative partners by sharing the resources in the applications of design and manufacturing. Feature-based data exchange (FBDE) has been one of the key issues in history of CPD and should be adapted in lasted CBDM-enabled CPD. Firstly this paper presents a service-oriented architecture for data exchange in CBDM. Within this architecture, FBDE was registered as service and FBDE users in the CBDM environment can acquire a set of FBDE services to replace the traditional FBDE functions among heterogeneous CAD systems. Secondly, in orderto put the philosophy of FBDE-as-a-Service into practice for CBDM, this paper proposes a peerto peer (P2P) approach for service-oriented FBDE, which revolutionizes the traditional centralized and neutral-file based approach. Thirdly, technique issues of FBDE-as-a-Service in P2P architecture are discussed in details, including constituting of the P2P FBDE service, procedure of service-oriented P2P FBDE, pre-P2P FBDE service, topological entity matching between pre/post-P2P service and post-P2P FBDE service. Finally, a case study of data exchange is tested to demonstrate the proposed idea of service-oriented FBDE for CBDM.
Yiqi Wu, Fazhi He, Dejun Zhang
IEEE Trans. Serv. Comput.3
2015 Feature-based data exchange as Service for Cloud Based Design and Manufacturing
abstract
Feature-based data exchange (FBDE) for heterogeneous CAD systems is one of the key issues in Collaborative Product Development (CPD) which is now enabled by Cloud-Based Design and Manufacturing (CBDM). Firstly this paper presents a FBDE-as-a-Service architecture for data exchange in CBDM. Within this architecture, FBDE users in the CBDM environment can acquire a set of Peer to Peer (P2P) services to realize the FBDE functions among heterogeneous CAD systems. Secondly, in order to integrate FBDE services into CBDM, we present a P2P approach for FBDE, which is totally different from traditional centralized and neutral file-based approach. Thirdly, some key issues of FBDE-as-a-Service, such as Pre-P2P FBDE service, Post-P2P FBDE service and topological entity matching are researched. Finally, a prototype system of FBDE-as-a-Service for data exchange is implemented to demonstrate the proposed ideas.
Yiqi Wu, Fazhi He, Dejun Zhang
CSCWD3
2014 Product data exchange of complex shape based on parametric curve
abstract
Product data exchange is one of most important key issues in Collaborative Product Development. Since feature-based parametric CAD systems have been dominated by industrial applications, Feature-Based Data Exchange (FBDE) is getting real growth. However the main feature-based method still lacks the ability to exchange complex shape among the heterogeneous CAD Systems. This paper attacks the problem by exchanging the parametric curve (such as Spline) which is sketched in 2D and will be used to prepare complex 3D shapes by various extrusion features. Therefore the data exchange of complex shape is divided into two layers: 3D extrusion layer and 2D sketch layer. The exchange of spline proceed in the 2D sketch layer between different CAD systems. We innovatively convert the problem of spline exchange into the problem of spline fitting, and employ the Genetic Algorithm (GA) to solve the problem. A new coding strategy in stage of initialization population is presented to improve the GA so it works well with the spline fitting among the heterogeneous CAD systems. Finally, a Hausdorff Distance (HD) is adopted to calculate the fitness. Experimental results demonstrate the effectiveness of our method.
Dejun Zhang, Fazhi He, Yiqi Wu, Xiantao Cai
CSCWD1
2013 A group Undo/Redo method in 3D collaborative modeling systems with performance evaluation
Yuan Cheng 0001, Fazhi He, Xiantao Cai, Dejun Zhang
J. Netw. Comput. Appl.4
2012 A selective undo/redo method in 3D collaborative modeling environment
abstract
In 3D collaborative modeling systems, users need a convenient mechanism to repeatedly modify the models they are operating on. In this paper, we contribute a selective undo/redo solution for users to select arbitrary operation to undo. With the consistency maintainence mechanism we proposed, operations need to be re-arranged on each site for after their arriving. Both history buffer and model state stream are adopted to present the arriving sequence of operations and their actual execution sequence. In case of concurrent undo/redo, undo state vector is proposed to make sure that an operation can only be undone once and redone by the designer who undoes it. Based on all the precautions we have made, an undo/redo algorithm is proposed. The algorithm has been verified in the prototype we implemented.
Yuan Cheng 0001, Xiantao Cai, Fazhi He, Dejun Zhang
CSCWD4
2012 Consistency maintenance based on the matching of topological entity
abstract
Consistency maintenance is one of the most important problems in collaborative CAD systems. However, existing consistency maintenance mechanisms limit multi-user interaction. This paper presents a consistency maintenance method to gain a less-constrained multi-user interaction. First, the causal relation between modeling operations is preserved using the state vector. Then, the concurrent deletion operations are checked to decide if the current operation is masked. If not, the solution for topological entities' matching is adopted to deal with the operations that use topological entities. Then, for those operations which do not use topological entities, the corresponding mechanism is adopted according to their types. By these mechanisms, the commutative, masked and conflicted relations between the concurrent operations, are explored and the conflicts are solved. The experiments prove that our method can support less-constrained multi-user interaction.
Fazhi He, Xiantao Cai, Dejun Zhang
CSCWD4