VLDB 2026 Research / reviewers in the wild / expert
Xiaoheng Jiang
dblp:159/8743
· DBLP profile ↗
53ranked-venue papers
11as first author
38since 2021 · last 2026
0000-0002-5770-0417ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 6 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 21 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deviation capture networks for anomaly detection
Jiawei Cheng, Yang Lu 0016, Wenjie Zhang 0008, Xiaoheng Jiang, Mingliang Xu 0001 |
Adv. Eng. Informatics | 6 |
| 2026 | Global context guided refinement and aggregation network for lightweight surface defect detection
Xiaoheng Jiang, Yang Lu 0016, Lisha Cui, Jiale Cao, Mingliang Xu 0001 |
Pattern Recognit. | 2 |
| 2026 | To the Best of Trust: Full-Stage Trusted Multi-Modal ClusteringabstractMulti-modal clustering aims to integrate complementary information from different modalities to uncover latent consistent structures and improve clustering performance. However, existing methods mainly rely on predictive (result) uncertainty to improve robustness, while often neglecting aleatoric (data) uncertainty introduced by sample noise and epistemic (model) uncertainty induced by model parameters and structural variations. To this end, we propose a novel Full-Stage Trusted Multi-modal Clustering (FSTMC) method. To the best of trust, we jointly utilize aleatoric, epistemic, and predictive uncertainties to optimize the model and learn more reliable feature representations and clustering results. In the representation learning phase, probabilistic modeling is used to capture stable latent representations and estimate aleatoric uncertainty, while structured random perturbations are present to estimate epistemic uncertainty. In the clustering stage, instead of conventional feature-level fusion, we design an evidence-based fusion strategy, where soft labels from each modality are first mapped into categorical evidence while cluster distributions are parameterized via a Dirichlet model, with finally dynamic multi-modal fusion achieved by Dempster-Shafer theory. To mitigate overconfidence and modal conflicts, prior constraints guided by aleatoric and epistemic uncertainty are imposed, resulting in calibrated predictive uncertainty. Finally, we exploit predictive uncertainty to selectively incorporate pseudo labels for optimization. Benchmark experiments on a number of multi-modal datasets demonstrate that our approach significantly improves accuracy compared to state-of-the-art methods. Shizhe Hu, Yucong Wu, Jinlan Wang, Xiaoheng Jiang, Pei Lv, Mingliang Xu 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | EchoDiffusion: Waveform Conditioned Diffusion Models for Echo-Based Depth EstimationabstractTo extract spatial information, depth estimation using conventional echo-based methods typically employs models with encoder-decoder architectures, such as UNet. However, these methods may face challenges in extracting fine details from echo waveforms and handling multi-scale feature extraction with high precision. To address these challenges, we introduce EchoDiffusion, a framework that incorporates diffusion models conditioned on waveform embeddings for echo-based depth estimation. This framework employs the Multi-Scale Adaptive Latent Feature Network (MALF-Net) to extract multi-scale spatial features and perform adaptive fusion, encoding the echo spectrograms into the latent space. Additionally, we propose the Echo Waveform Detail Embedder (EWDE), which leverages a pre-trained Wav2Vec model to extract detailed spatial information from echo waveforms, using these details as conditional inputs to guide the reverse diffusion process in the latent space. By embedding the echo waveforms into the reverse diffusion process, we can more accurately guide the generation of depth maps. Our extensive evaluations on the Replica and Matterport3D datasets demonstrate that EchoDiffusion establishes new benchmarks for state-of-the-art performance in echo-based depth estimation. Wenjie Zhang 0008, Xiaoheng Jiang, Zhen Tian 0004, Mingliang Xu 0001 |
AAAI | 5 |
| 2025 | LAGNet: A Location-Aware Guidance Network for Weak and Strip Defect Detection
Lisha Cui, Helong Jiao, Tengyue Liu, Chunyan Niu, Xiaoheng Jiang, Mingliang Xu 0001 |
CVM (1) | 6 |
| 2025 | Wavelet and Prototype Augmented Query-based Transformer for Pixel-level Surface Defect DetectionabstractAs an important part of intelligent manufacturing, pixel-level surface defect detection (SDD) aims to locate defect areas through mask prediction. Previous methods adopt the image-independent static convolution to indiscriminately classify per-pixel features for mask prediction, which leads to suboptimal results for some challenging scenes such as weak defects and cluttered backgrounds. In this paper, inspired by query-based methods, we propose a Wavelet and Prototype Augmented Query-based Transformer (WP-Former) for surface defect detection. Specifically, a set of dynamic queries for mask prediction is updated through the dual-domain transformer decoder. Firstly, a Wavelet-enhanced Cross-Attention (WCA) is proposed, which aggregates meaningful high-and low-frequency information of image features in the wavelet domain to refine queries. WCA enhances the representation of high-frequency components by capturing multi-scale relationships between different frequency components, enabling queries to focus more on defect details. Secondly, a Prototype-guided Cross-Attention (PCA) is proposed to refine queries through meta-prototypes in the spatial domain. The prototypes aggregate semantically meaningful tokens from image features, facilitating queries to aggregate crucial defect information under the cluttered backgrounds. Extensive experiments on three defect detection datasets (i.e., ESDIs-SOD, CrackSeg9k, and ZJU-Leaper) demonstrate that the proposed method achieves state-of-the-art performance in defect detection. The code will be available at https://github.com/yfhdm/WPFormer. Xiaoheng Jiang, Yang Lu 0016, Jiale Cao, Dong Chen 0017, Mingliang Xu 0001 |
CVPR | 2 |
| 2025 | CLIPeR: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic SegmentationabstractContrastive Language-Image Pre-training (CLIP) exhibits strong zero-shot classification ability on various image-level tasks, leading to the research to adapt CLIP for pixel-level open-vocabulary semantic segmentation without additional training. The key is to improve spatial representation of image-level CLIP, such as replacing self-attention map at last layer with self-self attention map or vision foundation model based attention map. In this paper, we present a novel hierarchical framework, named CLIPer, that hierarchically improves spatial representation of CLIP. The proposed CLIPer includes an early-layer fusion module and a fine-grained compensation module. We observe that, the embeddings and attention maps at early layers can preserve spatial structural information. Inspired by this, we design the early-layer fusion module to generate segmentation map with better spatial coherence. Afterwards, we employ a fine-grained compensation module to compensate the local details using the self-attention maps of diffusion model. We conduct the experiments on seven segmentation datasets. Our proposed CLIPer achieves the state-of-the-art performance on these datasets. For instance, using ViT-L, CLIPer has the mIoU of 69.8% and 43.3% on VOC and COCO Object, outperforming ProxyCLIP by 9.2% and 4.1% respectively. Jiale Cao, Jin Xie 0005, Xiaoheng Jiang, Yanwei Pang |
ICCV | 4 |
| 2025 | DefGAN-Im: Adversarial Industrial Defect Synthesis with Conditional Feature Disentanglement on Imbalanced Datasets
Xiaofei Nan, Jinyu Fan, Wenyang Li, Xiaoheng Jiang |
ICIC (11) | 6 |
| 2025 | Bounding Box-Derived Mask Guidance Network for Accurate Surface Defect Detection
Lisha Cui, Fengye Tian, Xiaoheng Jiang, Mingliang Xu 0001 |
PRCV (18) | 5 |
| 2025 | LGGFormer: A dual-branch local-guided global self-attention network for surface defect segmentation
Yang Lu 0016, Xiaoheng Jiang, Shaohui Jin, Shupan Li, Mingliang Xu 0001 |
Adv. Eng. Informatics | 3 |
| 2025 | RRGMambaFormer: A hybrid Transformer-Mamba architecture for radiology report generation
Hongzhao Li, Siwei Liu 0001, Xiaoheng Jiang, Mingyuan Jiu, Yang Lu 0016, Shupan Li, Mingliang Xu 0001 |
Expert Syst. Appl. | 4 |
| 2025 | NN2ViT: Neural Networks and Vision Transformers based approach for Visual Anomaly Detection in Industrial ImagesabstractEnsuring product quality through automated anomaly detection is crucial in manufacturing. Traditional methods often struggle to capture both local and global features effectively, relying heavily on predefined templates that limit their adaptability and accuracy. To address these challenges, this study propose NN2ViT, a novel approach that integrates a Single Shot Detector (SSD) for local feature detection and the Segment Anything Model (SAM) for global feature segmentation. This integration allows for a comprehensive analysis of anomalies in industrial images. Our method improves anomaly segmentation performance by fine-tuning SAM for precise segmentation in industrial product images. Experiments on the MVTec benchmark dataset demonstrate that NN2ViT outperforms traditional models and achieved the highest 95.54% and 96.23% Image AUROC and AP scores, respectively thus enhancing interpretability and adaptability to various anomaly patterns. This research presents a significant advancement in manufacturing quality control , contributing to improved product quality and operational efficiency. Junaid Abdul Wahid, Muhammad Ayoub, Mingliang Xu 0001, Xiaoheng Jiang, Lei Shi 0001, Shabir Hussain |
Neurocomputing | 4 |
| 2025 | Hierarchical Differential Attention for Multimodal Relation Extraction
Xiaoheng Jiang, Yang Lu 0016, Kunli Zhang, Mingliang Xu 0001 |
Knowl. Based Syst. | 2 |
| 2025 | From a perceptual perspective: No-Reference Image Quality Assessment using Dual Perception Hybrid Network
Yang Lu 0016, Zifan Yang, Zilu Zhou, Xiaoheng Jiang, Mingliang Xu 0001 |
Pattern Recognit. Lett. | 5 |
| 2025 | Enhanced multiscale attentional feature fusion model for defect detection on steel surfaces
Yongkai Xia, Yang Lu 0016, Xiaoheng Jiang, Mingliang Xu 0001 |
Pattern Recognit. Lett. | 3 |
| 2025 | Industrial product quality assessment using deep learning with defect attributes
Yang Lu 0016, Xiaoheng Jiang, Mingliang Xu 0001 |
Pattern Recognit. Lett. | 3 |
| 2025 | An Efficient Ungrouped Mask Method With two Learnable Parameters for 3D Object DetectionabstractIn 3D point cloud-based object detection, attention mechanism in Group-Free [1] learns direct relationships between proposals and all seed points, providing each proposal with a global context in the form of a cross-attention map. However, our analysis and experimental comparison show that the attention mechanism assigns inappropriately large attention weights to certain seed points far from a proposal, which is not conducive to detecting objects correctly. In this work, we alleviate the above problem by proposing a mask method. For an initial proposal, our method first calculates a spatial distance-based mask, which measures the spatial relationship between all seed points and the proposal. Then, we fuse the mask into cross-attention layers in stacked attention modules and get a refined cross-attention map. In essence, our mask gives each proposal a local context; after it is fused with the global context given by the attention mechanism, the refined cross-attention map could suppress the negative impact of some distant seed points on a proposal. We present two alternative strategies to compute the mask, a hard mask, and a soft mask. Experimental results demonstrate that the soft mask brings better performance. In the soft mask, for each initial proposal's 3D-box shape, we use a parametric approximate ellipsoid as the basis of the mask's calculation, which has only two learnable parameters. Experimental results show our work could outperform Group-Free 0.7 [email protected] at the cost of increasing inference time by less than 1%. The performance of our algorithm on the public dataset SUN RGB-D is 63.7 [email protected] and 45.5 [email protected], which is the best performance among algorithms that preserve the irregular of seed points. Shuai Guo 0004, Lei Shi 0001, Xiaoheng Jiang, Pei Lv, Qidong Liu 0001, Yazhou Hu, Rongrong Ji, Mingliang Xu 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | MGDefect: A Mask-Guided High-Quality Defect Image Generation Method for Improving Defect InspectionabstractDeep learning based defect inspection methods have achieved promising performance, which usually relies on a large number of well-labeled training samples. However, it requires much effort to obtain enough annotated samples especially pixel-level annotations in practical production. Generative adversarial networks (GANs) can be utilized to generate defect samples. However, training GANs typically requires a large amount of defect data and most of them cannot generate defect samples with pixel-level annotations. In this paper, we present a Mask-Guided Defect image generation method, called MGDefect, which can generate high-quality defect samples with pixel-level annotations and effectively improves the performance of downstream tasks. Specifically, MGDefect consists of a Mask-Guided Defect Generation GAN (MGDG-GAN) and a Defect Mask GAN (DM-GAN). MGDG-GAN generates images containing defects with specific locations, shapes, and sizes via mask guidance and the dual discrimination for defects at the region level and image level. DM-GAN aims to generate diverse and rational masks for MGDG-GAN. It also adopts region-level and image-level dual discrimination for masks to generate compatible masks with the target objects. MGDG-GAN mainly focuses on generating local defect regions and DM-GAN specializes in generating masks, which are both trained on limited defect samples and abundant normal samples. Experiments conducted on the MVTec AD, DAGM 2007, and KolektorSDD2 benchmark datasets demonstrate that our method achieves promising results compared with other state-of-the-art approaches. Meanwhile, the generated defect samples significantly improve the performance of defect inspection tasks including classification and segmentation. Specifically, our method achieves KID$\times 10^{3}$/IS scores of 48.35/2.27 on MVTec AD, 15.37/2.44 on DAGM 2007, and 19.70/2.01 on KolektorSDD2. Furthermore, our method improves mIoU by 10.59%, 2.20%, and 2.17% on these datasets, respectively, using U-Net as the segmentation model. Xiaoheng Jiang, Yang Lu 0016, Changsheng Xu, Mingliang Xu 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | DefectSAM: Hierarchically Adapting SAM for Pixel-Wise Surface Defect DetectionabstractSegment anything model (SAM) has recently demonstrated powerful segmentation ability for natural scene images (NSIs). However, the SAM exhibits limited performance in defect detection owing to the weak appearance of defects and cluttered backgrounds in industrial images. In this article, we propose a hierarchically adapting SAM for pixel-wise surface defect detection, named DefectSAM, which effectively modulates and decodes multilevel features of the encoder to capture defect information. Specifically, we introduce a learnable feature adaptation component between the image encoder and the decoder to modulate each level of features via the dual-feature adaptation unit. The dual-feature adaptation unit mainly includes the correlation-gated feature adaptation (CGFA) module and the mask-guided feature adaptation (MGFA) module. The CGFA exploits cross correlation spatial gating maps to adaptively incorporate a convolutional feature pyramid and Transformer features during feature adaptation, which is beneficial for capturing defect details. Moreover, the MGFA utilizes the mask prediction of high-level features as semantic guidance to select top-confidence foreground and background tokens for feature adaptation, focusing more on defect details and suppressing background noise. Extensive experiments on three defect detection datasets (i.e., MVTec AD, CrackSeg9k, ZJU-Leaper, and Magnetic tile) demonstrate that the proposed method achieves state-of-the-art performance with few learnable parameters, which greatly improves the generalization of SAM in defect detection. Xiaoheng Jiang, Yang Lu 0016, Jiale Cao, Mingliang Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Echo Depth Estimation via Attention-based Hierarchical Multi-scale Feature Fusion NetworkabstractIn environments where vision-based depth estimation systems, such as those utilizing infrared or imaging technologies, encounter limitations—particularly in low-light conditions—alternative approaches become essential. Echo depth estimation emerges as a compelling solution by leveraging the time delay of echoes to map the geometric structure of the surrounding environment. This method offers distinct advantages in specific scenarios, providing reliable data for accurate scene understanding and 3D reconstruction. Traditional echo depth estimation techniques primarily depend on spatial information captured by the encoder and depth predictions made by the decoder. However, these methods often fail to fully exploit the rich depth features present at different simultaneous frequencies. To address this challenge, we propose an echo depth estimation method via Attention-based Hierarchical Multi-scale Feature Fusion Network (AHMF-Net). This network is designed to extract spatial depth information from echo spectrograms across multiple scales and hierarchical levels, while fusing the most relevant information using an attention mechanism. AHMF-Net introduces two key modules in hierarchical levels: the Intra-layer Multi-scale Attention Feature Fusion (IMAF) module, which functions as the encoder to capture multi-scale features across varying granularities, and the Inter-layer Multi-Scale Detail Feature Fusion (IMDF) module, which integrates features from all encoding layers into the decoder to enable effective inter-layer multi-scale fusion. Additionally, the encoder incorporates an attention mechanism that enhances depth-related features by capturing channel dependencies at multiple scales. We evaluated AHMF-Net on the Replica, Matterport3D, and BatVision datasets, where it consistently outperformed state-of-the-art models in echo-based depth estimation, demonstrating superior accuracy and robustness. The source code is publicly available at https://github.com/wjzhang-ai/AHMF-Net . Wenjie Zhang 0008, Yibo Guo, Xiaoheng Jiang, Shaohui Jin, Mingliang Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | End-to-End Transformer Architecture with Novel Ensemble Learning Method Integrating CT Scans and Clinical Narratives for Brain Stroke Diagnosis
Junaid Abdul Wahid, Muhammad Ayoub, Mingliang Xu 0001, Xiaoheng Jiang, Lei Shi 0001, Lifeng Li, Shabir Hussain, Ashfaque Khawaja, Yufei Gao 0001 |
CGI (3) | 4 |
| 2024 | ID-like Prompt Learning for Few-Shot Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection methods often exploit auxiliary outliers to train model identifying OOD samples, especially discovering challenging outliers from auxiliary outliers dataset to improve OOD detection. However, they may still face limitations in effectively distinguishing between the most challenging OOD samples that are much like in-distribution (ID) data, i.e., ID-like samples. To this end, we propose a novel OOD detection framework that discovers ID-like outliers using CLIP [32]from the vicinity space of the ID samples, thus helping to identify these most challenging OOD samples. Then a prompt learning framework is proposed that utilizes the identified ID-like outliers to further leverage the capabilities of CLIP for OOD detection. Benefiting from the powerful CLIP, we only need a small number of ID samples to learn the prompts of the model without exposing other auxiliary outlier datasets. By focusing on the most challenging ID-like OOD samples and elegantly exploiting the capabilities of CLIP, our method achieves superior few-shot learning performance on various real-world image datasets (e.g., in 4-shot OOD detection on the ImageNet-1k dataset, our method reduces the average FPR95 by 12.16% and improves the average AUROC by 2.76%, compared to state-of-the-art methods). Code is available at https://github.com/ycfate/ID-like Yichen Bai, Zongbo Han, Bing Cao 0002, Xiaoheng Jiang, Qinghua Hu, Changqing Zhang 0002 |
CVPR | 4 |
| 2024 | Context Mutual Evolution Network for Weakly Supervised Surface Defect Detection
Xiaoheng Jiang, Penghui Xiao, Yang Lu 0016, Shaohui Jin, Mingliang Xu 0001 |
ICPR (10) | 1 |
| 2024 | Corner Detection: Passive Non-Lin-of-Sight Pedestrian Detection
Shaohui Jin, Xiaoheng Jiang, Jiyue Wang, Hao Liu 0125, Mingliang Xu 0001 |
PRCV (9) | 4 |
| 2024 | Adaptive Dual Attention Fusion Network for RGB-D Surface Defect Detection
Xiaoheng Jiang, Jingqi Liu, Yang Lu 0016, Shaohui Jin, Hao Liu 0125, Mingliang Xu 0001 |
PRCV (9) | 1 |
| 2024 | DiffuSaliency: Synthesizing Multi-object Images with Masks for Semantic Segmentation Using Diffusion and Saliency Detection
Yunhan Jiang, Xianglong Shi, Xiaoheng Jiang, Yang Lu 0016, Mingliang Xu 0001 |
PRCV (3) | 3 |
| 2024 | Sparse Context Transformer for Few-Shot Object Detection
Mingyuan Jiu, Hichem Sahbi, Xiaoheng Jiang, Mingliang Xu 0001 |
PRICAI (4) | 4 |
| 2024 | Hierarchical symmetric cross entropy for distant supervised relation extraction
Xiaoheng Jiang, Pengshuai Lv, Yang Lu 0016, Shupan Li, Kunli Zhang, Mingliang Xu 0001 |
Appl. Intell. | 2 |
| 2024 | Automated Audio Data Augmentation Network Using Bi-Level Optimization for Sound Event Localization and DetectionabstractIn sound event localization and detection (SELD), traditional methods often treat localization and detection algorithms separately from data augmentation. During the model training process, the strategy for data augmentation is typically implemented in a non-learnable manner. Existing audio data augmentation strategies struggle to find optimal parameter solutions for data augmentation that can be effectively applied to SELD systems. To address this challenge, we introduce an innovative network-based strategy, termed the Automated Audio Data Augmentation (AADA) network. This strategy employs bi-level optimization to synergistically integrate audio data augmentation techniques with SELD tasks. In the AADA network, the lower-level SELD task serves as a constraint for the higher-level data augmentation process. The audio data augmentation parameters are adaptively optimized by utilizing the transfer of intermediate feature information from the SELD tasks, thus obtaining optimal parameters for these tasks. Evaluation of our approach on the Sony-TAU Realistic Spatial Soundscapes 2023 dataset achieves a SELD score of 0.4801, significantly surpassing the performance metrics of all traditional data augmentation strategies for SELD. Wenjie Zhang 0008, Xiaoheng Jiang, Mingliang Xu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2023 | DFAR-Net: Dual-Input Three-Branch Attention Fusion Reconstruction Network for Polarized Non-Line-of-Sight Imaging
Hao Liu 0125, Ke Wang 0064, Shaohu Jin, Pengyun Chen, Xiaoheng Jiang, Mingliang Xu 0001 |
PRCV (6) | 7 |
| 2023 | Client Selection and Resource Allocation for Federated Learning in Digital-Twin-Enabled Industrial Internet of ThingsabstractThe technological advancements in digital twins (DTs) and next-generation networks boost the development of the Industrial Internet of Things (IIoT). However, the increasing concern for data privacy hinders the further development of IIoT. By manipulating its own dataset and training machine learning models locally, federated learning (FL) has been regarded as a promising technology to establish DT models for IIoT. Nevertheless, the vast amount of model parameters, limited computation and communication resources, and the dynamic industrial environments pose great challenges for FL. In this paper, we present a novel FL framework integrating with DTs in the resource-constrained and dynamic IIoT environment. The clustering algorithm is exploited to design an efficient client selection approach. Furthermore, we propose an efficient model aggregation approach to improve the training efficiency. A resource allocation problem is then formulated and solved based on deep reinforcement learning. Numerical results show that the proposed client selection and resource allocation approach reduces the time cost compared with the baseline schemes. Shuo He 0002, Tianxiang Ren, Xiaoheng Jiang, Mingliang Xu 0001 |
WCNC | 3 |
| 2023 | User-Guided Personalized Image Aesthetic Assessment Based on Deep Reinforcement LearningabstractPersonalized image aesthetic assessment (PIAA) has recently become a hot topic due to its wide applications, such as photography, film, television, e-commerce, fashion design, and so on. This task is more seriously affected by subjective factors and samples provided by users. In order to acquire precise personalized aesthetic distribution by small amount of samples, we propose a novel user-guided personalized image aesthetic assessment framework. This framework leverages user interactions to retouch and rank images for aesthetic assessment based on deep reinforcement learning (DRL), and generates personalized aesthetic distribution that is more in line with the aesthetic preferences of different users. It mainly consists of two stages. In the first stage, personalized aesthetic ranking is generated by interactive image enhancement and manual ranking, meanwhile, two policy networks will be trained. These two networks will be trained iteratively and alternatively to facilitate the final personalized aesthetic assessment. In the second stage, these modified images are labeled with aesthetic attributes by one style-specific classifier, and then the personalized aesthetic distribution is generated based on the multiple aesthetic attributes of these images, which conforms to the aesthetic preference of users better. Compared with other existing methods, our approach has achieved new state-of-the-art in the task of personalized image aesthetic assessment on the public AVA and FLICKR-AES datasets. Pei Lv, Jianqi Fan, Xixi Nie, Weiming Dong, Xiaoheng Jiang, Bing Zhou 0003, Mingliang Xu 0001, Changsheng Xu |
IEEE Trans. Multim. | 5 |
| 2022 | Focal and Global Spatial-Temporal Transformer for Skeleton-Based Action Recognition
Zhimin Gao, Peitao Wang, Pei Lv, Xiaoheng Jiang, Qidong Liu 0001, Pichao Wang, Mingliang Xu 0001, Wanqing Li 0001 |
ACCV (4) | 4 |
| 2022 | Trajectory distributions: A new description of movement for trajectory predictionabstractTrajectory prediction is a fundamental and challenging task for numerous applications, such as autonomous driving and intelligent robots. Current works typically treat pedestrian trajectories as a series of 2D point coordinates. However, in real scenarios, the trajectory often exhibits randomness, and has its own probability distribution. Inspired by this observation and other movement characteristics of pedestrians, we propose a simple and intuitive movement description called a trajectory distribution, which maps the coordinates of the pedestrian trajectory to a 2D Gaussian distribution in space. Based on this novel description, we develop a new trajectory prediction method, which we call the social probability method . The method combines trajectory distributions and powerful convolutional recurrent neural networks. Both the input and output of our method are trajectory distributions, which provide the recurrent neural network with sufficient spatial and random information about moving pedestrians. Furthermore, the social probability method extracts spatio-temporal features directly from the new movement description to generate robust and accurate predictions. Experiments on public benchmark datasets show the effectiveness of the proposed method. Pei Lv, Tianxin Gu, Xiaoheng Jiang, Bing Zhou 0003, Mingliang Xu 0001 |
Comput. Vis. Media | 5 |
| 2022 | Transferring priors from virtual data for crowd counting in real world
Xiaoheng Jiang, Hao Liu 0057, Li Zhang 0072, Geyang Li, Mingliang Xu 0001, Pei Lv, Bing Zhou 0003 |
Frontiers Comput. Sci. | 1 |
| 2022 | Context-Aware Block Net for Small Object DetectionabstractState-of-the-art object detectors usually progressively downsample the input image until it is represented by small feature maps, which loses the spatial information and compromises the representation of small objects. In this article, we propose a context-aware block net (CAB Net) to improve small object detection by building high-resolution and strong semantic feature maps. To internally enhance the representation capacity of feature maps with high spatial resolution, we delicately design the context-aware block (CAB). CAB exploits pyramidal dilated convolutions to incorporate multilevel contextual information without losing the original resolution of feature maps. Then, we assemble CAB to the end of the truncated backbone network (e.g., VGG16) with a relatively small downsampling factor (e.g., 8) and cast off all following layers. CAB Net can capture both basic visual patterns as well as semantical information of small objects, thus improving the performance of small object detection. Experiments conducted on the benchmark Tsinghua-Tencent 100K and the Airport dataset show that CAB Net outperforms other top-performing detectors by a large margin while keeping real-time speed, which demonstrates the effectiveness of CAB Net for small object detection. Lisha Cui, Pei Lv, Xiaoheng Jiang, Zhimin Gao, Bing Zhou 0003, Ling Shao 0001, Mingliang Xu 0001 |
IEEE Trans. Cybern. | 3 |
| 2021 | Multi-label correlation guided feature fusion network for abnormal ECG diagnosis
Zhaoyang Ge, Xiaoheng Jiang, Zhuang Tong, Panpan Feng, Bing Zhou 0003, Mingliang Xu 0001, Zongmin Wang, Yanwei Pang |
Knowl. Based Syst. | 2 |
| 2021 | Density-Aware Multi-Task Learning for Crowd CountingabstractIn this paper, we present a method called density-aware convolutional neural network (DensityCNN) to perform the crowd counting task in various crowded scenes. The key idea of the DensityCNN is to utilize high-level semantic information to provide guidance and constraint when generating density maps. To this end, we implement the DensityCNN by adopting a multi-task CNN structure to jointly learn density-level classification and density map estimation. The density-level classification task learns multi-channel semantic features that are aware of the density distributions of the input image. This task is accomplished via our specially designed group-based convolutional structure in a supervised learning manner. In the density map estimation task, these semantic features are deployed together with high-dimension convolutional features to generate density maps with lower count errors. Extensive experiments on four challenging crowd datasets (ShanghaiTech, UCF_CC_50, UCF-QNCF, and WorldExpo'10) and one vehicle dataset TRANCOS demonstrate the effectiveness of the proposed method. Xiaoheng Jiang, Li Zhang 0072, Tianzhu Zhang 0001, Pei Lv, Bing Zhou 0003, Yanwei Pang, Mingliang Xu 0001, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2020 | Attention Scaling for Crowd CountingabstractConvolutional Neural Network (CNN) based methods generally take crowd counting as a regression task by outputting crowd densities. They learn the mapping between image contents and crowd density distributions. Though having achieved promising results, these data-driven counting networks are prone to overestimate or underestimate people counts of regions with different density patterns, which degrades the whole count accuracy. To overcome this problem, we propose an approach to alleviate the counting performance differences in different regions. Specifically, our approach consists of two networks named Density Attention Network (DANet) and Attention Scaling Network (ASNet). DANet provides ASNet with attention masks related to regions of different density levels. ASNet first generates density maps and scaling factors and then multiplies them by attention masks to output separate attention-based density maps. These density maps are summed to give the final density map. The attention scaling factors help attenuate the estimation errors in different regions. Furthermore, we present a novel Adaptive Pyramid Loss (APLoss) to hierarchically calculate the estimation losses of sub-regions, which alleviates the training bias. Extensive experiments on four challenging datasets (ShanghaiTech Part A, UCF_CC_50, UCF-QNRF, and WorldExpo'10) demonstrate the superiority of the proposed approach. Xiaoheng Jiang, Li Zhang 0072, Mingliang Xu 0001, Tianzhu Zhang 0001, Pei Lv, Bing Zhou 0003, Xin Yang 0011, Yanwei Pang |
CVPR | 1 |
| 2020 | MDSSD: multi-scale deconvolutional single shot detector for small objects
Lisha Cui, Pei Lv, Xiaoheng Jiang, Zhimin Gao, Bing Zhou 0003, Mingliang Xu 0001 |
Sci. China Inf. Sci. | 4 |
| 2020 | Handwriting posture prediction based on unsupervised model
Bailin Yang, Zhenguang Liu, Xiaoheng Jiang, Mingliang Xu 0001 |
Pattern Recognit. | 4 |
| 2020 | Learning Multi-Level Density Maps for Crowd CountingabstractPeople in crowd scenes often exhibit the characteristic of imbalanced distribution. On the one hand, people size varies largely due to the camera perspective. People far away from the camera look smaller and are likely to occlude each other, whereas people near to the camera look larger and are relatively sparse. On the other hand, the number of people also varies greatly in the same or different scenes. This article aims to develop a novel model that can accurately estimate the crowd count from a given scene with imbalanced people distribution. To this end, we have proposed an effective multi-level convolutional neural network (MLCNN) architecture that first adaptively learns multi-level density maps and then fuses them to predict the final output. Density map of each level focuses on dealing with people of certain sizes. As a result, the fusion of multi-level density maps is able to tackle the large variation in people size. In addition, we introduce a new loss function named balanced loss (BL) to impose relatively BL feedback during training, which helps further improve the performance of the proposed network. Furthermore, we introduce a new data set including 1111 images with a total of 49 061 head annotations. MLCNN is easy to train with only one end-to-end training stage. Experimental results demonstrate that our MLCNN achieves state-of-the-art performance. In particular, our MLCNN reaches a mean absolute error (MAE) of 242.4 on the UCF_CC_50 data set, which is 37.2 lower than the second-best result. Xiaoheng Jiang, Li Zhang 0072, Pei Lv, Yibo Guo, Ruijie Zhu 0001, Yanwei Pang, Xi Li 0001, Bing Zhou 0003, Mingliang Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Disassembling a 3D mechanism for efficient packingabstractAbstract This paper introduces a disassemble‐and‐pack algorithm to disassemble a mechanical 3D model in groups that can be efficiently packed within a box, with the objective of reassembling them easily after delivery. Its key feature is that, mostly, the mechanism can be disassembled at the joint and each part can be an adjusted motion structure based on its joint type. Our system consists of two steps: disassembling the mechanical object into a group set and packing them within a box efficiently. The first step consists in the creation of a hierarchy of possible group set of parts that can be tightly packed within their minimum bounding boxes. Use the breadth‐first search algorithm to traverse the hierarchy of possible group set in order to disconnect the joints and get the group set. In the second step, according to the reverse order of volume, each group in the set is inserted into the specified box. The fact that mechanism disassembly and shape packing are both an NP‐complete problem justifies finding approximated solutions according to efficacy and efficiency. Experimental results show that our approach can really efficiently pack a range of mechanisms from a simple model to complex objects. Xiaoheng Jiang, Ning-Bo Gu, Weiwei Xu 0003, Junxiao Xue, Bing Zhou 0003, Mingliang Xu 0001 |
Comput. Animat. Virtual Worlds | 2 |
| 2019 | Depth Information Guided Crowd Counting for complex crowd scenes
Mingliang Xu 0001, Zhaoyang Ge, Xiaoheng Jiang, Gaoge Cui, Pei Lv, Bing Zhou 0003, Changsheng Xu |
Pattern Recognit. Lett. | 3 |
| 2019 | Traffic Simulation and Visual Verification in SmogabstractSmog causes low visibility on the road and it can impact the safety of traffic. Modeling traffic in smog will have a significant impact on realistic traffic simulations. Most existing traffic models assume that drivers have optimal vision in the simulations, making these simulations are not suitable for modeling smog weather conditions. In this article, we introduce the Smog Full Velocity Difference Model (SMOG-FVDM) for a realistic simulation of traffic in smog weather conditions. In this model, we present a stadia model for drivers in smog conditions. We introduce it into a car-following traffic model using both psychological force and body force concepts, and then we introduce the SMOG-FVDM. Considering that there are lots of parameters in the SMOG-FVDM, we design a visual verification system based on SMOG-FVDM to arrive at an adequate solution which can show visual simulation results under different road scenarios and different degrees of smog by reconciling the parameters. Experimental results show that our model can give a realistic and efficient traffic simulation of smog weather conditions. Mingliang Xu 0001, Shili Chu, Yong Gan, Xiaoheng Jiang, Bing Zhou 0003 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2019 | Motion-Aware Compression and Transmission of Mesh Animation SequencesabstractWith the increasing demand in using 3D mesh data over networks, supporting effective compression and efficient transmission of meshes has caught lots of attention in recent years. This article introduces a novel compression method for 3D mesh animation sequences, supporting user-defined and progressive transmissions over networks. Our motion-aware approach starts with clustering animation frames based on their motion similarities, dividing a mesh animation sequence into fragments of varying lengths. This is done by a novel temporal clustering algorithm, which measures motion similarity based on the curvature and torsion of a space curve formed by corresponding vertices along a series of animation frames. We further segment each cluster based on mesh vertex coherence, representing topological proximity within an object under certain motion. To produce a compact representation, we perform intra-cluster compression based on Graph Fourier Transform (GFT) and Set Partitioning In Hierarchical Trees (SPIHT) coding. Optimized compression results can be achieved by applying GFT due to the proximity in vertex position and motion. We adapt SPIHT to support progressive transmission and design a mechanism to transmit mesh animation sequences with user-defined quality. Experimental results show that our method can obtain a high compression ratio while maintaining a low reconstruction error. Bailin Yang, Luhong Zhang, Frederick W. B. Li, Xiaoheng Jiang, Zhigang Deng 0001, Meng Wang 0001, Mingliang Xu 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2018 | Deep neural networks with Elastic Rectified Linear Units for object recognition
Xiaoheng Jiang, Yanwei Pang, Xuelong Li 0001, Yinghong Xie |
Neurocomputing | 1 |
| 2018 | Cascaded Subpatch Networks for Effective CNNsabstractConventional convolutional neural networks use either a linear or a nonlinear filter to extract features from an image patch (region) of spatial size (typically, is small and is equal to , e.g., is 5 or 7). Generally, the size of the filter is equal to the size of the input patch. We argue that the representational ability of equal-size strategy is not strong enough. To overcome the drawback, we propose to use subpatch filter whose spatial size is smaller than . The proposed subpatch filter consists of two subsequent filters. The first one is a linear filter of spatial size and is aimed at extracting features from spatial domain. The second one is of spatial size and is used for strengthening the connection between different input feature channels and for reducing the number of parameters. The subpatch filter convolves with the input patch and the resulting network is called a subpatch network. Taking the output of one subpatch network as input, we further repeat constructing subpatch networks until the output contains only one neuron in spatial domain. These subpatch networks form a new network called the cascaded subpatch network (CSNet). The feature layer generated by CSNet is called the csconv layer. For the whole input image, we construct a deep neural network by stacking a sequence of csconv layers. Experimental results on five benchmark data sets demonstrate the effectiveness and compactness of the proposed CSNet. For example, our CSNet reaches a test error of 5.68% on the CIFAR10 data set without model averaging. To the best of our knowledge, this is the best result ever obtained on the CIFAR10 data set. Xiaoheng Jiang, Yanwei Pang, Manli Sun, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Convolution in Convolution for Network in NetworkabstractNetwork in network (NiN) is an effective instance and an important extension of deep convolutional neural network consisting of alternating convolutional layers and pooling layers. Instead of using a linear filter for convolution, NiN utilizes shallow multilayer perceptron (MLP), a nonlinear function, to replace the linear filter. Because of the powerfulness of MLP and convolutions in spatial domain, NiN has stronger ability of feature representation and hence results in better recognition performance. However, MLP itself consists of fully connected layers that give rise to a large number of parameters. In this paper, we propose to replace dense shallow MLP with sparse shallow MLP. One or more layers of the sparse shallow MLP are sparely connected in the channel dimension or channel-spatial domain. The proposed method is implemented by applying unshared convolution across the channel dimension and applying shared convolution across the spatial dimension in some computational layers. The proposed method is called convolution in convolution (CiC). The experimental results on the CIFAR10 data set, augmented CIFAR10 data set, and CIFAR100 data set demonstrate the effectiveness of the proposed CiC method. Yanwei Pang, Manli Sun, Xiaoheng Jiang, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Learning Pooling for Convolutional Neural Network
Manli Sun, Zhanjie Song, Xiaoheng Jiang, Yanwei Pang |
Neurocomputing | 3 |
| 2016 | Speed up deep neural network based pedestrian detection by sharing features across multi-scale modelsabstractDeep neural networks (DNNs) have now demonstrated state-of-the-art detection performance on pedestrian datasets. However, because of their high computational complexity, detection efficiency is still a frustrating problem even with the help of Graphics Processing Units (GPUs). To improve detection efficiency, this paper proposes to share features across a group of DNNs that correspond to pedestrian models of different sizes. By sharing features, the computational burden for extracting features from an image pyramid can be significantly reduced. Simultaneously, we can detect pedestrians of several different scales on one single layer of an image pyramid. Furthermore, the improvement of detection efficiency is achieved with negligible loss of detection accuracy. Experimental results demonstrate the robustness and efficiency of the proposed algorithm. Xiaoheng Jiang, Yanwei Pang, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2015 | Flexible sliding windows with adaptive pixel strides
Xiaoheng Jiang, Yanwei Pang, Xuelong Li 0001 |
Signal Process. | 1 |
| 2015 | Efficient object detection by prediction in 3D space
Yanwei Pang, Xiaoheng Jiang, Xuelong Li 0001 |
Signal Process. | 2 |