Caiyan Jia

dblp:67/1275 · DBLP profile ↗
← Back
60ranked-venue papers
3as first author
42since 2021 · last 2026
0000-0003-0650-9564ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 2 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 14 since 2021Databases, data management, data science and information retrieval · 13 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 5 since 2021Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2026 MDiff4STR: Mask Diffusion Model for Scene Text Recognition
abstract
Mask Diffusion Models (MDMs) have recently emerged as a promising alternative to auto-regressive models (ARMs) for vision-language tasks, owing to their flexible balance of efficiency and accuracy. In this paper, for the first time, we introduce MDMs into the Scene Text Recognition (STR) task. We show that vanilla MDM lags behind ARMs in terms of accuracy, although it improves recognition efficiency. To bridge this gap, we propose MDiff4STR, a Mask Diffusion model enhanced with two key improvement strategies tailored for STR. Specifically, we identify two key challenges in applying MDMs to STR: noising gap between training and inference, and overconfident predictions during inference. Both significantly hinder the performance of MDMs. To mitigate the first issue, we develop six noising strategies that better align training with inference behavior. For the second, we propose a token-replacement noise mechanism that provides a non-mask noise type, encouraging the model to reconsider and revise overly confident but incorrect predictions. We conduct extensive evaluations of MDiff4STR on both standard and challenging STR benchmarks, covering diverse scenarios including irregular, artistic, occluded, and Chinese text, as well as whether the use of pretraining. Across these settings, MDiff4STR consistently outperforms popular STR models, surpassing state-of-the-art ARMs in accuracy, while maintaining fast inference with only three denoising steps. Code: https://github.com/Topdu/OpenOCR.
Yongkun Du, Miaomiao Zhao, Songlin Fan, Zhineng Chen, Caiyan Jia, Yu-Gang Jiang 0001
AAAI5
2026 Mamba-integrated spatio-temporal attention graph convolutional network for session-based recommendation
Yafang Li, Baokai Zu, Caiyan Jia
Appl. Intell.4
2026 Contrastive social recommendation: Harnessing community structures for enhanced personalization
Yafang Li, Chenda Li, Baokai Zu, Caiyan Jia
Expert Syst. Appl.4
2026 Unsupervised contrastive domain adaptive rumor detection with test-time classifier adjustment
Hongyan Ran, Xiaohong Li 0012, Huifang Ma, Caiyan Jia, Yaogong Feng
Inf. Process. Manag.5
2026 DGFusion: Dual-Guided Fusion for Robust Multi-Modal 3D Object Detection
abstract
As a critical task in autonomous driving perception systems, 3D object detection is used to identify and track key objects, such as vehicles and pedestrians. However, detecting distant, small, or occluded objects (hard instances) remains a challenge, which directly compromises the safety of autonomous driving systems. We observe that existing multi-modal 3D object detection methods often follow a single-guided paradigm, failing to account for the differences in information density of hard instances between modalities. In this work, we propose DGFusion, based on the Dual-guided paradigm, which fully inherits the advantages of the Point-guide-Image paradigm and integrates the Image-guide-Point paradigm to address the limitations of the single paradigms. The core of DGFusion, the Difficulty-aware Instance Pair Matcher (DIPM), performs instance-level feature matching based on difficulty to generate easy and hard instance pairs, while the Dual-guided Modules exploit the advantages of both pair types to enable effective multi-modal feature fusion. Experimental results demonstrate that our DGFusion outperforms the baseline methods, with respective improvements of +1.0% mAP, +0.8% NDS, and +1.3% average recall on nuScenes. Extensive experiments demonstrate consistent robustness gains for hard instance detection across ego-distance, size, visibility, and small-scale training scenarios.
Feiyang Jia, Caiyan Jia, Ailin Liu, Shaoqing Xu, Qiming Xia, Lei Yang 0060, Ziying Song
IEEE Trans. Circuits Syst. Video Technol.2
2025 Out of Length Text Recognition with Sub-String Matching
abstract
Scene Text Recognition (STR) methods have demonstrated robust performance in word-level text recognition. However, in real applications the text image is sometimes long due to detected with multiple horizontal words. It triggers the requirement to build long text recognition models from readily available short (i.e., word-level) text datasets, which has been less studied previously. In this paper, we term this task Out of Length (OOL) text recognition. We establish the first Long Text Benchmark (LTB) to facilitate the assessment of different methods in long text recognition. Meanwhile, we propose a novel method called OOL Text Recognition with sub-String Matching (SMTR). SMTR comprises two cross-attention-based modules: one encodes a sub-string containing multiple characters into next and previous queries, and the other employs the queries to attend to the image features, matching the sub-string and simultaneously recognizing its next and previous character. SMTR can recognize text of arbitrary length by iterating the process above. To avoid being trapped in recognizing highly similar sub-strings, we introduce a regularization training to compel SMTR to effectively discover subtle differences between similar sub-strings for precise matching. In addition, we propose an inference augmentation strategy to alleviate confusion caused by identical sub-strings in the same text and improve the overall recognition efficiency. Extensive experimental results reveal that SMTR, even when trained exclusively on short text, outperforms existing methods in public short text benchmarks and exhibits a clear advantage on LTB.
Yongkun Du, Zhineng Chen, Caiyan Jia, Xieping Gao 0001, Yu-Gang Jiang 0001
AAAI3
2025 Towards Real-World Rumor Detection: Anomaly Detection Framework with Graph Supervised Contrastive Learning
abstract
Current rumor detection methods based on propagation structure learning predominately treat rumor detection as a class-balanced classification task on limited labeled data. However, real-world social media data exhibits an imbalanced distribution with a minority of rumors among massive regular posts. To address the data scarcity and imbalance issues, we construct two large-scale conversation datasets from Weibo and Twitter and analyze the domain distributions. We find obvious differences between rumor and non-rumor distributions, with non-rumors mostly in entertainment domains while rumors concentrate in news, indicating the conformity of rumor detection to an anomaly detection paradigm. Correspondingly, we propose the Anomaly Detection framework with Graph Supervised Contrastive Learning (AD-GSCL). It heuristically treats unlabeled data as non-rumors and adapts graph contrastive learning for rumor detection. Extensive experiments demonstrate AD-GSCL’s superiority under class-balanced, imbalanced, and few-shot conditions. Our findings provide valuable insights for real-world rumor detection featuring imbalanced data distributions.
Chaoqun Cui, Caiyan Jia
COLING2
2025 Enhancing Rumor Detection Methods with Propagation Structure Infused Language Model
abstract
Pretrained Language Models (PLMs) have excelled in various Natural Language Processing tasks, benefiting from large-scale pretraining and self-attention mechanism’s ability to capture long-range dependencies. However, their performance on social media application tasks like rumor detection remains suboptimal. We attribute this to mismatches between pretraining corpora and social texts, inadequate handling of unique social symbols, and pretraining tasks ill-suited for modeling user engagements implicit in propagation structures. To address these issues, we propose a continue pretraining strategy called Post Engagement Prediction (PEP) to infuse information from propagation structures into PLMs. PEP makes models to predict root, branch, and parent relations between posts, capturing interactions of stance and sentiment crucial for rumor detection. We also curate and release large-scale Twitter corpus: TwitterCorpus (269GB text), and two unlabeled claim conversation datasets with propagation structures (UTwitter and UWeibo). Utilizing these resources and PEP strategy, we train a Twitter-tailored PLM called SoLM. Extensive experiments demonstrate PEP significantly boosts rumor detection performance across universal and social media PLMs, even in few-shot scenarios. On benchmark datasets, PEP enhances baseline models by 1.0-3.7% accuracy, even enabling it to outperform current state-of-the-art methods on multiple datasets. SoLM alone, without high-level modules, also achieves competitive results, highlighting the strategy’s effectiveness in learning discriminative post interaction features.
Chaoqun Cui, Kunkun Ma, Caiyan Jia
COLING4
2025 Don't Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Driving
abstract
End-to-end autonomous driving frameworks enable seamless integration of perception and planning but often rely on one-shot trajectory prediction, which may lead to unstable control and vulnerability to occlusions in single-frame perception. To address this, we propose the Momentum-Aware Driving (MomAD) framework, which introduces trajectory momentum and perception momentum to stabilize and refine trajectory predictions. MomAD comprises two core components: (1) Topological Trajectory Matching (TTM) employs Hausdorff Distance to select the optimal planning query that aligns with prior paths to ensure coherence; (2) Momentum Planning Interactor (MPI) cross-attends the selected planning query with historical queries to expand static and dynamic perception files. This enriched query, in turn, helps regenerate long-horizon trajectory and reduce collision risks. To mitigate noise arising from dynamic environments and detection errors, we introduce robust instance denoising during training, enabling the planning model to focus on critical signals and improve its robustness. We also propose a novel Trajectory Prediction Consistency (TPC) metric to quantitatively assess planning stability. Experiments on the nuScenes dataset demonstrate that MomAD achieves superior long-term consistency (≥ 3s) compared to SOTA methods. Moreover, evaluations on the curated Turning-nuScenes shows that MomAD reduces the collision rate by 26% and improves TPC by 0.97m (33.45%) over a 6s prediction horizon, while closed- loop on Bench2Drive demonstrates an up to 16.3% improvement in success rate. The source code is available at https://github.com/adept-thu/MomAD.
Ziying Song, Caiyan Jia, Hongyu Pan, Shaoqing Xu, Lei Yang 0060, Yadan Luo
CVPR2
2025 SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text Recognition
abstract
Connectionist temporal classification (CTC)-based scene text recognition (STR) methods, e.g., SVTR, are widely employed in OCR applications, mainly due to their simple architecture, which only contains a visual model and a CTC-aligned linear classifier, and therefore fast inference. However, they generally exhibit worse accuracy than encoder-decoder-based methods (EDTRs) due to struggling with text irregularity and linguistic missing. To address these challenges, we propose SVTRv2, a CTC model endowed with the ability to handle text irregularities and model linguistic context. First, a multi-size resizing strategy is proposed to resize text instances to appropriate predefined sizes, effectively avoiding severe text distortion. Meanwhile, we introduce a feature rearrangement module to ensure that visual features accommodate the requirement of CTC, thus alleviating the alignment puzzle. Second, we propose a semantic guidance module. It integrates linguistic context into the visual features, allowing CTC model to leverage language information for accuracy improvement. This module can be omitted at the inference stage and would not increase the time cost. We extensively evaluate SVTRv2 in both standard and recent challenging benchmarks, where SVTRv2 is fairly compared to popular STR models across multiple scenarios, including different types of text irregularity, languages, long text, and whether employing pretraining. SVTRv2 surpasses most EDTRs across the scenarios in terms of accuracy and inference speed. Code: https://github.com/Topdu/OpenOCR.
Yongkun Du, Zhineng Chen, Hongtao Xie 0001, Caiyan Jia, Yu-Gang Jiang 0001
ICCV4
2025 GUIDE: Learnable Deep Contrastive Graph Clustering with Centrality Guidance
Yafang Li, Baokai Zu, Caiyan Jia
ICIC (21)4
2025 Contrastive learning of adaptive social information fusion for recommender systems
Yafang Li, Chenda Li, Caiyan Jia, Baokai Zu
Neurocomputing3
2025 Few-shot learning with distribution calibration for event-level rumor detection
Hongyan Ran, Caiyan Jia, Xiaohong Li 0012, Zhichang Zhang
Neurocomputing2
2025 Attention-based Graph Clustering Network with Dual Information Interaction
Xiumin Lin, Yafang Li, Caiyan Jia, Baokai Zu, Wanting Zhu
Knowl. Based Syst.3
2025 Context Perception Parallel Decoder for Scene Text Recognition
abstract
Scene text recognition (STR) methods have struggled to attain high accuracy and fast inference speed. Auto-Regressive (AR)-based models implement the recognition in a character-by-character manner, showing superiority in accuracy but with slow inference speed. Alternatively, Parallel Decoding (PD)-based models infer all characters in a single decoding pass, offering faster inference speed but generally worse accuracy. To realize the dual goals of "AR-level accuracy and PD-level speed", we propose a Context Perception Parallel Decoder (CPPD) to perceive the related context and predict the character sequence in a PD pass. CPPD devises a character counting module to infer the occurrence count of each character, and a character ordering module to deduce the content-free reading order and positions. Meanwhile, the character prediction task associates the positions with characters. They together build a comprehensive recognition context, which benefits the decoder to focus accurately on characters with the attention mechanism, thereby improving the recognition accuracy. We construct a series of CPPD models and also plug the proposed modules into existing STR decoders. Experiments on both English and Chinese benchmarks demonstrate that the CPPD models achieve highly competitive accuracy while running much faster than existing leading models. Moreover, the plugged models achieve significant accuracy improvements.
Yongkun Du, Zhineng Chen, Caiyan Jia, Xiaoting Yin, Chenxia Li, Yuning Du, Yu-Gang Jiang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Instruction-Guided Scene Text Recognition
abstract
Multi-modal models have shown appealing performance in visual recognition tasks, as free-form text-guided training evokes the ability to understand fine-grained visual content. However, current models cannot be trivially applied to scene text recognition (STR) due to the compositional difference between natural and text images. We propose a novel instruction-guided scene text recognition (IGTR) paradigm that formulates STR as an instruction learning problem and understands text images by predicting character attributes, e.g., character frequency, position, etc. IGTR first devises instruction triplets, providing rich and diverse descriptions of character attributes. To effectively learn these attributes through question-answering, IGTR develops a lightweight instruction encoder, a cross-modal feature fusion module and a multi-task answer head, which guides nuanced text image understanding. Furthermore, IGTR realizes different recognition pipelines simply by using different instructions, enabling a character-understanding-based text reasoning paradigm that differs from current methods considerably. Experiments on English and Chinese benchmarks show that IGTR outperforms existing models by significant margins, while maintaining a small model size and fast inference speed. Moreover, by adjusting the sampling of instructions, IGTR offers an elegant way to tackle the recognition of rarely appearing and morphologically similar characters, which were previous challenges.
Yongkun Du, Zhineng Chen, Caiyan Jia, Yu-Gang Jiang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Propagation Tree Is Not Deep: Adaptive Graph Contrastive Learning Approach for Rumor Detection
abstract
Rumor detection on social media has become increasingly important. Most existing graph-based models presume rumor propagation trees (RPTs) have deep structures and learn sequential stance features along branches. However, through statistical analysis on real-world datasets, we find RPTs exhibit wide structures, with most nodes being shallow 1-level replies. To focus learning on intensive substructures, we propose Rumor Adaptive Graph Contrastive Learning (RAGCL) method with adaptive view augmentation guided by node centralities. We summarize three principles for RPT augmentation: 1) exempt root nodes, 2) retain deep reply nodes, 3) preserve lower-level nodes in deep sections. We employ node dropping, attribute masking and edge dropping with probabilities from centrality-based importance scores to generate views. A graph contrastive objective then learns robust rumor representations. Extensive experiments on four benchmark datasets demonstrate RAGCL outperforms state-of-the-art methods. Our work reveals the wide-structure nature of RPTs and contributes an effective graph contrastive learning approach tailored for rumor detection through principled adaptive augmentation. The proposed principles and augmentation techniques can potentially benefit other applications involving tree-structured graphs.
Chaoqun Cui, Caiyan Jia
AAAI2
2024 GraphBEV: Towards Robust BEV Feature Alignment for Multi-modal 3D Object Detection
Ziying Song, Lei Yang 0060, Shaoqing Xu, Caiyan Jia, Feiyang Jia, Li Wang 0092
ECCV (26)6
2024 Dual Contrastive Learning Guided Pathological Image Re-Staining
abstract
Pathological virtual re-staining is a valuable research topic in AI-aided diagnosis, as it reduces the need for costly and time-consuming physical staining. However, existing methods still suffer from the insufficient ability to preserve tissue microstructure and cellular details, making the generated images less convincing. In this paper, we propose a CycleGAN-based dual contrastive learning re-staining method called DCLRStain. DCLRStain establishes dual contrastive learning between the source and re-stained image domains, conducting negative sampling within each image pair from both domains. It guides the model’s attention to finer content such as cellular details. Meanwhile, DCLRStain introduces a structural similarity-based loss term that further forces the tissue microstructure to be consistent between the source and re-stained images. Experimental results demonstrate that DCLRStain yields competitive quantitative scores compared to state-of-the-art models and maintains superior qualitative performance. Moreover, DCLRStain achieves higher accuracy in the downstream classification task.
Yuexiao Liang, Zhineng Chen, Caiyan Jia, Xiongjun Ye, Xieping Gao 0001
ICASSP4
2024 RoboFusion: Towards Robust Multi-Modal 3D Object Detection via SAM
Ziying Song, Guoxing Zhang, Lei Yang 0060, Shaoqing Xu, Caiyan Jia, Feiyang Jia, Li Wang 0092
IJCAI6
2024 Adversarially deep interative-fused embedding clustering via joint self-supervised networks
Yafang Li, Xiumin Lin, Caiyan Jia, Baokai Zu, Shaotao Zhu
Neurocomputing3
2024 A debiased self-training framework with graph self-supervised pre-training aided for semi-supervised rumor detection
Yuhan Qiao, Chaoqun Cui, Yiying Wang, Caiyan Jia
Neurocomputing4
2024 Attributed graph clustering under the contrastive mechanism with cluster-preserving augmentation
Yimei Zheng, Caiyan Jia, Jian Yu 0001
Inf. Sci.2
2024 GraphAlign++: An Accurate Feature Alignment by Graph Matching for Multi-Modal 3D Object Detection
abstract
LiDAR and camera are complementary sensors for 3D object detection in autonomous driving. However, it is challenging to explore the unnatural interaction between point clouds and images, and the critical factor is how to conduct feature alignment of these heterogeneous modalities. Currently, many methods achieve feature alignment through projection calibration, without accounting for the impact of sensors misalignment errors, resulting in sub-optimal performance. In this paper, we present GraphAlign++, a more accurate feature alignment framework for 3D object detection by graph matching. Specifically, we construct the nearest neighbor relationship by calculating Euclidean distances of point cloud features within the subspaces. Through the projection calibration between the image and point cloud pairs, we project the nearest neighbors of point cloud features onto the corresponding image. Then by matching the nearest neighbors of a single point-feature of the point cloud with multiple pixel-features of the image, we search for a more appropriate feature alignment. In addition, we provide a self-attention module to enhance the weights of significant relations to fine-tune the feature alignment between these two heterogeneous modalities. Extensive experiments on nuScenes benchmark demonstrate the effectiveness and efficiency of GraphAlign++. Notably, due to the more accurate feature alignment, which contributes to increase mAP by 3.10% on KITTI test hard level, our method is remarkably beneficial for long-range object detection.
Ziying Song, Caiyan Jia, Lei Yang 0060, Haiyue Wei
IEEE Trans. Circuits Syst. Video Technol.2
2024 SparseDet: A Simple and Effective Framework for Fully Sparse LiDAR-Based 3-D Object Detection
abstract
LiDAR-based sparse 3-D object detection plays a crucial role in autonomous driving applications due to its computational efficiency advantages. Existing methods either use the features of a single central voxel as an object proxy or treat an aggregated cluster of foreground points as an object proxy. However, the former cannot aggregate contextual information, resulting in insufficient information expression in object proxies. The latter relies on multistage pipelines and auxiliary tasks, which reduce the inference speed. To maintain the efficiency of the sparse framework while fully aggregating contextual information, in this work, we propose SparseDet that designs sparse queries as object proxies. It introduces two key modules: the local multiscale feature aggregation (LMFA) module and the global feature aggregation (GFA) module, aiming to fully capture the contextual information, thereby enhancing the ability of the proxies to represent objects. The LMFA module achieves feature fusion across different scales for sparse key voxels via coordinate transformations and using nearest neighbor relationships to capture object-level details and local contextual information, whereas the GFA module uses self-attention mechanisms to selectively aggregate the features of the key voxels across the entire scene for capturing scene-level contextual information. Experiments on nuScenes and KITTI demonstrate the effectiveness of our method. Specifically, SparseDet surpasses the previous best sparse detector VoxelNeXt (a typical method using voxels as object proxies) by 2.2% mean average precision (mAP) with 13.5 frames/s on nuScenes and outperforms VoxelNeXt by 1.12%$\text {AP}_{\text {3-D}}$on hard level tasks with 17.9 frames/s on KITTI. What is more, not only the mAP of SparseDet exceeds that of FSDV2 (a classical method using clusters of foreground points as object proxies) but also its inference speed is 1.3 times faster than FSDV2 on the nuScenes test set. The code has been released inhttps://github.com/liulin813/SparseDet.git.
Ziying Song, Qiming Xia, Feiyang Jia, Caiyan Jia, Lei Yang 0060, Hongyu Pan
IEEE Trans. Geosci. Remote. Sens.5
2024 Robustness-Aware 3D Object Detection in Autonomous Driving: A Review and Outlook
abstract
In the realm of modern autonomous driving, the perception system is indispensable for accurately assessing the state of the surrounding environment, thereby enabling informed prediction and planning. The key step to this system is related to 3D object detection that utilizes vehicle-mounted sensors such as LiDAR and cameras to identify the size, the category, and the location of nearby objects. Despite the surge in 3D object detection methods aimed at enhancing detection precision and efficiency, there is a gap in the literature that systematically examines their resilience against environmental variations, noise, and weather changes. This study emphasizes the importance of robustness, alongside accuracy and latency, in evaluating perception systems under practical scenarios. Our work presents an extensive survey of camera-only, LiDAR-only, and multi-modal 3D object detection algorithms, thoroughly evaluating their trade-off between accuracy, latency, and robustness, particularly on datasets like KITTI-C and nuScenes-C to ensure fair comparisons. Among these, multi-modal 3D detection approaches exhibit superior robustness, and a novel taxonomy is introduced to reorganize the literature for enhanced clarity. This survey aims to offer a more practical perspective on the current capabilities and the constraints of 3D object detection algorithms in real-world applications, thus steering future research towards robustness-centric advancements.
Ziying Song, Feiyang Jia, Yadan Luo, Caiyan Jia, Lei Yang 0060, Li Wang 0092
IEEE Trans. Intell. Transp. Syst.5
2024 ProtoMGAE: Prototype-Aware Masked Graph Auto-Encoder for Graph Representation Learning
abstract
Graph self-supervised representation learning has gained considerable attention and demonstrated remarkable efficacy in extracting meaningful representations from graphs, particularly in the absence of labeled data. Two representative methods in this domain are graph auto-encoding and graph contrastive learning. However, the former methods primarily focus on global structures, potentially overlooking some fine-grained information during reconstruction. The latter methods emphasize node similarity across correlated views in the embedding space, potentially neglecting the inherent global graph information in the original input space. Moreover, handling incomplete graphs in real-world scenarios, where original features are unavailable for certain nodes, poses challenges for both types of methods. To alleviate these limitations, we integrate masked graph auto-encoding and prototype-aware graph contrastive learning into a unified model to learn node representations in graphs. In our method, we begin by masking a portion of node features and utilize a specific decoding strategy to reconstruct the masked information. This process facilitates the recovery of graphs from a global or macro level and enables handling incomplete graphs easily. Moreover, we treat the masked graph and the original one as a pair of contrasting views, enforcing the alignment and uniformity between their corresponding node representations at a local or micro level. Last, to capture cluster structures from a meso level and learn more discriminative representations, we introduce a prototype-aware clustering consistency loss that is jointly optimized with the preceding two complementary objectives. Extensive experiments conducted on several datasets demonstrate that the proposed method achieves significantly better or competitive performance on downstream tasks, especially for graph clustering, compared with the state-of-the-art methods, showcasing its superiority in enhancing graph representation learning.
Yimei Zheng, Caiyan Jia
ACM Trans. Knowl. Discov. Data2
2023 Unsupervised Cross-Domain Rumor Detection with Contrastive Learning and Cross-Attention
abstract
Massive rumors usually appear along with breaking news or trending topics, seriously hindering the truth. Existing rumor detection methods are mostly focused on the same domain, thus have poor performance in cross-domain scenarios due to domain shift. In this work, we propose an end-to-end instance-wise and prototype-wise contrastive learning model with cross-attention mechanism for cross-domain rumor detection. The model not only performs cross-domain feature alignment, but also enforces target samples to align with the corresponding prototypes of a given source domain. Since target labels in a target domain are unavailable, we use a clustering-based approach with carefully initialized centers by a batch of source domain samples to produce pseudo labels. Moreover, we use a cross-attention mechanism on a pair of source data and target data with the same labels to learn domain-invariant representations. Because the samples in a domain pair tend to express similar semantic patterns especially on the people’s attitudes (e.g., supporting or denying) towards the same category of rumors, the discrepancy between a pair of source domain and target domain will be decreased. We conduct experiments on four groups of cross-domain datasets and show that our proposed model achieves state-of-the-art performance.
Hongyan Ran, Caiyan Jia
AAAI2
2023 GraphAlign: Enhancing Accurate Feature Alignment by Graph matching for Multi-Modal 3D Object Detection
abstract
LiDAR and cameras are complementary sensors for 3D object detection in autonomous driving. However, it is challenging to explore the unnatural interaction between point clouds and images, and the critical factor is how to conduct feature alignment of heterogeneous modalities. Currently, many methods achieve feature alignment by projection calibration only, without considering the problem of coordinate conversion accuracy errors between sensors, leading to sub-optimal performance. In this paper, we present GraphAlign, a more accurate feature alignment strategy for 3D object detection by graph matching. Specifically, we fuse image features from a semantic segmentation encoder in the image branch and point cloud features from a 3D Sparse CNN in the LiDAR branch. To save computation, we construct the nearest neighbor relationship by calculating Euclidean distance within the subspaces that are divided into the point cloud features. Through the projection calibration between the image and point cloud, we project the nearest neighbors of point cloud features onto the image features. Then by matching the nearest neighbors with a single point cloud to multiple images, we search for a more appropriate feature alignment. In addition, we provide a self-attention module to enhance the weights of significant relations to fine-tune the feature alignment between heterogeneous modalities. Extensive experiments on nuScenes benchmark demonstrate the effectiveness and efficiency of our GraphAlign.
Ziying Song, Haiyue Wei, Lei Yang 0060, Caiyan Jia
ICCV5
2023 Contrastive Learning with Cluster-Preserving Augmentation for Attributed Graph Clustering
Yimei Zheng, Caiyan Jia, Jian Yu 0001
ECML/PKDD (1)2
2023 Self-supervised variational autoencoder towards recommendation by nested contrastive learning
Jing Wang 0116, Jun Wu 0007, Caiyan Jia
Appl. Intell.3
2023 A metric-learning method for few-shot cross-event rumor detection
Hongyan Ran, Caiyan Jia, Jian Yu 0001
Neurocomputing2
2023 A Rumor Detection Model Incorporating Propagation Path Contextual Semantics and User Information
Xueming Han, Caiyan Jia
Neural Process. Lett.3
2023 Deep embedded clustering with distribution consistency preservation for attributed networks
Yimei Zheng, Caiyan Jia, Jian Yu 0001, Xuanya Li
Pattern Recognit.2
2023 VP-Net: Voxels as Points for 3-D Object Detection
abstract
3D object detection with LiDAR point clouds is a challenging problem which requires 3D scene understanding, yet this task is critical to autonomous driving. Existing voxel-based 3D object detectors are becoming increasingly popular but have several shortcomings. For example, during voxelization, features of distant sparse point clouds are largely discarded, which leads to the missing detection of objects. Additionally, the correlation of points between voxels and the importance of different voxels within a region are not well learned. Therefore, we present a robust network (VP-Net) that views voxels as points to accurately detect 3D objects in LiDAR point clouds and can capture objects’ internal relationships. 3D CNN processing shows the output features of VP-Net as key points. The relationship between key points is then constructed into local graphs to enhance object feature extraction via a self-attention mechanism. Finally, the Euclidean distance between the extracted features guides our model’s weight reassignment for strengthening the importance of neighbor points, thereby enhancing the internal feature aggregation of objects. Experiments on KITTI and nuScenes 3D object detection benchmarks demonstrate the efficiency of enhancing inter-voxel validity within object features and show that the proposed VP-Net can achieve state-of-the-art performance.
Ziying Song, Haiyue Wei, Caiyan Jia, Yongchao Xia, Xiaokun Li
IEEE Trans. Geosci. Remote. Sens.3
2023 VoxelNextFusion: A Simple, Unified, and Effective Voxel Fusion Framework for Multimodal 3-D Object Detection
abstract
LiDAR-camera fusion can enhance the performance of 3D object detection by utilizing complementary information between depth-aware LiDAR points and semantically rich images. Existing voxel-based methods face significant challenges when fusing sparse voxel features with dense image features in a one-to-one manner, resulting in the loss of the advantages of images, including semantic and continuity information, leading to sub-optimal detection performance, especially at long distances. In this paper, we present VoxelNextFusion, a multi-modal 3D object detection framework specifically designed for voxel-based methods, which effectively bridges the gap between sparse point clouds and dense images. In particular, we propose a voxel-based image pipeline that involves projecting point clouds onto images to obtain both pixel- and patch-level features. These features are then fused using a self-attention to obtain a combined representation. Moreover, to address the issue of background features present in patches, we propose a feature importance module that effectively distinguishes between foreground and background features, thus minimizing the impact of the background features. Extensive experiments were conducted on the widely used KITTI and nuScenes 3D object detection benchmarks. Notably, our VoxelNextFusion achieved around +3.20% in [email protected] improvement for car detection in hard level compared to the Voxel R-CNN baseline on the KITTI test dataset.
Ziying Song, Jun Xie 0003, Caiyan Jia, Shaoqing Xu, Zhepeng Wang 0002
IEEE Trans. Geosci. Remote. Sens.5
2022 SVTR: Scene Text Recognition with a Single Visual Model
abstract
Dominant scene text recognition models commonly contain two building blocks, a visual model for feature extraction and a sequence model for text transcription. This hybrid architecture, although accurate, is complex and less efficient. In this study, we propose a Single Visual model for Scene Text recognition within the patch-wise image tokenization framework, which dispenses with the sequential modeling entirely. The method, termed SVTR, firstly decomposes an image text into small patches named character components. Afterward, hierarchical stages are recurrently carried out by component-level mixing, merging and/or combining. Global and local mixing blocks are devised to perceive the inter-character and intra-character patterns, leading to a multi-grained character component perception. Thus, characters are recognized by a simple linear prediction. Experimental results on both English and Chinese scene text recognition tasks demonstrate the effectiveness of SVTR. SVTR-L (Large) achieves highly competitive accuracy in English and outperforms existing methods by a large margin in Chinese, while running faster. In addition, SVTR-T (Tiny) is an effective and much smaller model, which shows appealing speed at inference. The code is publicly available at https://github.com/PaddlePaddle/PaddleOCR.
Yongkun Du, Zhineng Chen, Caiyan Jia, Xiaoting Yin, Tianlun Zheng, Chenxia Li, Yuning Du, Yu-Gang Jiang 0001
IJCAI3
2022 Fast Detection of Multi-Direction Remote Sensing Ship Object Based on Scale Space Pyramid
abstract
Ships in remote sensing images are usually arranged in arbitrary direction, small in size, and densely arranged. As a result, existing object detection algorithms cannot detect ships quickly and accurately. In order to solve the above problems, a lightweight object detection network for fast detection of ships is proposed. The network is composed of backbone network, four-scale fusion network and rotation branch. First, a lightweight network unit S-LeanNet is designed and used to build a low-computing and accurate backbone network. Then, a four-scale feature fusion module is designed to generate a four-scale feature pyramid, which contains more features such as ship shape and texture, and at the same time is conducive to the detection of small ships. Finally, a novel rotation branch module is designed, using balance L1 loss function and R-NMS for post-processing, to realize the precise positioning and regression of the rotating bounding box in one step. Experimental results show that the detection precision of our method in the DOT A remote sensing data set is compared with the latest SCRDet detection method, the precision is increased by 1.1%, and the operating speed is increased by 8 times, which can meet the fast detection requirements of ships.
Ziying Song, Li Wang 0092, Caiyan Jia, Jiangfeng Bi, Haiyue Wei, Yongchao Xia, Lijun Zhao 0003
MSN4
2022 MGAT-ESM: Multi-channel graph attention neural network with event-sharing module for rumor detection
Hongyan Ran, Caiyan Jia, Xuanya Li
Inf. Sci.2
2021 Self-supervised Variational Autoencoder for Recommender Systems
abstract
Variational autoencoder (VAE) is considered as an emerging model for ensuring competitive performance in recom-mender systems. However, its performance is severely limited by the amount of training examples and, as a result, existing VAE models may fail to provide satisfactory recommendation results in presence of highly sparse user-item interactions. In this paper, we propose a self-supervised VAE model, SSVAE in short, to improve the generalization ability of VAE model on the sparse interaction datasets. Concretely, we first build multiple views for each user by data augmentation, and then design a pretext task to align the representations learned from different views of each user. Particularly, SSVAE aims to optimize a combined objective of recommendation task and pretext task, making them to rein-force each other during the learning process. Our encouraging experimental results on three real-world benchmarks validate the superiority of our SSVAE model to state-of-the-art VAE style recommendation techniques.
Jing Wang 0116, Gangdu Liu, Jun Wu 0007, Caiyan Jia
ICTAI4
2021 Bag of Tricks for Building an Accurate and Slim Object Detector for Embedded Applications
abstract
Object detection is an essential computer vision task that possesses extensive application prospects in on-road applications. Copious novel methods have been proposed in this branch recently. However, the majority of them have high computational cost, making them intractable to be deployed on embedded devices. In this paper, taking YOLOv5s, the smallest model in the YOLOv5 family, as the baseline, we explore a bag of tricks that improve the detection performance for a specified on-road application, under the premise of ensuring that it does not increase the computational cost of YOLOv5s. Specifically, we introduce relevantly external data to deal with the problems of sample imbalance. Meanwhile, knowledge distillation is employed to transfer knowledge from a cumbersome model to a compact model, where a united distillation scheme is developed to enhance the effectiveness. In addition, a pseudo-label based training strategy is utilized to further learn from the biggest YOLOv5 model. We have applied the above tricks to the Embedded Deep Learning Object Detection Model Compression Competition for Traffic in Asian Countries held in conjunction with ICMR 2021. The experiments have shown that all the tricks are useful. Their combination have built an accurate and slim detection model. It is highly competitive and has been ranked 2nd place in the competition. We believe the tricks are also meaningful for building other application-oriented object detectors.
Yongkun Du, Zhineng Chen, Caiyan Jia, Xuanya Li, Yu-Gang Jiang 0001
ICMR3
2021 A lightweight propagation path aggregating network with neural topic model for rumor detection
Hongyan Ran, Caiyan Jia, Xuanya Li, Xueming Han
Neurocomputing3
2020 Signet Ring Cell Detection with Classification Reinforcement Detection Network
Caiyan Jia, Zhineng Chen, Xieping Gao 0001
ISBRA2
2020 Traffic Flow Prediction via Spatial Temporal Graph Neural Network
abstract
Traffic flow analysis, prediction and management are keystones for building smart cities in the new era. With the help of deep neural networks and big traffic data, we can better understand the latent patterns hidden in the complex transportation networks. The dynamic of the traffic flow on one road not only depends on the sequential patterns in the temporal dimension but also relies on other roads in the spatial dimension. Although there are existing works on predicting the future traffic flow, the majority of them have certain limitations on modeling spatial and temporal dependencies. In this paper, we propose a novel spatial temporal graph neural network for traffic flow prediction, which can comprehensively capture spatial and temporal patterns. In particular, the framework offers a learnable positional attention mechanism to effectively aggregate information from adjacent roads. Meanwhile, it provides a sequential component to model the traffic flow dynamics which can exploit both local and global temporal dependencies. Experimental results on various real traffic datasets demonstrate the effectiveness of the proposed framework.
Yao Ma 0001, Yiqi Wang 0001, Wei Jin 0009, Xin Wang 0035, Jiliang Tang, Caiyan Jia, Jian Yu 0001
WWW7
2020 Integrated network analysis of symptom clusters across disease conditions
Kezhi Lu, Kuo Yang 0001, Edouard Niyongabo, Zixin Shu, Kai Chang, Qunsheng Zou, Jiyue Jiang, Caiyan Jia, Baoyan Liu, Xuezhong Zhou
J. Biomed. Informatics9
2019 psDirector: An Automatic Director for Watching View Generation from Panoramic Soccer Video
Caiyan Jia, Zhineng Chen, Xiaoyan Gu 0001, Hongyun Bao
MMM (2)2
2019 A generative model for exploring structure regularities in attributed networks
Zhenhai Chang, Caiyan Jia, Xianjun Yin, Yimei Zheng
Inf. Sci.2
2019 Locally Weighted Fusion of Structural and Attribute Information in Graph Clustering
abstract
Attributed graphs have attracted much attention in recent years. Different from conventional graphs, attributed graphs involve two different types of heterogeneous information, i.e., structural information, which represents the links between the nodes, and attribute information on each of the nodes. Clustering on attributed graphs usually requires the fusion of both types of information in order to identify meaningful clusters. However, most of existing works implement the combination of these two types of information in a "global" manner by treating all nodes equally and learning a global weight for the information fusion. To address this issue, this paper proposed a novel weighted K -means algorithm with "local" learning for attributed graph clustering, called adaptive fusion of structural and attribute information (Adapt-SA) and analyzed the convergence property of the algorithm. The key advantage of this model is to automatically balance the structural connections and attribute information of each node to learn a fusion weight, and get densely connected clusters with high attribute semantic similarity. Experimental study of weights on both synthetic and real-world data sets showed that the weights learned by Adapt-SA were reasonable, and they reflected which one of these two types of information was more important to decide the membership of a node. We also compared Adapt-SA with the state-of-the-art algorithms on the real-world networks with varieties of characteristics. The experimental results demonstrated that our method outperformed the other algorithms in partitioning an attributed graph into a community structure or other general structures.
Yafang Li, Caiyan Jia, Xiangnan Kong, Liu Yang 0010, Jian Yu 0001
IEEE Trans. Cybern.2
2019 Structure-Aware Deep Learning for Product Image Classification
abstract
Automatic product image classification is a task of crucial importance with respect to the management of online retailers. Motivated by recent advancements of deep Convolutional Neural Networks (CNN) on image classification, in this work we revisit the problem in the context of product images with the existence of a predefined categorical hierarchy and attributes, aiming to leverage the hierarchy and attributes to improve classification accuracy. With these structure-aware clues, we argue that more advanced deep models could be developed beyond the flat one-versus-all classification performed by conventional CNNs. To this end, novel efforts of this work include a salient-sensitive CNN that gazes into the product foreground by inserting a dedicated spatial attention module; a multiclass regression-based refinement that is expected to predict more accurately by merging prediction scores from multiple preceding CNNs, each corresponding to a distinct classifier in the hierarchy; and a multitask deep learning architecture that effectively explores correlations among categories and attributes for categorical label prediction. Experimental results on nearly 1 million real-world product images basically validate the effectiveness of the proposed efforts individually and jointly, from which performance gains are observed.
Zhineng Chen, Shanshan Ai, Caiyan Jia
ACM Trans. Multim. Comput. Commun. Appl.3
2018 Clustering Uncertain Graphs with Node Attributes
abstract
Graph clustering has attracted much attention in recent years, which has wide applications in social and biological networks. Recent approaches on graph clustering mainly focus on either certain graphs with node attributes or uncertain graphs without node attributes. However, many real-world graphs have both uncertainty on the edges and attributes on the nodes. We refer to such networks as \emph{attributed uncertain graphs}. Different from conventional graphs, attributed uncertain graphs post two major challenges for graph clustering: 1) uncertainty on the edges, which makes it difficult to extract reliable clusters; 2) high dimensional attributes on the nodes, which contain irrelevant and noisy information. In this paper, we study the problem of node clustering on attributed uncertain graphs, where we exploit both the uncertain edges and a set of important attributes for graph clustering. The uncertain edges can help identify the set of relevant attributes in the nodes, which are called focus attributes. While the focus attributes can help reduce the uncertainty in edges for graph clustering. We propose two novel approaches: AUG-I based upon integrated attribute induced graphs and AUG-U based upon the unified partition over possible worlds of a uncertain graph. Extensive empirical studies on real-world datasets demonstrate the effectiveness of our approaches for clustering tasks on attributed uncertain graphs.
Yafang Li, Xiangnan Kong, Caiyan Jia, Jianqiang Li 0002
ACML3
2018 General Structure Preserving Network Embedding
Sinan Zhu, Caiyan Jia
IDEAL (1)2
2018 FFDet: a Fully Convolutional Network for Coral Reef Fish Detection by Layer Fusion
abstract
Underwater coral reef fish detection is topic receiving increasingly attention due to its importance in various applications like fish biodiversity monitoring, marine resource managements, etc. However, compared with studies on generic object detection, existing methods on this task are not mature so far where advanced deep models and technologies are seldom considered. This paper presents FFDet, a fully convolutional network for coral reef fish detection by layer fusion. FFDet consists of a single shot multibox detector (SSD)-based backbone, but different with SSD, it devises a novel feature fusion module to aggregate adjacent prediction layers for enhanced feature representation. Thus, instead of using the prediction layers one-by-one, the enhanced features each combining information from multiple layers, are leveraged to detect fishes at different scales. We argue that the proposed module is capable of encoding both strong semantics and detail context information. Experimental results on SeaCLEF dataset show that FFDet not only outperforms SSD in performance by sacrificing only a little efficiency, but also better than another two popular end-to-end deep models in both detection performance and speed, especially on detecting large-sized fishes.
Cuncun Shi, Caiyan Jia, Zhineng Chen
VCIP2
2018 Concept decompositions for short text clustering by identifying word communities
Caiyan Jia, Matthew B. Carson, Jian Yu 0001
Pattern Recognit.1
2017 Large-Scale Product Classification via Spatial Attention Based CNN Learning and Multi-class Regression
Shanshan Ai, Caiyan Jia, Zhineng Chen
MMM (1)2
2015 Improving Automatic Name-Face Association using Celebrity Images on the Web
abstract
This paper investigates the task of automatically associating faces appearing in images (or videos) with their names. Our novelty lies in the use of celebrity Web images to facilitate the task. Specifically, we first propose a method named Image Matching (IM), which uses the faces in images returned from name queries over an image search engine as the gallery set of the names, and a probe face is classified as one of the names, or none of them, according to their matching scores and compatibility characterized by a proposed Assigning-Thresholding (AT) pipeline. Noting IM could provide guidance for association for the well-established Graph-based Association (GA), we further propose two methods that jointly utilize the two kinds of complementary cues. They are: the early fusion of IM and GA (EF-IMGA) that takes the IM score as an additional information source to help the association in GA, and the late fusion of IM and GA (LF-IMGA) that combines the scores from both IM and GA obtained individually to make the association. Evaluations on datasets of captioned news images and Web videos both show the proposed methods, especially the two fused ones, provide significant improvements over GA.
Zhineng Chen, Bailan Feng, Chong-Wah Ngo, Caiyan Jia, Xiangsheng Huang
ICMR4
2013 A fast weak motif-finding algorithm based on community detection in graphs
abstract
BACKGROUND: Identification of transcription factor binding sites (also called 'motif discovery') in DNA sequences is a basic step in understanding genetic regulation. Although many successful programs have been developed, the problem is far from being solved on account of diversity in gene expression/regulation and the low specificity of binding sites. State-of-the-art algorithms have their own constraints (e.g., high time or space complexity for finding long motifs, low precision in identification of weak motifs, or the OOPS constraint: one occurrence of the motif instance per sequence) which limit their scope of application. RESULTS: In this paper, we present a novel and fast algorithm we call TFBSGroup. It is based on community detection from a graph and is used to discover long and weak (l,d) motifs under the ZOMOPS constraint (zero, one or multiple occurrence(s) of the motif instance(s) per sequence), where l is the length of a motif and d is the maximum number of mutations between a motif instance and the motif itself. Firstly, TFBSGroup transforms the (l, d) motif search in sequences to focus on the discovery of dense subgraphs within a graph. It identifies these subgraphs using a fast community detection method for obtaining coarse-grained candidate motifs. Next, it greedily refines these candidate motifs towards the true motif within their own communities. Empirical studies on synthetic (l, d) samples have shown that TFBSGroup is very efficient (e.g., it can find true (18, 6), (24, 8) motifs within 30 seconds). More importantly, the algorithm has succeeded in rapidly identifying motifs in a large data set of prokaryotic promoters generated from the Escherichia coli database RegulonDB. The algorithm has also accurately identified motifs in ChIP-seq data sets for 12 mouse transcription factors involved in ES cell pluripotency and self-renewal. CONCLUSIONS: Our novel heuristic algorithm, TFBSGroup, is able to quickly identify nearly exact matches for long and weak (l, d) motifs in DNA sequences under the ZOMOPS constraint. It is also capable of finding motifs in real applications. The source code for TFBSGroup can be obtained from http://bioinformatics.bioengr.uic.edu/TFBSGroup/.
Caiyan Jia, Matthew B. Carson, Jian Yu 0001
BMC Bioinform.1
2011 Data Clustering by Scaled Adjacency Matrix
Jian Yu 0001, Caiyan Jia
KSEM2
2010 Affinity Propagation on Identifying Communities in Social and Biological Networks
Caiyan Jia, Yawen Jiang, Jian Yu 0001
KSEM1
2009 Convergence Analysis of Affinity Propagation
Jian Yu 0001, Caiyan Jia
KSEM2
2007 An Exact Data Mining Method for Finding Center Strings and All Their Instances
abstract
Common substring problems allowing errors are known to be NP-hard. The main challenge of the problems lies in the combinatorial explosion of potential candidates. In this paper, we propose and study a generalized center string (GCS) problem, where not only all models (center strings) of any length, but also the positions of all their (degenerative) instances in input sequences are searched for. Inspired by frequent pattern mining techniques in data mining field, we present an exact and efficient method to solve GCS. First, a highly parallelized Trie-like structure, consensus tree, is proposed. Based on this structure, we present three Bpriori algorithms step by step. Bpriori algorithms can solve GCS with reasonable time and/or space complexities. We have proved that GCS is fixed parameter tractable with respect to fixed symbol set size and fixed length of input sequences. Experiment results on both artificial and real data have shown the correctness of the algorithms and the validity of our complexity analysis. A comparison with some current algorithms for solving common approximate substring problems is also given
Ruqian Lu, Caiyan Jia, Shaofang Zhang, Lusheng Chen
IEEE Trans. Knowl. Data Eng.2