VLDB 2026 Research / reviewers in the wild / expert
Tong Jia 0001
dblp:56/7768-1
· DBLP profile ↗
54ranked-venue papers
2as first author
49since 2021 · last 2026
0000-0003-1424-798XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 1 first-author · 28 since 2021Artificial intelligence and machine learning · 19 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BASE: A boundary-aware adaptive semantic evidential framework for open-set skeleton-based action recognition
Dongyue Chen 0001, Dingyu Xue, Shizhuo Deng, Tong Jia 0001 |
Expert Syst. Appl. | 5 |
| 2026 | Conditional diffusion models for X-ray security image synthesis: Addressing data scarcity in security inspection tasks
Da Cai, Tong Jia 0001, Hao Wang 0073, Mingyuan Li 0003, Dongyue Chen 0001 |
Neurocomputing | 2 |
| 2026 | WeCo-OSAR: Weighted Contrastive Learning for Open-Set Skeleton-based Action Recognition with pseudo-OOD samples
Dongyue Chen 0001, Dingyu Xue, Shizhuo Deng, Tong Jia 0001 |
Knowl. Based Syst. | 5 |
| 2026 | DCART: A dual contrastive alignment residual transformer model for visual grounding
Dongyue Chen 0001, Hao Wang 0073, Tong Jia 0001, Shizhuo Deng |
Pattern Recognit. | 4 |
| 2026 | CGFMamba-PCR: Color-Geometric Fusion Mamba-Based Color Point Cloud RegistrationabstractRecently, color point cloud registration has begun to receive attention. Unlike geometric-only point clouds, color point clouds incorporate additional color information. Therefore, color point cloud registration can achieve higher accuracy than geometric-only point cloud registration. Despite the success, existing methods are computationally intensive due to the high resource demands of Transformer. In this paper, we propose a Mamba architecture based registration algorithm, CGFMamba-PCR, with color-geometric fusion. Specifically, we propose CGFMamba, a novel color-geometric fusion Mamba network for color point cloud registration, which enhances feature representation through color-geometric guided point ordering and positional encoding. For the input color point cloud pair, they are passed through CGFMamba based feature extraction module to obtain their corresponding features. Then, these features are passed through feature matching and outlier rejection modules to obtain final registration result. Furthermore, an ordering method for the Mamba architecture is proposed that clusters color hue and sorts spatial coordinates. Experiments on Color3DMatch and Color3DLoMatch datasets demonstrate that the proposed algorithm outperforms the state-of-the-art (SOTA) methods. The code of the proposed algorithm will be open-sourced upon acceptance of this paper. Shiyi Guo, Tong Jia 0001, Bi Yang, Yihong Wu 0002, Hao Wei 0008, Nannan Liu, Ning An 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | GAA-TSO: Geometry-Aware-Assisted Depth Completion for Transparent and Specular ObjectsabstractTransparent and specular objects are frequently encountered in daily life, factories, and laboratories. However, due to the unique optical properties, the depth information on these objects is usually incomplete and inaccurate, which poses significant challenges for downstream robotics tasks. Therefore, it is crucial to accurately restore the depth information of transparent and specular objects. Previous depth completion methods for these objects usually generate structure-less or ambiguous depth predictions. To address these issues, we propose a geometry-aware assisted depth completion method for transparent and specular objects, which focuses on exploring the 3D structural cues of the scene. Specifically, besides extracting 2D features from RGB-D input, we back-project the input depth to a point cloud and build the 3D branch to extract hierarchical scene-level 3D structural features. To exploit 3D geometric information, we design several gated cross-modal fusion modules to effectively propagate multi-level 3D geometric features to the image branch. In addition, we propose an adaptive correlation aggregation strategy to appropriately assign 3D features to the corresponding 2D features. Extensive experiments on ClearGrasp, OOD, TransCG, and STD datasets show that our method outperforms other state-of-the-art methods. We further demonstrate that our method significantly enhances the performance of downstream robotic grasping tasks. The code will be available at: https://github.com/lyz3356/GAA-TSO. Yizhe Liu, Tong Jia 0001, Jiahui Wei, Hao Wang 0073, Dongyue Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Motion-Guided Disentanglement for Point Cloud Masked AutoencodersabstractMasked autoencoders have been extended beyond images, but random masking often fails to capture non-uniform motion regions in inherently disordered and irregular data like point cloud videos. In this paper, we propose a motion-guided disentanglement method (MGD) to improve masked autoencoders for point cloud video representation learning. Specifically, we begin by estimating motion intensity using an optimal transport approach, which guides the separate masking of dynamic and static regions. This motion-guided masking ensures balanced coverage, addressing the limitations of random masking in capturing non-uniformly distributed motion regions. Furthermore, we disentangle the prediction tasks into motion prediction for high-motion point tubes and appearance reconstruction for low-motion ones. This disentanglement enables the model to more effectively capture both motion and appearance in point cloud videos. We conducted experiments on four widely used point cloud video datasets—NTU RGB+D, MSR-Action3D, NvGesture, and SHREC’17—which demonstrate that our approach consistently improves masked autoencoders for point cloud video representation learning, achieving new state-of-the-art results. Code will be publicly available on GitHub. Haoran Wang 0001, Shaqing Song, Baosheng Yu, Tong Jia 0001, Dongyue Chen 0001, Chunfeng Yuan, Weiming Hu 0004, Haibin Ling |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Contextual Style Coherence Network for X-Ray Prohibited Item Image SynthesisabstractProhibited item detection in X-Ray baggage images plays a crucial role for preventing the social security and stability. Well annotated X-Ray prohibited item training samples show necessity in achieving high detection performance for X-Ray inspection system. While collection of massive samples is extremely laborious and costly, especially for those X-Ray images, which need professional inspection machine. Synthesizing X-Ray images through Threat Image Projection (TIP) is a promising solution to overcome the data insufficient limitation in prohibited item detection. However, TIP based methods rarely consider the contextual style coherence between the foreground prohibited items and background images, resulting in generating low realistic X-Ray security images. For improving image quality and diversity, we propose a Contextual Style Coherence Network for X-Ray Prohibited item Image Synthesis. Specifically, we first propose a style fusion module to guarantee the style coherence and consistency between the foreground prohibited items and background images. We transfer the threat image projection from image space to feature space, and an affine transformation matrix is applied to uniformly sample the location, ratio and scale of the prohibited items to improve the sample diversity. We further normalize the features of the foreground prohibited item by implementing the style transfer through Gram matrix. Then, a mask partial convolution is designed for inpainting the non-object regions of the foreground prohibited items to achieve a better style transition, especially for the boundary parts. The whole network follows the adversarial training pipeline in an unsupervised manner guided by the incorporation of adversarial loss and total variation regularization. We evaluate the synthetic images generated by our method from different evaluating metrics including image quality and object detection performance on various prohibited item detection datasets. The results verify that our method can effectively generate realistic X-Ray prohibited item images and improve the detection performance. Hao Wang 0073, Tong Jia 0001, Dongyue Chen 0001, Shizhuo Deng |
IEEE Trans. Image Process. | 2 |
| 2026 | Global and Local Visual-Textual Alignment for Open Vocabulary Object DetectionabstractRecently, with the development of the Vision-Language Model (VLM), adopting such VLM (e.g., CLIP) into object detection framework has gradually become a promising and attractive research direction, and the resulted open vocabulary object detection methods can effectively alleviate the limitations in those close-set ones, making the detectors perceive the unseen world. The core issue in open vocabulary object detection is to design an effective and efficient alignment between the visual (e.g., image) and textual (e.g., caption) features in the semantic space, so that the detectors can capture more information around the open-set scene. Current approaches deploy extra uncurated image-text pairs to pre-train a detector for obtaining a better visual-textual alignment in the feature space. Besides, knowledge distillation technology is also adopted to design an appropriate information transferring flow for aligning the visual-textual knowledge. However, large-scale image-text pairs are not always available to obtain, and the pre-training process will inevitable introduce much more computation overhead. While knowledge distillation methods focus on aligning between the local region visual feature in RoI and the textual features of VLM, neglecting the global information alignment between the image and text. For addressing the dilemmas in these alignment manners, we propose a Global and Local Visual-Textual Alignment for Open Vocabulary Object Detection in this paper. Specifically, our proposed method integrates global image-caption and local region-prompt alignments into a unified learning paradigm. The global alignment takes the whole image and caption as the visual and textual inputs, respectively, and matches the image and caption representations from the detector and the text encoder in CLIP by contrastive learning from the overall perspective. Different from global alignment, the local one concentrates on the accordance between regions and prompts from the aspect of portion description. It extracts and aligns the embeddings for the visual patch RoIs from the image encoder in CLIP and discriminating textual token prompts from the text encoder. Moreover, we also design a prompt tuning strategy, which contains global and local components corresponding to the alignment procedure, for better adapting CLIP to downstream task object detection in a parameter-efficient learning manner. By implementation on Faster R-CNN, we conduct experiments on open vocabulary benchmarks OV-COCO and OV-LVIS, respectively. The results verify that our proposed method can achieve clear improvement over counterparts on novel categories, while performing favorably against state-of-the-arts. Hao Wang 0073, Tong Jia 0001, Shizhuo Deng, Dongyue Chen 0001, Qilong Wang 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 2 |
| 2026 | SLNeRF: Joint Optimization of Structured Light and NeRFabstractRecently, based on multi-view stereo (MVS) methods, utilizing stereo prior information to guide novel-view synthesis has become an important approach to addressing the generalization problem of neural radiance fields (NeRF). However, this approach faces challenges in handling certain difficult scenarios, such as weak texture regions, where it struggles to effectively extract geometric features information. As a result, it encounters limitations in stereo image feature representation and inaccuracies in prior depth estimation. To tackle these problems, we propose the first generalizable novel-view synthesis method that jointly optimizes structured light and NeRF. Considering that active stereo methods based on structured light optimization can enhance geometric feature extraction capability in weak texture regions through the addition of a texture layer, we integrate active stereo vision into novel-view synthesis. To this end, we propose a novel framework, dubbed SLNeRF. First, we design a structured light generation scheme based on Fourier transform and establish a differentiable imaging model using geometric optics and the Lambertian model to generate active stereo images. Then, we obtain stereo image features, as well as prior depth information through the feature extractor, which are used to construct 3D feature volumes. Finally, we accomplish the novel-view synthesis task through a neural renderer. Compared to state-of-the-art generalizable NeRF methods, our method reports encouraging results on public datasets as well as in real-world scenarios. Tong Jia 0001, Shuyang Lin, Dongyue Chen 0001, Ping Xiao, Cuiwei Liu |
IEEE Trans. Multim. | 2 |
| 2025 | Efficient Indoor Depth Completion Network Using Mask-adaptive Gated ConvolutionabstractMost indoor depth completion tasks rely on convolutional auto-encoders to reconstruct depth images, especially in areas with significant missing values. While traditional convolution treats valid and missing pixels equally, Partial Convolution (PConv) has mitigated this limitation. However, PConv fails to distinguish the varying degree of invalidity across different missing areas, which highlights the need for a more refined strategy. To solve this problem, we propose a novel system for indoor depth completion tasks that leverages Mask-adaptive Gated Convolution (MagaConv). MagaConv utilizes gated signals to selectively apply convolution kernels based on the characteristics of missing depth data. These gating signals are generated using shared convolution kernels that jointly process depth features and corresponding masks, ensuring coherent weight optimization. Additionally, the mask undergoes iterative updates according to predefined rules. To improve the fusion of depth and color information, we introduce a Bi-directional Aligning Projection (Bid-AP) module, which utilizes a bi-directional projection scheme with global spatial-channel attention mechanisms to filter out depth-irrelevant features from other modalities. Extensive experiments on popular benchmarks, including NYU-Depth V2, DIML, and SUN RGB-D, demonstrate that our model outperforms state-of-the-art methods in both accuracy and efficiency. Tingxuan Huang, Shizhuo Deng, Tong Jia 0001, Dongyue Chen 0001 |
AAAI | 4 |
| 2025 | DUPL: Domain-agnostic Unknown-aware Prompt Learning for Threshold-free Open-set Domain GeneralizationabstractOpen-set domain generalization (OSDG) aims to recognize known categories in unseen target domains without fine-tuning, while rejecting unknown categories. Existing methods show limited practicality for the requirement of extra generated or collected unknown samples, or for the assumption that the known categories appear in all source domains. Additionally, they need to determine an optimal threshold for distinguishing between known and unknown samples during testing, which is impractical in OSDG, as the target domain is unavailable during training. Besides, they are usually studied with conventional CNNs and shows unsatisfactory generalizability. To address these issues, we harness the transferable property of the pre-trained vision-language model CLIP, and propose domain-agnostic unknown-aware prompt learning (DUPL) framework to achieve threshold-free OSDG. Specifically, we train unknown tokens (UT) to enable threshold-free unknown rejection through unknown-aware prompt learning (UAPL) with only known data, and then introduce a Fourier-based data augmentation (FDA) strategy to obtain domain-agnostic prompts via domain-agnostic semantic consistency (DASC) regularization. Extensive experiments show that our method achieves state-of-the-art OSDG performance. Code is available at https://github.com/X-funbean/DUPL. Fangbin Xu, Dongyue Chen 0001, Shizhuo Deng, Tong Jia 0001, Hao Wang 0073 |
ICME | 4 |
| 2025 | Ensemble CLIPs: Effective Zero-shot Classification with Hundreds of Multi-modal CLIPs
Shizhuo Deng, Zehua Gan, Da Teng, Dongyue Chen 0001, Tong Jia 0001 |
ICMR | 6 |
| 2025 | Noise-Aware Decoding with Salient Region Enhancing for Zero-Shot Image CaptioningabstractImage captioning aims to create natural language descriptions of images. Recent advancements in image captioning have explored text-only training methods that eliminate the need for image annotations. However, these methods are prone to generate descriptions that include objects that do not actually appear in the image, but are instead drawn from the retrieval texts or hard prompts-resulting in object hallucinations. To address this issue, we propose synergistic prompting mechanism called NASCap (Noise-Aware Decoding with Salient Region Enhancing for Zero-Shot Image Captioning). Our method further improves the accuracy of generated captions by designing a fusion model with mask attention that isolate integrates retrieved captions with input features. The hard prompt mixed with negative entities is designed to further improve model's robust to wrong information. Additionally, we introduce a training-free multi-granularity fusion strategy that dynamically perceive and enhance salient regions into global representation. Extensive experiments demonstrate that NASCap sets a new state-of-the art cross-domain (transferable) captioning and performs Through extensive experiments, our straightforward yet powerful approach has demonstrated its efficacy, outperforming the state-of-the-art methods by a significant margin in image captioning compared to zero-shot captioning based on text-only training. Dongyue Chen 0001, Tong Jia 0001, Shizhuo Deng |
ACM Multimedia | 4 |
| 2025 | CSPCL: Category Semantic Prior Contrastive Learning for Deformable DETR-Based Prohibited Item DetectorsabstractProhibited item detection based on X-ray images is one of the most effective security inspection methods. However, the foreground-background feature coupling caused by the overlapping phenomenon specific to X-ray images makes general detectors designed for natural images perform poorly. To address this issue, we propose a Category Semantic Prior Contrastive Learning (CSPCL) mechanism, which aligns the class prototypes perceived by the classifier with the content queries to correct and supplement the missing semantic information responsible for classification, thereby enhancing the model sensitivity to foreground features. To achieve this alignment, we design a specific contrastive loss, CSP loss, which comprises the Intra-Class Truncated Attraction (ITA) loss and the Inter-Class Adaptive Repulsion (IAR) loss, and outperforms classic contrastive losses. Specifically, the ITA loss leverages class prototypes to attract intra-class content queries and preserves essential intra-class diversity via a gradient truncation function. The IAR loss employs class prototypes to adaptively repel inter-class content queries, with the repulsion strength scaled by prototype-prototype similarity, thereby improving inter-class discriminability, especially among similar categories. CSPCL is general and can be easily integrated into Deformable DETR-based models. Extensive experiments on the PIXray, OPIXray, PIDray, and CLCXray datasets demonstrate that CSPCL significantly enhances the performance of various state-of-the-art models without increasing inference complexity. The code is publicly available at https://github.com/Limingyuan001/CSPCL. Mingyuan Li 0003, Tong Jia 0001, Hao Wang 0073, Shiyi Guo, Da Cai, Dongyue Chen 0001 |
NeurIPS | 2 |
| 2025 | Weighted Evidential Continual Learning with Logits-Angle Knowledge Distillation
Dingyu Xue, Shizhuo Deng, Tong Jia 0001, Dongyue Chen 0001 |
PRCV (1) | 4 |
| 2025 | Detection of novel prohibited item categories for real-world security inspection
Shuyang Lin, Tong Jia 0001, Hao Wang 0073, Mingyuan Li 0003, Dongyue Chen 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | RASP: Robot Active Scene Perception With Joint Viewpoint Planning and Depth Completion in Cluttered EnvironmentsabstractCapturing dense visual perception in cluttered environments using depth sensors is crucial for downstream robotics tasks. However, occlusions between objects and unreliable depth data make it very challenging for robots to perceive comprehensive and accurate scene information. To address these issues, we propose a novel robot active scene perception method called RASP, which is composed of two parts. First, we introduce a temporal attention-based view planning algorithm, which actively plans the minimum feasible viewpoint sequence based on latent dependencies in all previous observations to maximize the information perception of the cluttered environments. Subsequently, for the unreliable depth data obtained by the depth sensor, especially caused by transparent and specular objects in the scene, we design a geometry-guided depth completion network that fully utilizes the 3D scene information during the progressive perception process. Specifically, multi-level scene geometric features are extracted and projected into the image space, combining with the image features to guide the depth completion step. These two parts are learned jointly to achieve consistent and accurate results. Extensive experiments demonstrate that our method outperforms state-of-the-art methods. Furthermore, we show that our method significantly improves the performance of downstream grasping tasks. Yizhe Liu, Tong Jia 0001, Hao Wang 0073, Dongyue Chen 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Automatic Label Assignment for Object DetectionabstractLabel assignment, which aims to classify region proposals as positive or negative samples depending on the correlations between their classification and localization predictions with the corresponding ground truth, is recognized as an essential ingredient in object detection and strongly affects the detection performance. Recently, some dynamic label assignment methods have been proposed to overcome the limitations of the static methods and achieve promising performance improvement. Despite eliminating the restrictions of the human prior sampling knowledge in static methods, existing dynamic principles usually suffer from two weaknesses. First, most of them deploy mixture models or implicit branch in prediction head to coarsely estimate the spatial distribution of the positive samples for objects. They give little attention to the effect of appearance information of the objects. Furthermore, these methods still cannot perceive the quality distribution of the positive samples, and these low-quality samples lead to adverse effects on the detection performance. To address issues, this paper presents a novel automatic label assignment for object detection. Specifically, our method first introduces an instance property branch into object detection pipeline to distinguish the foreground from the background. Then, an objectness prediction module which is composed by the confidence and weight mechanisms is developed to generate the positive and negative weight maps for the objects. The instance property branch and objectness prediction module can provide a coarse-to-fine optimization framework to make our method realize the appearance of the objects. Finally, a positive sample selection strategy is proposed to explore the quality statistical distribution of the positive samples, which are trained by different designed label targets. We evaluate our method on the MS COCO dataset and we achieve 48.4%, 47.9%, 48.0% and 49.3% on ResNet-101, ResNeXt-101, DCN-ResNet-101 and DCN-ResNeXt-101 in terms of AP0.5:0.95, respectively. We evaluate the timing complexity of ALA by calculating the inference speed and the frame per second (FPS) for these four backbones are 11.9, 10.4, 9.9 and 8.0, respectively. The experiment results demonstrate that we can obtain clear improvement over the competing methods with favorable performance compared to the state-of-the-arts. Hao Wang 0073, Tong Jia 0001, Qilong Wang 0001, Wangmeng Zuo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Open-Vocabulary Prohibited Item Detection for Real-World X-Ray Security InspectionabstractComputer-aided prohibited item detection is applied in X-ray security inspection to maintain public safety. However, existing prohibited item detectors are limited to a small set of categories in current X-ray datasets, posing potential risks to public security. Since constructing bigger datasets and annotating hundreds of categories is time-consuming and labor-intensive, scaling detectors to more categories with minimal supervision is of great importance. To this end, in this paper, we adopt an open-vocabulary object detection (OVOD) method to detect arbitrary unlabeled novel categories of prohibited item. OVOD methods typically rely on datasets with caption annotations, which are lacking in the domain of prohibited item detection. To support the research on OVOD in X-ray security inspection scenarios, we contribute PIXray Caption dataset, the first X-ray dataset with image-caption pair annotations, which could benchmark and facilitate researches in the community. Further, we propose a novel Open-Vocabulary Prohibited Item Detection (OVPID) network to leverage textual information from captions. OVPID contains two core modules, i.e., Interference Resistant Module (IRM) and Prediction Module (PM). Specifically, IRM includes two submodules, namely Edge Perception (EP) and Foreground Activation (FA), which are designed to address the dilemma of interference caused by overlapping problem and complex background in X-ray images. PM consists of two branches for classification and localization. In classification branch, PM generates more accurate prompts for X-ray dataset via large multimodal model (LMM). In localization branch, PM aligns the student embeddings with both teacher and caption embeddings. Extensive experiments on PIXray Caption dataset demonstrate that OVPID outperforms other OVOD methods by delivering a higher accuracy on novel categories. Shuyang Lin, Tong Jia 0001, Hao Wang 0073, Mingyuan Li 0003 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Low-Overlap Point Cloud Registration by Semiglobal Block MatchingabstractIn recent years, the local feature-based point cloud registration methods have attracted more attention for their high robustness. The state-of-the-art local feature-based pipeline consists of local feature extraction, feature matching, and outlier rejection. However, the accuracy of this pipeline decreases significantly between low-overlap point clouds. In this article, we propose a novel hand-crafted framework to realize efficient low-overlap point cloud registration through point cloud partitioning, semiglobal point cloud block matching, and best transformation selection with refinement. Experimental results on benchmark datasets show that the proposed algorithm presents the superior performance over the previous methods on low-overlap scenes under different features and achieves very competitive accuracy on non-low-overlap data. Shiyi Guo, Yihong Wu 0002, Binjian Xie, Bingxi Liu 0001, Tong Jia 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Meta-TIP: An Unsupervised End-to-End Fusion Network for Multi-Dataset Style-Adaptive Threat Image ProjectionabstractThreat Image Projection (TIP) is a convenient and effective means to expand X-ray baggage images, which is essential for training both security personnel and computer-aided screening systems. Existing methods are primarily divided into two categories: X-ray imaging principle-based methods and GAN-based generative methods. The former cast prohibited items acquisition and projection as two individual steps and rarely consider the style consistency between the source prohibited items and target X-ray images from different datasets, making them less flexible and reliable for practical applications. Although GAN-based methods can directly generate visually consistent prohibited items on target images, they suffer from unstable training and lack of interpretability, which significantly impact the quality of the generated items. To overcome these limitations, we present a conceptually simple, flexible and unsupervised end-to-end TIP framework, termed as Meta-TIP, which superimposes the prohibited item distilled from the source image onto the target image in a style-adaptive manner. Specifically, Meta-TIP mainly applies three innovations: 1) reconstruct a pure prohibited item from a cluttered source image with a novel foreground-background contrastive loss; 2) a material-aware style-adaptive projection module learns two modulation parameters pertinently based on the style of similar material objects in the target image to control the appearance of prohibited items; 3) a novel logarithmic form loss is well-designed based on the principle of TIP to optimize synthetic results in an unsupervised manner. We comprehensively verify the authenticity and training effect of the synthetic X-ray images on four public datasets, i.e., SIXray, OPIXray, PIXray, and PIDray dataset, and the results confirm that our framework can flexibly generate very realistic synthetic images without any limitations. Tong Jia 0001, Hao Wang 0073, Dongyue Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | WS-SAM: Generalizing SAM to Weakly Supervised Object Detection With Category LabelabstractBuilding an effective object detector usually depends on large well-annotated training samples. While annotating such dataset is extremely laborious and costly, where box-level supervision which contains both accurate classification category and localization coordinate is required. Compared to above box-level supervised annotation, those weakly supervised learning manners (e.g,, category, point and scribble) need relatively less laborious annotation cost, and provide a feasible way to mitigate the reliance on the dataset. Because of the lack of sufficient supervised information, current weakly supervised methods cannot achieve satisfactory detection performance. Recently, Segment Anything Model (SAM) has appeared as a task-agnostic foundation model and shown promising performance improvement in many related works due to its powerful generalization and data processing abilities. The properties of the SAM inspire us to adopt such basic benchmark to weakly supervised object detection field to compensate the deficiencies in supervised information. However, directly deploying SAM on weakly supervised object detection task meets with two issues. Firstly, SAM needs meticulously-designed prompts, and such expert-level prompts restrict their applicability and practicality. Besides, SAM is a category unawareness model, and it cannot assign the category labels to the generated predictions. To solve above issues, we propose WS-SAM, which generalizes Segment Anything Model (SAM) to weakly supervised object detection with category label. Specifically, we design an adaptive prompt generator to take full advantages of the spatial and semantic information from the prompt. It employs in a self-prompting manner by taking the output of SAM from the previous iteration as the prompt input to guide the next iteration, where the prompts can be adaptively generated based on the classification activation map. We also develop a segmentation mask refinement module and formulate the label assignment process as a shortest path optimization problem by considering the similarity between each location and prompts. Furthermore, a bidirectional adapter is also implemented to resolve the domain discrepancy by incorporating domain-specific information. We evaluate the effectiveness of our method on several detection datasets (e.g., PASCAL VOC and MS COCO), and the experiment results show that our proposed method can achieve clear improvement over state-of-the-art methods, while performing favorably against state-of-the-arts. Hao Wang 0073, Tong Jia 0001, Qilong Wang 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 2 |
| 2025 | AO-DETR: Anti-Overlapping DETR for X-Ray Prohibited Items DetectionabstractProhibited item detection in X-ray images is one of the most essential and highly effective methods widely employed in various security inspection scenarios. Considering the significant overlapping phenomenon in X-ray prohibited item images, we propose an anti-overlapping detection transformer (AO-DETR) based on one of the state-of-the-art (SOTA) general object detectors, DETR with improved denoising anchor boxes (DINO). Specifically, to address the feature coupling issue caused by overlapping phenomena, we introduce the category-specific one-to-one assignment (CSA) strategy to constrain category-specific object queries in predicting prohibited items of fixed categories, which can enhance their ability to extract features specific to prohibited items of a particular category from the overlapping foreground-background features. To address the edge blurring problem caused by overlapping phenomena, we propose the look forward densely (LFD) scheme, which improves the localization accuracy of reference boxes in mid-to-high-level decoder layers and enhances the ability to locate blurry edges of the final layer. Similar to DINO, our AO-DETR provides two different versions with distinct backbones, tailored to meet diverse application requirements. Extensive experiments on the PIXray, OPIXray, and HIXray datasets demonstrate that the proposed method surpasses the SOTA object detectors, indicating its potential applications in the field of prohibited item detection. The source code will be available at: https://github.com/Limingyuan001/AO-DETR. Mingyuan Li 0003, Tong Jia 0001, Hao Wang 0073, Shuyang Lin, Da Cai, Dongyue Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Self-Supervised Federated Learning for Personalized Human Activity RecognitionabstractPersonalized Human Activity Recognition (PHAR) based on wearable sensors is crucial in the medical, sports, industrial and other fields. PHAR faces challenges of privacy leakage and a shortage of labeled data. Therefore, we propose a framework called self-supervised federated learning for personalized human activity recognition (SSF-HAR) to implement private PHAR. To protect user privacy, our framework integrates federated learning (FL) to achieve the transmission of only model parameters between the cloud and clients, rather than user data. Besides, we propose a strategy of weighted aggregation to update the cloud model with the client models. To overcome the lack of labeled data, our framework introduces self-supervised learning (SSL) tasks to pretrain a feature extractor in the cloud. The proxy task of SSL transforms data and provides pseudo-labels in three forms. We test the performance on the benchmark datasets MotionSense and WIDSM. The experiments show that SSF-HAR outperforms other FL frameworks for PHAR. Shizhuo Deng, Da Teng, Zhubao Guo, Dongyue Chen 0001, Tong Jia 0001, Hao Wang 0073 |
ICME | 6 |
| 2024 | APPN: An Attention-based Pseudo-label Propagation Network for few-shot learning with noisy labels
Shizhuo Deng, Da Teng, Dongyue Chen 0001, Tong Jia 0001, Hao Wang 0073 |
Neurocomputing | 5 |
| 2024 | A lightweight Transformer-based visual question answering network with Weight-Sharing Hybrid Attention
Dongyue Chen 0001, Tong Jia 0001, Shizhuo Deng |
Neurocomputing | 3 |
| 2024 | Ensembling disentangled domain-specific prompts for domain generalization
Fangbin Xu, Shizhuo Deng, Tong Jia 0001, Xiaosheng Yu 0001, Dongyue Chen 0001 |
Knowl. Based Syst. | 3 |
| 2024 | Conjoined triple deep network for video anomaly detection
Xingya Chang, Yunhe Wu, Shizhuo Deng, Tong Jia 0001, Dongyue Chen 0001 |
Multim. Tools Appl. | 4 |
| 2024 | Toward Dual-View X-Ray Baggage Inspection: A Large-Scale Benchmark and Adaptive Hierarchical Cross Refinement for Prohibited Item DiscoveryabstractDual-view baggage inspection has been widely applied in real-world scenarios, where orthogonal viewpoints are deployed to capture diverse and complementary information. Compared with single-view, it can effectively improve the identification performance when rotation and overlay hinder the viewability of the objects. However, this topic has not been rigorously explored due to the scarcity of datasets. To overcome this limitation, we contribute the first fully public large-scale Dual-view X-ray dataset. Our dataset, named DvXray, contains 16,000 pairs, 32,000 X-ray images, in which 15 common classes of 5,496 prohibited items are manually labeled. Besides, we propose an approach named Adaptive Hierarchical Cross Refinement (AHCR) to establish a strong baseline for prohibited item discovery in dual-view X-ray images. AHCR hypothesizes that each input pair is sampled from one mixture distribution, hence gathering the non-overlapping and position-aware cues along the shared axis and complementarily delivering to the other in a hierarchical structure to enrich the feature discriminability of the objects of interest from background overlaps. Upon this structure, we propose an adaptive control strategy and a confidence-weighted view fusion term to make it robust to difficult samples. Extensive experiments on DvXray show that AHCR not only brings significant classification gains over various backbones, such as recent Swin Transformer and ConvNeXt, but also exhibits an impressively better ability to localize objects. In addition, AHCR performs favorably against the counterparts and some recent multi-view learning approaches, moving a step closer towards potential application in practice. Dataset and code are available at https://github.com/Mbwslib/DvXray. Tong Jia 0001, Mingyuan Li 0003, Songsheng Wu, Hao Wang 0073, Dongyue Chen 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Delving Into Cluttered Prohibited Item Detection for Security Inspection SystemabstractProhibited item detection in X-Ray baggage images can efficiently prevent the social security and stability. With the development of deep learning, applying specific methods in computer vision tasks to prohibited item detection has shown promising perspectives, and the resulted intelligent security inspection system can effectively address the limitations of human inspection. Despite several deep learning based methods have been proposed to flourish this researching field, there are still two issues which have not been fully explored. First, most of them suffer from the shortage of dataset and collection of massive and well-annotated samples is extremely laborious and costly. Second, the designation of backbone modules for current prohibited item detection methods shares similar idea with the ones in object detection. While little consideration has been paid to the properties of prohibited item, where most of them are over-lapped and cluttered. To overcome limitations, this paper proposes a cluttered prohibited item detection method for security inspection system. Specifically, our method first generates synthetic X-Ray images through cut-and-paste strategy from the training samples in each training mini-batch, where the strategy effectively and efficiently augments the dataset and quality of the synthetic samples can be guaranteed. Then, a high-order dilated convolution module is developed for enriching the representation ability of the feature, further promoting the localization ability for over-lapped and cluttered prohibited items. Experiments show that our proposed method can be well generalized to various datasets, and achieve clear improvement over state-of-the-art methods. Hao Wang 0073, Tong Jia 0001, Dongyue Chen 0001, Shizhuo Deng |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Relation Knowledge Distillation by Auxiliary Learning for Object DetectionabstractBalancing the trade-off between accuracy and speed for obtaining higher performance without sacrificing the inference time is a challenging topic for object detection task. Knowledge distillation, which serves as a kind of model compression techniques, provides a potential and feasible way to handle above efficiency and effectiveness issue through transferring the dark knowledge from the sophisticated teacher detector to the simple student one. Despite demonstrating promising solutions to make harmonies between accuracy and speed, current knowledge distillation for object detection methods still suffer from two limitations. Firstly, most of the methods are inherited or refereed from the frameworks in image classification task, and deploy an implicit manner by imitating or constraining the features from the intermediate layers or the output predictions between the teacher and student models. While little consideration has been raised to the intrinsic relevance of the classification and localization predictions in object detection task. Besides, these methods fail to investigate the relationship between detection and distillation tasks in knowledge distillation pipeline, and they train the whole network by simply integrating losses from these two different tasks through hand-crafted designation parameters. For addressing the aforementioned issues, we propose a novel Relation Knowledge Distillation by Auxiliary Learning for Object Detection (ReAL) method in this paper. Specifically, we first design a prediction relation distillation module which makes the student model directly mimic the output predictions from the teacher one, and conduct self and mutual relation distillation losses to excavate the relation information between teacher and student models. Moreover, for better devolving into the relationship between different tasks in distillation pipeline, we introduce the auxiliary learning into knowledge distillation for object detection and develop a dynamic weight adaptation strategy. Through regarding detection task as primary task and treating distillation task as auxiliary task in auxiliary learning framework, we dynamically adjust and regularize the corresponding weights of the losses for these tasks during the training process. Experiments on MS COCO dataset are conducted using various detector combinations of teacher and student models and the results show that our proposed ReAL can achieve obvious improvement on different distillation model configurations, while performing favorably against state-of-the-arts. Hao Wang 0073, Tong Jia 0001, Qilong Wang 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 2 |
| 2024 | LHAR: Lightweight Human Activity Recognition on Knowledge DistillationabstractSensor-based Human Activity Recognition (HAR) is widely used in daily life and is the basic-level bridge to virtual healthcare in the metaverse. The current challenge is the low recognition accuracy for personalized users on smart wearable devices. The limited resource cannot support large deep learning models updated locally. Besides, integrating and transmitting sensor data to the cloud would reduce the efficiency. Considering the tradeoff between performance and complexity, we propose a Lightweight Human Activity Recognition (LHAR) framework. In LHAR, we combine the cross-people HAR task with the lightweight model task. LHAR framework is designed on the teacher-student architecture and the student network consists of multiple depthwise separable convolution layers to achieve fewer parameters. The dark knowledge distilled from the complex teacher model enhances the generalization ability of LHAR. To achieve effective knowledge distillation, we propose two optimization methods. Firstly, we train the teacher model by ensemble learning to promote teacher performance. Secondly, a multi-channel data augmentation method is proposed for the diversity of the dataset, which is a plug-in operation for the ensemble teacher model. In the experiments, we compare LHAR with state-of-art models in comparison evaluation, ablation study and the hyperparameter analysis, which proves the better performance of LHAR in efficiency and effectiveness. Shizhuo Deng, Da Teng, Chuangui Yang, Dongyue Chen 0001, Tong Jia 0001, Hao Wang 0073 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | CBDMoE: Consistent-but-Diverse Mixture of Experts for Domain GeneralizationabstractMachine learning models often suffer from severe performance degradation due to distributional shifts between testing and training data. To address this issue, researchers have focused on domain generalization (DG), which aims to generalize a model trained on source domains to arbitrary unseen target domains. Recently, ensemble learning has emerged as a popular strategy for addressing the DG problem, and domain-specific experts are typically involved. However, the existing methods do not sufficiently consider the generalizability of individual experts or leverage the consistency and diversity among them, thus limiting the generalizability of the constructed models. In this paper, we propose a consistent-but-diverse mixture of experts (CBDMoE) algorithm, which is an improved MoE framework that effectively harnesses ensemble learning for solving the DG problem. Specifically, we introduce individual expert learning (IEL), which incorporates a novel domain-class-balanced subset division (DCBSD)-based sampling strategy to facilitate a generalizable expert learning process. Additionally, we present consistent-but-diverse learning (CBDL), which employs two regularizing losses to encourage consistency and diversity in the predictions of the experts. Our proposed strategy significantly enhances the generalizability of the MoE framework. Extensive experiments conducted on three popular DG benchmark datasets demonstrate that our method outperforms the state-of-the-art approaches. Fangbin Xu, Dongyue Chen 0001, Tong Jia 0001, Shizhuo Deng, Hao Wang 0073 |
IEEE Trans. Multim. | 3 |
| 2024 | Bmsmlet: boosting multi-scale information on multi-level aggregated features for salient object detection
Tong Jia 0001, Yunhe Wu, Zhikang Zeng |
Vis. Comput. | 2 |
| 2023 | AGG-Net: Attention Guided Gated-convolutional Network for Depth Image CompletionabstractRecently, stereo vision based on lightweight RGBD cameras has been widely used in various fields. However, limited by the imaging principles, the commonly used RGB-D cameras based on TOF, structured light, or binocular vision acquire some invalid data inevitably, such as weak reflection, boundary shadows, and artifacts, which may bring adverse impacts to the follow-up work. In this paper, we propose a new model for depth image completion based on the Attention Guided Gated-convolutional Network (AGG-Net), through which more accurate and reliable depth images can be obtained from the raw depth maps and the corresponding RGB images. Our model employs a UNet-like architecture which consists of two parallel branches of depth and color features. In the encoding stage, an Attention Guided Gated-Convolution (AG-GConv) module is proposed to realize the fusion of depth and color features at different scales, which can effectively reduce the negative impacts of invalid depth data on the reconstruction. In the decoding stage, an Attention Guided Skip Connection (AG-SC) module is presented to avoid introducing too many depth-irrelevant features to the reconstruction. The experimental results demonstrate that our method outperforms the state-of-the-art methods on the popular benchmarks NYU-Depth V2, DIML, and SUN RGB-D. https://github.com/htx0601/AGG-Net Dongyue Chen 0001, Tingxuan Huang, Zhimin Song, Shizhuo Deng, Tong Jia 0001 |
ICCV | 5 |
| 2023 | Fast Personalized Human Activity Recognition on Heuristic Parameter EstimationabstractPersonalized human activity recognition (HAR) on wearable sensor data is crucial for healthcare, industry and sport. Personalized HAR has two challenges: very few high-quality labeled personalized data and limited terminal computing resource. Most previous domain adaptation models ignore the transfer efficiency on the terminal, which is also one criterion of Artificial Intelligence of Things. Therefore, we propose a fast transfer HAR, FastTrans, to improve the efficiency and get a trade-off with recognition effectiveness. A fusion feature extraction module is designed to learn multi-scale features on the improved hybrid loss. The proposed heuristic parameter estimation method learns the approximate solutions of the classification weights in FastTrans by scanning the adaptation data only once. Besides, an efficient time series data augmentation is proposed as a plugin for dataset variety. The results on benchmark datasets show the dramatically competitive performance of FastTrans on transferring efficiency with close accuracy to other models. Shizhuo Deng, Chuangui Yang, Zhubao Guo, Boqian Lin, Dongyue Chen 0001, Tong Jia 0001 |
ICME | 6 |
| 2023 | Self-relation attention networks for weakly supervised few-shot activity recognition
Shizhuo Deng, Zhubao Guo, Da Teng, Boqian Lin, Dongyue Chen 0001, Tong Jia 0001, Hao Wang 0073 |
Knowl. Based Syst. | 6 |
| 2023 | Multi-information Constraint Learning for Unsupervised Domain Adaptive Person Re-identification
Dongyue Chen 0001, Bing Haozhe, Chunren Tang, Miaoting Tian, Tong Jia 0001 |
Neural Process. Lett. | 5 |
| 2023 | Fully Cascade Consistency Learning for One-Stage Object DetectionabstractObject detection is usually solved by deploying one single prediction head including classification and localization branches to obtain the final results. Recently proposed works utilize several prediction heads in a cascade learning manner to improve the detection performance. Despite achieving promising performance, existing cascade learning manner methods still meet with two inconsistency issues. Firstly, most of them refine the bounding boxes in different prediction heads only by depending on the localization accuracy (i.e., IoU), while ignoring the inconsistency between classification confidence and localization accuracy. Moreover, simply increasing the IoU threshold by experience to select positive samples makes the inconsistency issue even worse. Secondly, little consideration has been paid on the feature inconsistency between detection-specific features from different prediction heads and detection-generalized ones from backbone model. The extracted feature from backbone model contains the general representation for the whole images. While prediction heads need to be carefully designed to have specific ability which contains more discriminative expressions for the two sub-tasks classification and regression. The different contexture representations of the output features from these two parts lead to the feature inconsistency between backbone model and prediction head in cascade learning architecture. To solve these two inconsistency issues, this paper proposes a novel cascade consistency learning method for one-stage detector. Specifically, a feature adaptation module is firstly developed to calibrate features from different prediction heads and backbone model for solving the feature inconsistency. Then, we design an automatic positive sample threshold selection strategy for further solve the inconsistency between the classification and localization predictions. Moreover, the quality of bounding boxes in cascade learning manner are evaluated by taking both the classification confidence and localization accuracy into consideration. Experiments on MS COCO show that our proposed cascade consistency learning manner (dubbed$\text{C}^{2}\text{L}$) can achieve clear improvement over counterparts based on several different one-stage detectors, while performing favorably against state-of-the-arts. Hao Wang 0073, Tong Jia 0001, Qilong Wang 0001, Wangmeng Zuo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Automated Segmentation of Prohibited Items in X-Ray Baggage Images Using Dense De-Overlap Attention SnakeabstractProhibited item segmentation has a wide range of applications in the security check field, such as computer-aided screening, threat image projection and material discrimination. However, the severe object overlapping in X-ray baggage images restricts the performance of common CNN-based segmentation methods greatly. Worse, no public dataset can be used to promote research in this challenging and promising area. In this paper, to cope with these problems, we present the first Prohibited Item X-ray segmentation dataset named PIXray. PIXray comprises 5,046 X-ray images, in which 15 classes of 15,201 prohibited items are annotated as instance-level masks. Besides, we contribute a dense de-overlap attention snake (DDoAS) in the context of deep learning for automated and real-time prohibited item segmentation. DDoAS mainly includes a dense de-overlap module (DDoM) and an attention deforming module (ADM). Specifically, DDoM is designed to infer prohibited item information accurately from extreme background overlaps through dense reversed connections. ADM aims to improve the low learning efficiency introduced by large variations in shapes and sizes among different prohibited items. Comprehensive evaluation on the PIXray shows the effectiveness and superiority of DDoM and ADM. DDoM excels at recognizing prohibited items from complex backgrounds than other in-domain methods and achieves consistent performance gain over various network backbones, extending the idea of tackling overlapping images data. ADM can ease the model training and further refine the mask quality. Furthermore, out-of-domain experiments prove that DDoAS can also be applied to natural images and achieves comparable performance to the state-of-the-art methods, which implies its potential applications in other fields. The dataset and source code are available athttps://github.com/Mbwslib/DDoAS. Tong Jia 0001, Dongyue Chen 0001, Yichun Zhang |
IEEE Trans. Multim. | 2 |
| 2023 | 3D human body reconstruction based on SMPL model
Dongyue Chen 0001, Yuanyuan Song, Fangzheng Liang, Tong Jia 0001 |
Vis. Comput. | 6 |
| 2023 | Two-stage salient object detection based on prior distribution learning and saliency consistency optimization
Yunhe Wu, Xingya Chang, Dongyue Chen 0001, Tong Jia 0001 |
Vis. Comput. | 5 |
| 2023 | Publisher Correction: Two-stage salient object detection based on prior distribution learning and saliency consistency optimization
Yunhe Wu, Xingya Chang, Dongyue Chen 0001, Tong Jia 0001 |
Vis. Comput. | 5 |
| 2023 | Dual-stream stereo network for depth estimation
Yangyang Zhong, Tong Jia 0001, Kaiqi Xi, Wenhao Li 0012, Dongyue Chen 0001 |
Vis. Comput. | 2 |
| 2022 | HOB-net: high-order block network via deep metric learning for person re-identification
Dongyue Chen 0001, Tong Jia 0001, Fangbin Xu |
Appl. Intell. | 3 |
| 2022 | Weakly Supervised Anomaly Detection Based on Two-Step Cyclic Iterative PU Learning Strategy
Dongyue Chen 0001, Xinyue Tantai, Xingya Chang, Miaoting Tian, Tong Jia 0001 |
Neural Process. Lett. | 5 |
| 2021 | Salient object detection via a boundary-guided graph structure
Yunhe Wu, Tong Jia 0001, Jiaduo Sun, Dingyu Xue |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | NM-GAN: Noise-modulated generative adversarial network for video anomaly detection
Dongyue Chen 0001, Lingyi Yue, Xingya Chang, Ming Xu 0007, Tong Jia 0001 |
Pattern Recognit. | 5 |
| 2020 | Anomaly detection in surveillance video based on bidirectional prediction
Dongyue Chen 0001, Lingyi Yue, Tong Jia 0001 |
Image Vis. Comput. | 5 |
| 2016 | Visual saliency detection: From space to frequency
Dongyue Chen 0001, Tong Jia 0001, Chengdong Wu 0001 |
Signal Process. Image Commun. | 2 |
| 2016 | Scene Depth Perception Based on Omnidirectional Structured LightabstractA depth perception method combining omnidirectional images and encoding structured light was proposed. First, a new structured light pattern was presented by using monochromatic light. The primitive of the pattern consists of four-direction sand clock-like (FDSC) image. FDSC can provide more robust and accurate position compared with conventional pattern primitive. Second, on the basis of multiple reference planes, a calibration method of projector was proposed to significantly simplify projector calibration in the constructed omnidirectional imaging system. Third, a depth point cloud matching algorithm based on the principle of prior constraint iterative closest point under mobile condition was proposed to avoid the effect of occlusion. The experimental results demonstrated that the proposed method can acquire omnidirectional depth information about large-scale scenes. The error analysis of 16 groups of depth data reported a maximum measuring error of 0.53 mm and an average measuring error of 0.25 mm. Tong Jia 0001, ZhongXuan Zhou, Haixiu Meng |
IEEE Trans. Image Process. | 1 |
| 2015 | 3D depth information extraction with omni-directional camera
Tong Jia 0001, Yan Shi 0006, ZhongXuan Zhou, Dongyue Chen 0001 |
Inf. Process. Lett. | 1 |
| 2012 | Gradient Vector Flow Based on Anisotropic Diffusion
Xiaosheng Yu 0001, Chengdong Wu 0001, Dongyue Chen 0001, Tong Jia 0001 |
ISNN (2) | 5 |