EDBT 2026 Demo / reviewers in the wild / expert
Zhuangzhuang Chen
dblp:224/1766
· DBLP profile ↗
33ranked-venue papers
12as first author
29since 2021 · last 2026
0000-0002-0336-1181ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 7 since 2021Computer networks · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoGenSAM: Codebook-Interactive Generative Labeling for Adapting SAM to Crack SegmentationabstractThe goal of this work is to adapt Segment Anything Models (SAM) into crack segmentation tasks via automatic label generation, thus eliminating manual annotation cost. In this regard, an intuitive approach is to extract edges of crack samples and generate labels via the dilation and erosion processes for fine-tuning SAM. However, this simple solution cannot guarantee the quality of generated labels, as crack regions will be corrupted due to the imperfect edge detection. To this end, this paper proposes CoGenSAM, a novel Codebook-interactive Generative Labeling framework that enables an annotation-free SAM fine-tuning. To achieve this, in the first stage, we pre-train a vector-quantized variational auto-encoder (VQVAE) by reconstructing the synthesized crack-like structures for learning crack-aware priors within the codebook. In the second stage, these priors help another VQVAE serve as the restoration model to restore the randomly corrupted structures into uncorrupted ones. Specifically, we propose the crack-aware contrastive-interaction to maximize the mutual information with the above priors via codebook interaction. Then, high-quality labels can be generated by restoring corrupted labels from edge detection, contributing to an annotation-free SAM fine-tuning. We collect a new dataset, Bridge2025, to address the limited availability of related bridge-oriented benchmarks. Experiments show that our performance is close to fully-supervised methods. Zhuangzhuang Chen, Dachong Li, Zhiliang Lin, Xingyu Feng 0001, Jie Chen 0027, Jianqiang Li 0001 |
AAAI | 1 |
| 2026 | CaPro: Curvilinear-aware Prompt Learning with Single Unlabeled Image for Cost-effective Curvilinear Structure SegmentationabstractCurvilinear structure segmentation (CSS) plays a vital role in industrial applications, including medical imaging and structural health monitoring. Recently, the strong capacity of the Segment Anything Model (SAM) has inspired its downstream application in CSS tasks. To adapt SAM to CSS tasks, previous methods heavily rely on a certain number of samples and costly pixel-level annotation, which are hard to access for a new scenario. Considering this, the goal of our work is to adapt SAM in a very cost-effective setting where only a single unlabeled image is given. This is far more challenging than the typical supervised, unsupervised, or self-supervised learning manner that needs a large number of training samples. To tackle this problem, we propose a finetuning-free SAM for curvilinear structure segmentation, called curvilinear-aware prompt learning (CaPro), which aims to automatically learn visual prompts via a single unlabeled image. In the first stage, we generate extensive curvilinear structures and oriented sub-curvilinear box annotations. To increase the realism of generated curvilinear structures, we adapt these structures into real image domains via the Fourier Transform using a single real-world unlabeled image. Now, these adapted images can be used to train our oriented sub-curvilinear detector. In the second stage, we propose the curvilinear-aware discrete representation matching to filter those unreliable detection results. Afterward, these reliable detection results can be converted into informative prompts, contributing to the cost-effective SAM adaptation to CSS tasks. Experiments demonstrate the effectiveness of CaPro on medical image and crack segmentation tasks. Zhuangzhuang Chen, Qiangyu Chen, Chubin Ou, Xiaomeng Li 0001 |
AAAI | 1 |
| 2026 | Beyond feature mapping: Dual-heterogeneous knowledge distillation with mamba for industrial anomaly detection
Muhao Xu, Zihan Nie, Baochen Fu, Zhuangzhuang Chen, Hua Wei 0007, Yi Wan 0002, Weiye Song |
Expert Syst. Appl. | 4 |
| 2026 | A multilevel alignment and cross-fusion knowledge distillation framework for vision transformer-based medical image segmentation
Pengchen Liang, Jianguo Chen 0001, Renkai Wu, Zhuangzhuang Chen, Bin Pu, Qing Chang 0004, Guo Ran |
Future Gener. Comput. Syst. | 5 |
| 2026 | Dual dynamic graph attention network driven deep reinforcement learning for flexible job-Shop scheduling
Yan Kang 0003, Tianjing Li, Lei Zhao 0013, Zhuangzhuang Chen, Bin Pu |
Knowl. Based Syst. | 6 |
| 2026 | MTLQ-ViT: Multi-granularity Tail-enhanced Logarithmic Quantization for Vision Transformers
Yan Kang 0003, Shouhao Xu, Qika Lin, Kai He 0001, Zhuangzhuang Chen, Bin Pu |
Pattern Recognit. | 6 |
| 2026 | Anomaly-aware transitions for reward-free offline imitation learning
Zhiliang Lin, Zhuangzhuang Chen, Guanming Zhu, Jie Chen 0027 |
Pattern Recognit. | 2 |
| 2026 | Decompose-Compose Feature Augmentation for Imbalanced Crack Recognition in Industrial ScenariosabstractAutomated crack recognition has achieved remarkable progress in the past decades as a critical task in structure health monitoring, to ensure safety and durability in many industrial scenarios. However, imbalanced crack recognition remains challenging due to the scarcity of crack samples and the consequential limited diversity. To resolve this, Artificial Intelligence Generated Content (AIGC) has been gradually adopted to generate synthetic data and reduce reliance on large amounts of labeled crack samples. This paper assumes that a crack sample in the feature space can be regarded as a combination of crack and background semantics. Then, the decompose-compose feature augmentation framework (DeCo) is proposed to perform crack data synthesis in the feature space by randomly composing crack and background semantic-relevant features. Specifically, the contrastive learning-based decomposing loss is proposed to enforce two encoders to separately learn crack and background semantics from crack samples with the theoretical guarantee. After that, an effective cross-instance feature union strategy is proposed to synthesize diverse crack samples by composing the crack-relevant features from a crack sample and background-relevant features across other training samples. To address the limited availability of related benchmarks, we collect INPP2022 and IRC2022 datasets from real-world applications in nuclear power plants and road pavement. Experimental results show that DeCo performs favorably against state-of-the-art competitors in imbalanced crack recognition tasks. Zhuangzhuang Chen, Chengqi Xu, Tao Hu 0023, Li Wang 0093, Jie Chen 0027, Jianqiang Li 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | Toward Reliable Imitation Learning With Limited Expert Demonstrations via Search-Based Inverse Dynamic LearningabstractImitation learning (IL) shows its superiority in faster strategy optimization by leveraging expert demonstrations. However, as one of the powerful IL methods, behavior cloning (BC) suffers from covariate shift problems, where the agent’s policy drifts away from the expert’s, leading to compounding errors and decreased generalization performance. To this end, we propose Search-based Inverse Dynamics Imitation Learning, namely SIDIL, to enhance the robustness of imitation learning by augmenting expert demonstrations via trajectory perturbation and stitching. Specifically, SIDIL first employs a nearest-neighbor search method to find the closest points between expert and perturbed data, generating new actions near the expert data via a stitching strategy. By exploiting this, the agent can recover from deviations and complete tasks under a wider range of conditions. Experimental results on various robotic tasks show that SIDIL outperforms baseline algorithms with higher success rates across multiple tasks. Meanwhile, SIDIL is allowed to expand the attractive region around expert demonstrations by stitching states from expert demonstrations, enjoying a higher task completion success rate even for states outside the expert distribution. These results highlight our potential to enhance humanoid locomotion and dexterous robotic hand manipulation, making it particularly suitable for industrial manufacturing automation where adaptability, robustness, and safety are essential. Zhiliang Lin, Zhuangzhuang Chen, Guanming Zhu, Li Wang 0093, Jianqiang Li 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | Prioritizing Expert-Like Transitions for Reward-Free Offline Imitation LearningabstractOffline imitation learning (IL) enables embodied agents of AI-driven automation control systems to acquire policies from demonstrations without interacting with the environment. However, collecting extensive expert demonstrations that rely on a manually designed reward function is labor-intensive, and even impractical for dynamic real-world settings. To mitigate the reliance on large amounts of reward-annotated expert data, previous works have made great efforts by leveraging both expert and suboptimal demonstrations. In this paper, we reveal that existing methods still suffer from suboptimal results for two reasons: (i) reward signals in suboptimal demonstrations are not reliable, due to these rewards derived from inexperienced annotators, thus can not reflect the actual rewards, and (ii) those suboptimal demonstrations, which are significantly distant from expert behaviors, result in low reconstructed rewards. Moreover, motivated by the recent diffusion-based foundation model, we propose Expert-Like Transition Imitation Learning (ELTIL), a novel generative trajectory augmentation method that effectively leverages diffusion models. ELTIL re-designs the reward function based on expert proximity, and then proposes amplified return conditioning for achieving high-quality trajectory generation. More specifically, instead of relying on environment-provided rewards or manually designed rewards, ELTIL first serves as a good reward labeler for transitions by measuring their distance to expert states. These re-designed rewards then serve as guidance in the diffusion process, with amplified returns contributing as conditioning values to bias generation toward expert-consistent behaviors. Extensive experiments on diverse D4RL benchmarks demonstrate that ELTIL consistently improves the performance and stability of reward learning and offline IL methods, maintaining reliability under varying noise levels, reward sparsity, and dataset distributions. These characteristics make ELTIL well-suited for deployment in robust and reliable automated control systems such as robotic manipulation and autonomous driving. Zhiliang Lin, Zhuangzhuang Chen, Guanming Zhu, Li Wang 0093, Jianqiang Li 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | GALC: Guided Amplified Learning With Lipschitz Constraint for Robust Trajectory GenerationabstractOffline reinforcement learning (RL) demonstrated remarkable performance in learning valid policies by benefiting from high-quality offline datasets. However, collecting such a dataset is labor-intensive, especially for humanoid locomotion. For this reason, many data augmentation techniques have been proposed to improve the quality of offline datasets through noise injection or data synthesis. However, existing data augmentation methods are noise-sensitive, resulting in limited capability in complex robotic environments. To address these issues, we propose guided amplified learning with Lipschitz constraint (GALC), a novel trajectory augmentation method that employs the reward-amplification-guided conditional diffusion model for noise-insensitive data augmentation. Specifically, we introduce a local Lipschitz continuity constraint to regulate the reverse denoising process from the offline dataset. Consequently, the exploration of the diffusion model can be restricted within the local continuity region of the original dataset, thereby generating high-reward trajectories. Moreover, the generated trajectories are also enforced to be noise-insensitive to perturbations, thus enjoying robustness. Notably, our proposed method can prevent the generation of unsafe actions that do not align with the environment dynamics. Extensive experiments on sparse reward scenarios and high-dimensional robotic tasks show that our proposed GALC achieves significant improvements in both the augmented trajectories and policy performance. Zhiliang Lin, Zhuangzhuang Chen, Guanming Zhu, Jiaxian Chen, Jianqiang Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2026 | E$^{2}$2LLM: Structure-Guided Efficient Inference for LLMs in Distributed Edge-IoT EnvironmentsabstractLarge language models (LLMs) are increasingly deployed in edge computing environments to reduce latency and preserve privacy. However, their inference process presents fundamental challenges for resource-constrained IoT devices. LLM inference involves computationally asymmetric stages: parallelizable prompt processing and sequential token decoding. This asymmetry creates deployment bottlenecks where IoT devices lack capacity for prompt processing while edge nodes suffer from inefficient sequential decoding. This paper presentsE$^{2}$LLM, an efficient distributed inference framework for large language models in heterogeneous edge-IoT environments.E$^{2}$LLMleverages high-capacity edge devices for structural planning and introduces auxiliary lightweight models to generate segment-specific key-value (KV) caches. These minimal inference artifacts enable collaborative parallel decoding across IoT devices without requiring full model instantiation. The framework employs static-dynamic KV cache separation to minimize communication overhead while maintaining semantic coherence through structure-guided coordination. Extensive evaluation on realistic edge testbeds demonstrates significant performance improvements. Under diverse deployment settings,E$^{2}$LLMachieves 74%–87.7% end-to-end latency reduction compared with several state-of-the-art baselines, while maintaining comparable generation quality; meanwhile, it also delivers a 34.6%–72.2% reduction in communication overhead, improves 9-12 × in energy efficiency. The framework exhibits strong scalability under bandwidth-limited conditions, enabling efficient LLM deployment across heterogeneous edge-IoT environments. Xingyu Feng 0001, Huanqi Yang, Zhuangzhuang Chen, Chengwen Luo 0001, Zhangbing Zhou, Weitao Xu, Victor C. M. Leung |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Attack-inspired Calibration Loss for Calibrating Crack RecognitionabstractDeep neural networks (DNNs) have substantially achieved high predictive accuracy in many vision tasks. However, we find that they are poorly calibrated for crack recognition tasks, as these DNNs tend to produce both under-confident and over-confident predictions in such safety-critical applications, thereby limiting their practical use in real-world scenarios. To address this issue, we propose a novel attack-inspired calibration loss (AICL) that explicitly regularizes class probabilities to be better confidence estimation. Specifically, we first propose the attack-inspired correctness estimation method (ACE) that aims to estimate the correctness degree of each sample via adversarial attacks. Then, we propose Correctness-aware Distribution Guidance, which starts from a distribution perspective that enforces the ordinal ranking of the predicted confidence referring to the estimated correctness degree. The proposed method can be conveniently implemented on top of any DNNs-based crack recognition model by serving as a plug-and-play loss function. To address the limited availability of related benchmarks, we collect a fully annotated dataset, namely, Bridge2024, which involves inconsistent cracks and noisy backgrounds in real-world bridges. Our AICL outperforms the state-of-art calibration methods on various benchmark datasets including CRACK2019, SDNET2018, and our BRIDGE2024. Zhuangzhuang Chen, Qiangyu Chen, Zhiliang Lin, Xingyu Feng 0001, Jie Chen 0027, Jianqiang Li 0001 |
AAAI | 1 |
| 2025 | Anatomical Knowledge Mining and Matching for Semi-supervised Medical Multi-structure DetectionabstractIn medical image analysis, detecting multiple structures is crucial for evaluations and diagnosis but is often limited by the lack of high-quality annotations. Semi-supervised object detection emerges as a potent methodology to enhance model performance and generalization by leveraging a vast pool of unlabeled data alongside a minimal set of labeled data. A striking observation is that both unlabelled and labeled medical images contain a priori anatomical knowledge from human screening. In this work, we introduce a novel semi-supervised approach named Semi-akmm for mining and matching anatomical knowledge in ultrasound images. We develop an Adaptive Prior Knowledge Transfer (APKT) module to mine and explore the distribution and knowledge of potential proposal boxes by proposal proportion constraint. Furthermore, within a teacher-student learning framework, we put forward an Anatomical Structure Matching (ASM) module to facilitate co-learning consistent topological prior knowledge between the student and teacher models. To our knowledge, this marks the inception of an efficient semi-supervised medical multi-structure detection model. Our experiments across five publicly available ultrasound datasets demonstrate that Semi-akmm sets a new benchmark in performance with solid results that outperform existing methods. Bin Pu, Liwen Wang 0002, Jiewen Yang, Xingbo Dong, Benteng Ma, Zhuangzhuang Chen, Lei Zhao 0013, Shengli Li 0001, Kenli Li 0001 |
AAAI | 6 |
| 2025 | MuTri: Multi-view Tri-alignment for OCT to OCTA 3D Image TranslationabstractOptical coherence tomography angiography (OCTA) shows its great importance in imaging microvascular networks by providing accurate 3D imaging of blood vessels, but it relies upon specialized sensors and expensive devices. For this reason, previous works show the potential to translate the readily available 3D Optical Coherence Tomography (OCT) images into 3D OCTA images. However, existing OCTA translation methods directly learn the mapping from the OCT domain to the OCTA domain in continuous and infinite space with guidance from only a single view, i.e., the OCTA project map, resulting in suboptimal results. To this end, we propose the multi-view Tri-alignment framework for OCT to OCTA 3D image translation in discrete and finite space, named MuTri. In the first stage, we pre-train two vector-quantized variational auto-encoder (VQ-VAE) by reconstructing 3D OCT and 3D OCTA data, providing semantic prior for subsequent multi-view guidances. In the second stage, our multi-view tri-alignment facilitates another VQVAE model to learn the mapping from the OCT domain to the OCTA domain in discrete and finite space. Specifically, a contrastive-inspired semantic alignment is proposed to maximize the mutual information with the pre-trained models from OCT and OCTA views, to facilitate codebook learning. Meanwhile, a vessel structure alignment is proposed to minimize the structure discrepancy with the pre-trained models from the OCTA project map view, benefiting from learning the detailed vessel structure information. We also collect the first large-scale dataset, namely, OCTA2024, which contains a pair of OCT and OCTA volumes from 846 subjects. Our codes and datasets are available at: https://github.com/xmed-lab/MuTri. Zhuangzhuang Chen, Hualiang Wang, Chubin Ou, Xiaomeng Li 0001 |
CVPR | 1 |
| 2025 | ShiftwiseConv: Small Convolutional Kernel with Large Kernel EffectabstractLarge kernels make standard convolutional neural networks (CNNs) great again over transformer architectures in various vision tasks. Nonetheless, recent studies meticulously designed around increasing kernel size have shown diminishing returns or stagnation in performance. Thus, the hidden factors of large kernel convolution that affect model performance remain unexplored. In this paper, we reveal that the key hidden factors of large kernels can be summarized as two separate components: extracting features at a certain granularity and fusing features by multiple pathways. To this end, we leverage the multi-path long-distance sparse dependency relationship to enhance feature utilization via the proposed Shiftwise (SW) convolution operator with a pure CNN architecture. In a wide range of vision tasks such as classification, segmentation, and detection, SW surpasses state-of-the-art transformers and CNN architectures, including SLaK and UniRepLKNet. More importantly, our experiments demonstrate that 3×3 convolutions can replace large convolutions in existing large kernel CNNs to achieve comparable effects, which may inspire follow-up works. Code and all the models at https://github.com/lidc54/shift-wiseConv. Dachong Li, Zhuangzhuang Chen, Jianqiang Li 0001 |
CVPR | 3 |
| 2025 | EA-KD: Entropy-Based Adaptive Knowledge DistillationabstractKnowledge distillation (KD) enables a smaller 'student' model to mimic a larger 'teacher' model by transferring knowledge from the teacher's output or features. However, most KD methods treat all samples uniformly, overlooking the varying learning value of each sample and thereby limiting effectiveness. In this paper, we propose Entropy- based Adaptive Knowledge Distillation (EA-KD), a simple yet effective plug-and-play KD method that prioritizes learning from valuable samples. EA-KD quantifies each sample's learning value by strategically combining the entropy of the teacher and student output, then dynamically reweights the distillation loss to place greater emphasis on high-entropy samples. Extensive experiments across diverse KD frameworks and tasks-including image classification, object detection, and large language model (LLM) distillation-demonstrate that EA-KD consistently enhances performance, achieving state-of-the-art results with negligible computational cost. Our code is available at https://github.com/cpsu00/EA-KD. Chi-Ping Su, Ching-Hsun Tseng, Bin Pu, Lei Zhao 0013, Jiewen Yang, Zhuangzhuang Chen, Shin-Jye Lee |
ICCV | 6 |
| 2025 | KI-GCNN: Knowledge-Informed Graph Convolutional Neural Network for Multiclass Trajectory Prediction at Signalized Intersections
Zhuangzhuang Chen |
QRS | 1 |
| 2025 | Self-Adaptive Fourier Augmentation Framework for Crack Segmentation in Industrial ScenariosabstractCrack segmentation receives extensive attention in structure health monitoring for many industrial scenarios, e.g., bridges, highways, and nuclear power plants. The current deep learning-based crack segmentation models enjoy the ability to extract discriminative crack features by training with an extensive labeled crack dataset. However, collecting extensive crack samples with accurate annotations from experts for a new scenario is labor-intensive, thereby limiting the effectiveness of these deep models in practical applications. To address this problem, the existing Fourier-based augmentation adopts a vanilla amplitude fusion process, i.e., the portion of amplitude components is fixed or randomly selected, failing to guarantee augmented samples’ semantics consistency, and diversity concerning the original sample. To fill this, this article proposes a self-adaptive Fourier augmentation framework that efficiently synthesizes diverse crack samples for training crack segmentation models. Our proposed framework advances Fourier transformation in an adversarial learning manner, alternating between self-adaptive Fourier-based data augmentation and teacher–student learning. The former aims to guarantee the diversity and semantics consistency of Fourier-based augmented samples, while the latter progressively updates the student network by observing these augmented samples for extracting discriminative features via a knowledge distillation mechanism. It is worth noting that the proposed method is only applied in the training stage without extra computation and memory during inference. Extensive experiments demonstrate the superiority of our method over the existing methods. Zhuangzhuang Chen, Tao Hu 0023, Chengqi Xu, Jie Chen 0027, Houbing Song, Li Wang 0093, Jianqiang Li 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | MambaSAM: A Visual Mamba-Adapted SAM Framework for Medical Image SegmentationabstractThe Segment Anything Model (SAM) has shown exceptional versatility in segmentation tasks across various natural image scenarios. However, its application to medical image segmentation poses significant challenges due to the intricate anatomical details and domain-specific characteristics inherent in medical images. To address these challenges, we propose a novel VMamba adapter framework that integrates a lightweight, trainable Visual Mamba (VMamba) branch with the pre-trained SAM ViT encoder. The VMamba adapter accurately captures multi-scale contextual correlations, integrates global and local information, and reduces ambiguities arising from local features only. Specifically, we propose a novel cross-branch attention (CBA) mechanism to facilitate effective interaction between the SAM and VMamba branches. This mechanism enables the model to learn and adapt more efficiently to the nuances of medical images, extracting rich, complementary features that enhance its representational capacity. Beyond architectural enhancements, we streamline the segmentation workflow by eliminating the need for prompt-driven input mechanisms. This results in an autonomous prediction model that reduces manual input requirements and improves operational efficiency. In addition, our method introduces only minimal additional trainable parameters, offering an efficient solution for medical image segmentation. Extensive evaluations of four medical image datasets demonstrate that our VMamba adapter framework achieves state-of-the-art performance. Specifically, on the ACDC dataset with limited training data, our method achieves an average Dice coefficient improvement of 0.18 and reduces the Hausdorff distance by 20.38 mm compared to the AutoSAM. Pengchen Liang, Leijun Shi, Bin Pu, Renkai Wu, Jianguo Chen 0001, Lite Xu, Zhuangzhuang Chen, Qing Chang 0004 |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | LLM-CoSen: Revisiting Collaborative Sensing With Large Language Models (LLMs)abstractCollaborative sensing has emerged as a novel sensing paradigm, entailing multi-sensor data sharing and multimodal modeling to collaboratively understand sensing behaviors. However, current solutions, i.e., data-level and decision-level fusion methods, fall short of generality, expert knowledge, and holistic/chronic perspective. In this paper, we proposeLLMCoSento revisit collaborative sensing with Large Language Models (LLMs). Specifically,LLM-CoSendesigns a semantic-level fusion approach for inference results for collaborative sensing. Such an approach is characterized by its generality, making it applicable to any heterogeneous devices, and its expert knowledge incorporation, which provides chronic, holistic, and insightful perspectives on the inference results. Regarding inference absence challenges, we propose a personalized model design method to constrain inference time, and a voting-based two-pass prompt engineering strategy for token completion. Regarding inference error challenges, we propose an accuracy restoration strategy for personalized models, and a two-level error estimator coupled with self-correction. Experimental results of human digital system use case on four corresponding benchmark datasets showLLM-CoSencan decrease inference absence by 72.83% and inference errors by 7.65% on average. Xingyu Feng 0001, Zehua Sun, Zhuangzhuang Chen, Chengwen Luo 0001, Zhangbing Zhou, Victor C. M. Leung, Weitao Xu |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Leveraging Segment Anything Model for Source-Free Domain Adaptation via Dual Feature Guided Auto-PromptingabstractSource-free domain adaptation (SFDA) for segmentation aims at adapting a model trained in the source domain to perform well in the target domain with only the source model and unlabeled target data. Inspired by the recent success of Segment Anything Model (SAM) which exhibits the generality of segmenting images of various modalities and in different domains given human-annotated prompts like bounding boxes or points, we for the first time explore the potentials of Segment Anything Model for SFDA via automatedly finding an accurate bounding box prompt. We find that the bounding boxes directly generated with existing SFDA approaches are defective due to the domain gap. To tackle this issue, we propose a novel Dual Feature Guided (DFG) auto-prompting approach to search for the box prompt. Specifically, the source model is first trained in a feature aggregation phase, which not only preliminarily adapts the source model to the target domain but also builds a feature distribution well-prepared for box prompt search. In the second phase, based on two feature distribution observations, we gradually expand the box prompt with the guidance of the target model feature and the SAM feature to handle the class-wise clustered target features and the class-wise dispersed target features, respectively. To remove the potentially enlarged false positive regions caused by the over-confident prediction of the target model, the refined pseudo-labels produced by SAM are further postprocessed based on connectivity analysis. Experiments on 3D and 2D datasets indicate that our approach yields superior performance compared to conventional methods. Code is available at https://github.com/xmed-lab/DFG. Zheang Huai, Yi Li 0050, Zhuangzhuang Chen, Xiaomeng Li 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Mind marginal non-crack regions: Clustering-inspired representation learning for crack segmentationabstractCrack segmentation datasets make great efforts to ob-tain the ground truth crack or non-crack labels as clearly as possible. However, it can be observed that ambiguities are still inevitable when considering the marginal non-crack re-gion, due to low contrast and heterogeneous texture. To solve this problem, we propose a novel clustering-inspired representation learning framework, which contains a two-phase strategy for automatic crack segmentation. In the first phase, a pre-process is proposed to localize the marginal non-crack region. Then, we propose an ambiguity-aware segmentation loss (Aseg Loss) that enables crack segmentation models to capture ambiguities in the above regions via learning segmentation variance, which allows us to further localize ambiguous regions. In the second phase, to learn the discriminative features of the above regions, we propose a clustering-inspired loss (CI Loss) that alters the supervision learning of these regions into an unsupervised clus-tering manner. We demonstrate that the proposed method could surpass the existing crack segmentation models on various datasets and our constructed CrackSeg5k dataset. Zhuangzhuang Chen, Zhuonan Lai, Jie Chen 0027, Jianqiang Li 0001 |
CVPR | 1 |
| 2024 | Implicit Gradient-Modulated Semantic Data Augmentation for Deep Crack RecognitionabstractCrack detection has attracted extensive attention in an intelligent transportation system (ITS). Despite the substantial progress of deep learning technology on crack recognition tasks, due to the various limitations in traffic, equipment, and time, it is hard to collect copious samples for training deep models. Considering this, implicitly semantic data augmentation (ISDA) tries to augment the training set in the feature space. However, when applying it to crack recognition tasks, our empirical studies reveal that those poor-classified augmented samples have little semantic relevance to the crack class, resulting in a non-negligible negative effect on training deep models. Since the augmented features follow the multivariate normal distribution, it is computationally inefficient to explicitly sample those features and filter out the hard-classified augmented features. To this end, we propose the implicit gradient-modulated semantic data augmentation (IGMSDA) for addressing the above problems. Concretely, this paper first proposes gradient-modulated (GM) loss to dynamically modulate the gradient of those poor-classified augmented samples by reshaping the standard cross-entropy loss. And then, in the feature space, we derive an upper bound of the expected GM loss on the augmented training set to avoid the costly explicit sampling process. Experiments show that IGMSDA improves the generalization performance of the existing deep models on crack recognition datasets. Zhuangzhuang Chen, Ronghao Lu, Jie Chen 0027, Houbing Song, Jianqiang Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | The Devil is in the Crack Orientation: A New Perspective for Crack DetectionabstractCracks are usually curve-like structures that are the focus of many computer-vision applications (e.g., road safety inspection and surface inspection of the industrial facilities). The existing pixel-based crack segmentation methods rely on time-consuming and costly pixel-level annotations. And the object-based crack detection methods exploit the horizontal box to detect the crack without considering crack orientation, resulting in scale variation and intra-class variation. Considering this, we provide a new perspective for crack detection that models the cracks as a series of sub-cracks with the corresponding orientation. However, the vanilla adaptation of the existing oriented object detection methods to the crack detection tasks will result in limited performance, due to the boundary discontinuity issue and the ambiguities in sub-crack orientation. In this paper, we propose a first-of-its-kind oriented sub-crack detector, dubbed as CrackDet, which is derived from a novel piecewise angle definition, to ease the boundary discontinuity problem. And then, we propose a multi-branch angle regression loss for learning sub-crack orientation and variance together. Since there are no related benchmarks, we construct three fully annotated datasets, namely, ORC, ONPP, and OCCSD, which involve various cracks in road pavement and industrial facilities. Experiments show that our approach outperforms state-of-the-art crack detectors. Zhuangzhuang Chen, Jin Zhang 0013, Zhuonan Lai, Guanming Zhu, Zun Liu, Jie Chen 0027, Jianqiang Li 0001 |
ICCV | 1 |
| 2022 | Geometry-Aware Guided Loss for Deep Crack RecognitionabstractDespite the substantial progress of deep models for crack recognition, due to the inconsistent cracks in varying sizes, shapes, and noisy background textures, there still lacks the discriminative power of the deeply learned features when supervised by the cross-entropy loss. In this paper, we propose the geometry-aware guided loss (GAGL) that enhances the discrimination ability and is only applied in the training stage without extra computation and memory during inference. The GAGL consists of the feature-based geometry-aware projected gradient descent method (FGA-PGD) that approximates the geometric distances of the features to the class boundaries, and the geometry-aware update rule that learns an anchor of each class as the approximation of the feature expected to have the largest geometric distance to the corresponding class boundary. Then the discriminative power can be enhanced by minimizing the distances between the features and their corresponding class anchors in the feature space. To address the limited availability of related benchmarks, we collect a fully annotated dataset, namely, NPP2021, which involves inconsistent cracks and noisy backgrounds in real-world nuclear power plants. Our proposed GAGL outperforms the state of the arts on various benchmark datasets including CRACK2019, SDNET2018, and our NPP2021. Zhuangzhuang Chen, Jin Zhang 0013, Zhuonan Lai, Jie Chen 0027, Zun Liu, Jianqiang Li 0001 |
CVPR | 1 |
| 2022 | When Active Learning Meets Implicit Semantic Data Augmentation
Zhuangzhuang Chen, Jin Zhang 0013, Jie Chen 0027, Jianqiang Li 0001 |
ECCV (25) | 1 |
| 2022 | Integrated Air-Ground Vehicles for UAV Emergency Landing Based on Graph Convolution NetworkabstractWith unmanned aerial vehicle (UAV) technologies advanced rapidly, many applications have emerged in cities. However, those applications do not widely spread as the safety consideration hinders the UAV from integrating into the civilian environment. This work focuses on investigating the UAV emergency landing problem which is a critical safety functionality of UAV. This work proposed a graph convolution network (GCN)-based decision network to learn by imitating the human pilots’ landing strategy. To alleviate the needs of a large amount of real-world data for model training, the proposed model allows to be trained in a simulated environment and then transferred to the real-world scenario due to the separation of domain-specific terrain classes and domain-independent topological structures among down-looking camera images. The GCN-based decision network can be coupled with a topological heuristic to improve the performance of action prediction in an emergency situation. To evaluate the proposed method, this work implemented a simulation environment for collecting data and testing the UAV emergency landing. The empirical results in both simulated and real-world scenarios show that the proposed methods can outperform the state-of-the-art counterparts in terms of predictive accuracy and success landing rate. Jie Chen 0027, Jianqiang Li 0001, Weiming Du, Zhuangzhuang Chen, Zun Liu, Huihui Wang 0001, Victor C. M. Leung |
IEEE Internet Things J. | 5 |
| 2021 | Diversity-Sensitive Generative Adversarial Network for Terrain Mapping Under Limited Human InterventionabstractIn a collaborative air-ground robotic system, the large-scale terrain mapping using aerial images is important for the ground robot to plan a globally optimal path. However, it is a challenging task in a novel and dynamic field without historical human supervision. To alleviate the reliance on human intervention, this article presents a novel framework that integrates active learning and generative adversarial networks (GANs) to effectively exploit small human-labeled data for terrain mapping. In order to model the diverse terrain patterns, this article designs two novel diversity-sensitive GAN models which can capture fine-grained terrain classes among aerial image patches. The proposed approaches are tested in two real-world scenarios using our collaborative air-ground robotic platform. The empirical results show that our methods can outperform their counterparts in the predictive accuracy of terrain classification, visual quality of terrain mapping, and average length of the planned ground path. In practice, the proposed terrain mapping framework is especially valuable when the budget in time or labor cost is very limited. Jianqiang Li 0001, Zhuangzhuang Chen, Jie Chen 0027, Qiuzhen Lin |
IEEE Trans. Cybern. | 2 |
| 2020 | A novel edge-enabled SLAM solution using projected depth image information
Jianqiang Li 0001, Zhuangzhuang Chen, Jia Wang 0008, Chengwen Luo 0001, Huihui Wang 0001 |
Neural Comput. Appl. | 3 |
| 2020 | Using Weighted Extreme Learning Machine Combined With Scale-Invariant Feature Transform to Predict Protein-Protein Interactions From Protein Evolutionary InformationabstractProtein-Protein Interactions (PPIs) play an irreplaceable role in biological activities of organisms. Although many high-throughput methods are used to identify PPIs from different kinds of organisms, they have some shortcomings, such as high cost and time-consuming. To solve the above problems, computational methods are developed to predict PPIs. Thus, in this paper, we present a method to predict PPIs using protein sequences. First, protein sequences are transformed into Position Weight Matrix (PWM), in which Scale-Invariant Feature Transform (SIFT) algorithm is used to extract features. Then Principal Component Analysis (PCA) is applied to reduce the dimension of features. At last, Weighted Extreme Learning Machine (WELM) classifier is employed to predict PPIs and a series of evaluation results are obtained. In our method, since SIFT and WELM are used to extract features and classify respectively, we called the proposed method SIFT-WELM. When applying the proposed method on three well-known PPIs datasets of Yeast, Human and Helicobacter.pylori, the average accuracies of our method using five-fold cross validation are obtained as high as 94.83, 97.60 and 83.64 percent, respectively. In order to evaluate the proposed approach properly, we compare it with Support Vector Machine (SVM) classifier and other recent-developed methods in different aspects. Moreover, the training time of our method is greatly shortened, which is obviously superior to the previous methods, such as SVM, ACC, PCVMZM and so on. Jianqiang Li 0001, Zhu-Hong You, Zhuangzhuang Chen, Qiuzhen Lin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2019 | Automatic Classification of Fetal Heart Rate Based on Convolutional Neural NetworkabstractFetal heart rate (FHR) is very significant to evaluate the status of fetus. However, based on traditional classification criteria is not accurate. With the rapid development of computer information technology, computer technology is vital for the analysis of FHR in electronic fetal monitoring (EFM). FHR is divided into three classes as: 1) normal; 2) suspicious; and 3) abnormal. Through the cooperation with the hospital, we got 4473 records, including 3012 normal, 1024 suspicious, 437 abnormal records by our EFM system. In order to improve the accuracy of fetal status assessment, high 1-D FHR records are divided into ten d-window segments, and then use convolutional neural network (CNN) to process the data in parallel. Finally, we use the voting method to determine the class of FHR records. We also made a comparative experiment, the feature extraction method based on basic statistics is used to extract the features of FHR. And then the features were applied as the input to support vector machine (SVM) and multilayer perceptron (MLP) to classify. According to the results of the experiment, the accuracy of classification of SVM, MLP, and CNN are 79.66%, 85.98%, and 93.24%, respectively. Jianqiang Li 0001, Zhuangzhuang Chen, Luxiang Huang, Xianghua Fu, Huihui Wang 0001, Qingguo Zhao |
IEEE Internet Things J. | 2 |
| 2018 | Using Weighted Extreme Learning Machine Combined with Scale-Invariant Feature Transform to Predict Protein-Protein Interactions from Protein Evolutionary Information
Jianqiang Li 0001, Zhu-Hong You, Zhuangzhuang Chen, Qiuzhen Lin |
ICIC (1) | 4 |