Kang Dang

dblp:141/9958 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0003-0613-2787ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 CNText2Sign and CNSign: Unified Chinese Sign Language Datasets for Bidirectional Accessibility
abstract
Sign language is the primary communication mode for 72 million hearing-impaired individuals worldwide, necessitating effective bidirectional Sign Language Production and Sign Language Translation systems. However, functional bidirectional systems require a unified linguistic environment, hindered by the lack of suitable unified datasets, particularly those providing the necessary pose information for accurate Sign Language Production (SLP) evaluation. Concurrently, current SLP evaluation methods like back-translation ignore pose accuracy, and high-quality coordinated generation remains challenging. To create this crucial environment and overcome these challenges, we introduce CNText2Sign and CNSign, which together constitute the first unified dataset aimed at supporting bidirectional accessibility systems for Chinese sign language; CNText2Sign provides 15,000 natural language-to-sign mappings and standardized skeletal keypoints for 8,643 vocabulary items supporting pose assessment. Building upon this foundation, we propose the AuraLLM model, which leverages a decoupled architecture with CNText2Sign's pose data for novel direct gesture accuracy assessment. The model employs retrieval augmentation and Cascading Vocabulary Resolution to handle semantic mapping and out-of-vocabulary words, and achieves all-scenario production with controllable coordination of gestures and facial expressions via pose-conditioned video synthesis. Concurrently, our Sign Language Translation model SignMST-C employs targeted self-supervised pretraining for dynamic feature capture, achieving new SOTA results on PHOENIX2014-T with BLEU-4 scores up to 32.08. AuraLLM establishes a strong performance baseline on CNText2Sign with a BLEU-4 score of 50.41 under direct evaluation.
Yulong Li 0002, Zhixiang Lu, Haochen Xue, Jianghao Wu 0001, Mian Zhou, Kang Dang, Yifang Wang 0006, Muhammad Imran Razzak, Jionglong Su
KDD (1)10
2026 HolistAno: Retinal anomaly detection with holistic feature modeling
abstract
Early detection of retinal lesions is critical for preventing vision loss. While supervised learning has shown promise, existing methods often rely on extensive labeled data, which is costly and difficult to obtain in medical applications. Unsupervised anomaly detection provides an attractive alternative by requiring only healthy retinal images and no abnormal annotations. However, current methods face significant challenges in modeling the complex structures of normal retinal anatomy, learning discriminative features for detecting subtle lesions, and capturing multi-scale features to handle anomalies of varying sizes – highlighting the need for holistic feature modeling that comprehensively represents both retinal anatomy and pathology. To address these challenges, we propose HolistAno, a novel unsupervised anomaly detection framework with holistic retinal modeling. HolistAno adopts a two-stage network architecture, incorporating a novel anomaly generator and a Balanced Mamba Scale Fusion (BMSF) module to effectively learn comprehensive retinal feature representations. This enables accurate detection of subtle lesions, diverse lesion types, and anomalies across multiple scales. Extensive experiments on five benchmark datasets demonstrate that HolistAno achieves state-of-the-art performance in both anomaly classification and localization tasks, with superior generalization and robustness across multiple datasets and cross-dataset scenarios compared to existing methods.
Jingqi Niu, Kang Dang, Nan Xi, Junsong Yuan 0001, Yanjing Liu, Mian Zhou, Jionglong Su
Expert Syst. Appl.2
2026 Enhancing decision boundaries in continual learning through a decoupled Gaussian framework
abstract
The goal of continual learning (CL) is to acquire new knowledge while retaining previously learned information. CNN-based and prompt-based CL methods have achieved remarkable progress in recent years. However, most prior work has primarily focused on reducing forgetting from the perspective of the model itself. In this paper, we investigate CL from the perspective of decision boundaries, analyzing the impact of instance-level feature overlap. To address this issue, we propose a generic Decoupled Gaussian Softmax Classifier that enhances class discriminability during CL process. Specifically, we decouple the features extracted by the backbone into multiple Gaussian distributions, which are directly fused into the feature space through weighted integration. A regularization term is introduced to penalize the overlap of similar features, while an adaptive decision boundary is assigned to each class to encourage inter-class separation and intra-class compactness. Experiments on 4 widely used continual learning datasets and 12 CL scenarios show that our method has good plug-and-play capability. It improves the average accuracy by 1%–2.63% over the baseline models, while effectively reducing both the forgetting rate and the Expected Calibration Error. Our code is available at: https://anonymous.4open.science/r/DGSC-main-310D .
Zhikun Feng, Liu Yu 0001, Ping Kuang, Mian Zhou, Kang Dang, Yakun Ju
Inf. Process. Manag.7
2026 TIPS: Two-level prompt selection for more stability-plasticity balance in continual learning
Zhikun Feng, Kang Dang, Mian Zhou, Ping Kuang, Mingyu Wu 0011, Liu Yu 0001, Jionglong Su
Pattern Recognit.3
2025 KD-MSLRT: Lightweight Sign Language Recognition Model Based on Mediapipe and 3D to 1D Knowledge Distillation
abstract
Artificial intelligence has achieved notable results in sign language recognition and translation. However, relatively few efforts have been made to significantly improve the quality of life for the 72 million hearing-impaired people worldwide. Sign language translation models, relying on video inputs, involves with large parameter sizes, making it time-consuming and computationally intensive to be deployed. This directly contributes to the scarcity of human-centered technology in this field. Additionally, the lack of datasets in sign language translation hampers research progress in this area. To address these, we first propose a cross-modal multi-knowledge distillation technique from 3D to 1D and a novel end-to-end pre-training text correction framework. Compared to other pre-trained models, our framework achieves significant advancements in correcting text output errors. Our model achieves a decrease in Word Error Rate (WER) of at least 1.4% on PHOENIX14 and PHOENIX14T datasets compared to the state-of-the-art CorrNet. Additionally, the TensorFlow Lite (TFLite) quantized model size is reduced to 12.93 MB, making it the smallest, fastest, and most accurate model to date. We have also collected and released extensive Chinese sign language datasets, and developed a specialized training vocabulary. To address the lack of research on data augmentation for landmark data, we have designed comparative experiments on various augmentation methods. Moreover, we performed a simulated deployment and prediction of our model on Intel platform CPUs and assessed the feasibility of deploying the model on other platforms.
Yulong Li 0002, Bolin Ren, Changyuan Liu, Zhengyong Jiang, Kang Dang, Jionglong Su
AAAI6
2025 Modality-Aware Shot Relating and Comparing for Video Scene Detection
abstract
Video scene detection involves assessing whether each shot and its surroundings belong to the same scene. Achieving this requires meticulously correlating multi-modal cues, e.g., visual entity and place modalities, among shots and comparing semantic changes around each shot. However, most methods treat multi-modal semantics equally and do not examine contextual differences between the two sides of a shot, leading to sub-optimal detection performance. In this paper, we propose the Modality-Aware Shot Relating and Comparing approach (MASRC), which enables relating shots per their own characteristics of visual entity and place modalities, as well as comparing multi-shots similarities to have scene changes explicitly encoded. Specifically, to fully harness the potential of visual entity and place modalities in modeling shot relations, we mine long-term shot correlations from entity semantics while simultaneously revealing short-term shot correlations from place semantics. In this way, we can learn distinctive shot features that consolidate coherence within scenes and amplify distinguishability across scenes. Once equipped with distinctive shot features, we further encode the relations between preceding and succeeding shots of each target shot by similarity convolution, aiding in the identification of scene ending shots. We validate the broad applicability of the proposed components in MASRC. Extensive experimental results on public benchmark datasets demonstrate that the proposed MASRC significantly advances video scene detection.
Jiawei Tan, Hongxing Wang 0001, Kang Dang, Zhilong Ou
AAAI3
2025 Decoding the Flow: CauseMotion for Emotional Causality Analysis in Long-form Conversations
abstract
Long-sequence causal reasoning seeks to uncover causal relationships within extended time series data but is hindered by complex dependencies and the challenges of validating causal links. To address the limitations of large-scale language models (e.g., GPT-4) in capturing intricate emotional causality within extended dialogues, we propose CauseMotion, an innovative framework combining emotional causal dynamic mapping with multimodal feature fusion. CauseMotion implements dynamic mapping through a sliding window mechanism and fusion strategies, while integrating audio features—vocal emotion, intensity, and speech rate—to enrich semantic representations. This design enables efficient retrieval of contextually relevant information and precise inference of emotional causal chains spanning multiple conversational turns. We constructed the first benchmark dataset for long-sequence emotional causal reasoning, featuring dialogues with over 70 turns. Experimental results show that CauseMotion significantly enhances emotional understanding and causal inference capabilities in large language models. A GLM-4 integrated with CauseMotion achieves an 8.7% improvement in causal accuracy over the original model and surpasses GPT-4o by 1.2%. On the DiaASQ dataset, CauseMotion-GLM-4 achieves state-of-the-art results in accuracy, F1 score, and causal reasoning accuracy.
Yulong Li 0002, Zichen Yu, Zhixiang Lu, Haochen Xue, Zhaodong Wu, Kang Dang, Muhammad Imran Razzak, Jionglong Su
AVSS9
2025 PathVLG: A Vision-Language Framework for Domain Generalization in Cross-Organ Adenocarcinoma Segmentation
abstract
Domain shifts in computational pathology, caused by variations in staining, imaging devices, and tissue morphology, challenge model performance in segmentation tasks. Existing domain generalization methods, such as style transfer and feature alignment, often fail to account for organ-level morphological differences. In this paper, we propose PathVLG, a visionlanguage model designed to improve domain generalization for adenocarcinoma segmentation. PathVLG leverages a CONCHbased encoder with three key innovations: the Text-informed Content Query Reformer (TCQR), Text-driven Style Augmentor (TSA), and Style Regeneration Decoder (SRD). These components help the model adapt across domains by incorporating text embeddings, generating diverse styles, and combining source and target domain features. Experimental results show that PathVLG outperforms existing methods in cross-domain generalization.
Biwen Meng, Wanrong Yang, Kang Dang, Yalin Zheng, Jingxin Liu 0005
BIBM5
2025 Anchor-Aware Similarity Cohesion in Target Frames Enables Predicting Temporal Moment Boundaries in 2D
abstract
Video moment retrieval aims to locate specific moments from a video according to the query text. This task presents two main challenges: i) aligning the query and video frames at the feature level, and ii) projecting the query-aligned frame features to the start and end boundaries of the matching interval. Previous work commonly involves all frames in feature alignment, easy to cause aligning irrelevant frames with the query. Furthermore, they forcibly map visual features to interval boundaries but ignoring the information gap between them, yielding suboptimal performance. In this study, to reduce distraction from irrelevant frames, we designate an anchor frame as that with the maximum query-frame relevance measured by the established Vision-Language Model. Via similarity comparison between the anchor frame and the others, we produce a semantically compact segment around the anchor frame, which serves as a guide to align features of query and related frames. We observe that such a feature alignment will make similarity cohesive between target frames, which enables us to predict the interval boundaries by a single point detection in the 2D semantic similarity space of frames, thus well bridging the information gap between frame semantics and temporal boundaries. Experimental results across various datasets demonstrate that our approach significantly improves the alignment between queries and video frames while effectively predicting temporal moment boundaries. Especially, on QVHighlights Test and ActivityNet Captions datasets, our proposed approach achieves 3.8% and 7.4% respectively higher than current state-of-the-art [email protected] performance. The code is available at https://github.com/ExMorgan-Alter/AFAFSGD.
Jiawei Tan, Hongxing Wang 0001, Junwu Weng, Zhilong Ou, Kang Dang
CVPR6
2025 Decoupling Overlapped Feature Spaces: When Continual Learning Meets Fine-Grain Classification
abstract
The goal of Class Incremental Learning (CIL) is to continuously learn new classes while preventing forgetting of old ones. Most previous works focused on reducing catastrophic forgetting from model’s perspective. However, the model is not the only factor contributing to forgetting. In this paper, we take the perspective of class instances and find that fine-grained class increments can lead to feature overlap between classes, further reducing instance margins. We call this interesting phenomenon as Fine-grained class confusion effect in CIL. Since preserving instance margins is crucial for resisting forgetting, it is beneficial to maintain the margin amount as much as possible. To achieve this, we propose a general Gaussian decoupling classifier to enhance the discriminability of similar classes during incremental learning. Specifically, we decouple the features of different classes extracted by the backbone network into multiple independent Gaussian distributions. By directly integrating them into the features with weighted fusion, we introduce a regularization penalty that encourages minimizing the overlap of similar features, thus increasing the feature distance between classes. Extensive experiments show that our method effectively improves class separation and better preserves instance margins, ultimately alleviating forgetting. The improved model achieves better performance on CUB-200 and CARS-196.
Zhikun Feng, Mingyu Wu 0011, Ping Kuang, Kang Dang, Mian Zhou, Liu Yu 0001
ICME4
2025 EndoFlow-SLAM: Real-Time Endoscopic SLAM with Flow-Constrained Gaussian Splatting
Taoyu Wu, Yiyi Miao, Zhuoxiao Li, Haocheng Zhao, Kang Dang, Jionglong Su, Limin Yu, Haoang Li
MICCAI (9)5
2025 MSWAL: 3D Multi-class Segmentation of Whole Abdominal Lesions Dataset
Zhaodong Wu, Qiaochu Zhao, Yulong Li 0002, Haochen Xue, Zhengyong Jiang, Angelos Stefanidis, Muhammad Imran Razzak, ZongYuan Ge, Junjun He, Yu Qiao 0001, Kang Dang, Jionglong Su
MICCAI (2)15
2024 Adaptive knowledge transfer for class incremental learning
Zhikun Feng, Mian Zhou, Angelos Stefanidis, Jionglong Su, Kang Dang, Chuanhui Li
Pattern Recognit. Lett.6
2024 DualStreamFoveaNet: A Dual Stream Fusion Architecture With Anatomical Awareness for Robust Fovea Localization
abstract
Accurate fovea localization is essential for analyzing retinal diseases to prevent irreversible vision loss. While current deep learning-based methods outperform traditional ones, they still face challenges such as the lack of local anatomical landmarks around the fovea, the inability to robustly handle diseased retinal images, and the variations in image conditions. In this paper, we propose a novel transformer-based architecture called DualStreamFoveaNet (DSFN) for multi-cue fusion. This architecture explicitly incorporates long-range connections and global features using retina and vessel distributions for robust fovea localization. We introduce a spatial attention mechanism in the dual-stream encoder to extract and fuse self-learned anatomical information, focusing more on features distributed along blood vessels and significantly reducing computational costs by decreasing token numbers. Our extensive experiments show that the proposed architecture achieves state-of-the-art performance on two public datasets and one large-scale private dataset. Furthermore, we demonstrate that the DSFN is more robust on both normal and diseased retina images and has better generalization capacity in cross-dataset experiments.
Sifan Song, Jinfeng Wang 0008, Zilong Wang 0006, Hongxing Wang 0001, Jionglong Su, Kang Dang
IEEE J. Biomed. Health Informatics7
2023 ReSynthDetect: A Fundus Anomaly Detection Network with Reconstruction and Synthetic Features
Jingqi Niu, Qinji Yu, Shiwen Dong, Zilong Wang 0006, Kang Dang
BMVC5
2023 Source-Free Domain Adaptation for Medical Image Segmentation via Prototype-Anchored Feature Alignment and Contrastive Learning
Qinji Yu, Nan Xi, Junsong Yuan 0001, Kang Dang
MICCAI (7)5
2022 Weakly Supervised Online Action Detection for Infant General Movements
Tongyi Luo, Chuncao Zhang, Siheng Chen, Guangjun Yu, Kang Dang
MICCAI (2)7
2022 ADAM Challenge: Detecting Age-Related Macular Degeneration From Fundus Images
abstract
Age-related macular degeneration (AMD) is the leading cause of visual impairment among elderly in the world. Early detection of AMD is of great importance, as the vision loss caused by this disease is irreversible and permanent. Color fundus photography is the most cost-effective imaging modality to screen for retinal disorders. Cutting edge deep learning based algorithms have been recently developed for automatically detecting AMD from fundus images. However, there are still lack of a comprehensive annotated dataset and standard evaluation benchmarks. To deal with this issue, we set up the Automatic Detection challenge on Age-related Macular degeneration (ADAM), which was held as a satellite event of the ISBI 2020 conference. The ADAM challenge consisted of four tasks which cover the main aspects of detecting and characterizing AMD from fundus images, including detection of AMD, detection and segmentation of optic disc, localization of fovea, and detection and segmentation of lesions. As part of the ADAM challenge, we have released a comprehensive dataset of 1200 fundus images with AMD diagnostic labels, pixel-wise segmentation masks for both optic disc and AMD-related lesions (drusen, exudates, hemorrhages and scars, among others), as well as the coordinates corresponding to the location of the macular fovea. A uniform evaluation framework has been built to make a fair comparison of different models using this dataset. During the ADAM challenge, 610 results were submitted for online evaluation, with 11 teams finally participating in the onsite challenge. This paper introduces the challenge, the dataset and the evaluation methods, as well as summarizes the participating methods and analyzes their results for each task. In particular, we observed that the ensembling strategy and the incorporation of clinical domain knowledge were the key to improve the performance of the deep learning models.
Huihui Fang, Fei Li 0021, Huazhu Fu, Xu Sun 0006, Xingxing Cao, Fengbin Lin, Jaemin Son, Gwenolé Quellec, Sarah Matta, Sharath M. Shankaranarayana, Chuen-heng Wang, Nisarg A. Shah, Chia-Yen Lee, Chih-Chung Hsu, Hai Xie, Bai Ying Lei, Ujjwal Baid, Shubham Innani, Kang Dang, Wenxiu Shi, Ravi Kamble, Nitin Singhal, Ching-Wei Wang, Shih-Chang Lo, José Ignacio Orlando, Hrvoje Bogunovic, Xiulan Zhang, Yanwu Xu 0001
IEEE Trans. Medical Imaging21
2021 Context Model for Pedestrian Intention Prediction Using Factored Latent-Dynamic Conditional Random Fields
abstract
Smooth handling of pedestrian interactions is a key requirement for Autonomous Vehicles (AV) and Advanced Driver Assistance Systems (ADAS). Such systems call for early and accurate prediction of a pedestrian’s crossing/not-crossing behaviour in front of the vehicle. Existing approaches to pedestrian behaviour prediction make use of pedestrian motion, his/her location in a scene and static context variables such as traffic lights, zebra crossings etc. We stress on the necessity of early prediction for smooth operation of such systems. We introduce the influence of vehicle interactions on pedestrian intention for this purpose. In this paper, we show a discernible advance in prediction time aided by the inclusion of such vehicle interaction context. We apply our methods to two different datasets, one in-house collected - NTU dataset and another public real-life benchmark - JAAD dataset. We also propose a generalization of the Latent-Dynamic Conditional Random Fields (LDCRF), called Factored LDCRF (FLDCRF), for improved sequence prediction performance. FLDCRF outperforms Long Short-Term Memory (LSTM) networks across the datasets over identical time-series features. While the existing best system predicts pedestrian stopping behaviour with 70% accuracy 0.38 seconds before the actual events, our system achieves such accuracy at least 0.9 seconds on an average before the actual events across datasets.
Satyajit Neogi, Michael Hoy, Kang Dang, Hang Yu 0002, Justin Dauwels
IEEE Trans. Intell. Transp. Syst.3
2018 Actor-Action Semantic Segmentation with Region Masks
Kang Dang, Chunluan Zhou, Zhigang Tu 0001, Michael Hoy, Justin Dauwels, Junsong Yuan 0001
BMVC1
2017 Real-time hierarchical fusion system for semantic segmentation in offroad scenes
abstract
Semantic segmentation is an important task for autonomous vehicle navigation in off road environments. However, several natural factors make this problem uniquely challenging. For example, road segmentation is often difficult under heavy shadow or steel terrain, and dangerous muddy water puddles may have the similar visual appearance to dirt road surfaces (and thus are hard to identify). To tacule these challenges, we present a semantic segmentation system based on a two-stage hierarchical fusion pipeline. The first stage improves the road segmentation by effectively fusing information from camera and 3D Lidar point cloud. The second stage is dedicated to detecting water puddles, based on the results from the first stage. Due to the parallelized architecture, our system can be deployed for real-time applications. We achieved an F1 score of around 93% for road segmentation and 80% for water puddle segmentation at more than 10 Hz.
Kang Dang, Michael Hoy, Justin Dauwels, Junsong Yuan 0001
FUSION1
2017 Learning location constrained pixel classifiers for image parsing
Kang Dang, Junsong Yuan 0001
J. Vis. Commun. Image Represent.1
2015 Adaptive Exponential Smoothing for Online Filtering of Pixel Prediction Maps
abstract
We propose an efficient online video filtering method, called adaptive exponential filtering (AES) to refine pixel prediction maps. Assuming each pixel is associated with a discriminative prediction score, the proposed AES applies exponentially decreasing weights over time to smooth the prediction score of each pixel, similar to classic exponential smoothing. However, instead of fixing the spatial pixel location to perform temporal filtering, we trace each pixel in the past frames by finding the optimal path that can bring the maximum exponential smoothing score, thus performing adaptive and non-linear filtering. Thanks to the pixel tracing, AES can better address object movements and avoid over-smoothing. To enable real-time filtering, we propose a linear-complexity dynamic programming scheme that can trace all pixels simultaneously. We apply the proposed filtering method to improve both saliency detection maps and scene parsing maps. The comparisons with average and exponential filtering, as well as state-of-the-art methods, validate that our AES can effectively refine the pixel prediction maps, without using the original video again.
Kang Dang, Junsong Yuan 0001
ICCV1
2014 Height Gradient Histogram (HIGH) for 3D Scene Labeling
abstract
RGB-D (color + 3D point cloud) based scene labeling has received much attention due to the affordable RGB-D sensors such as Microsoft Kinect. To fully utilize the RGB-D data, it is critical to develop robust features that can reliably describe the 3D shape information of the point cloud data. Previous work has proposed to extract SIFT-like features from the depth dimension data directly while ignored the important height dimension data of the 3D point cloud. In this paper, we propose to describe 3D scene using height gradient information and propose a new compact point cloud feature called Height Gradient Histogram (HIGH). Using Text on Boost as the pixel classifier, the experiments on two benchmarked 3D scene labeling datasets show that HIGH feature can well handle the intra-category variations of object class, and significantly improve class-average accuracy compared with the state-of-the-art results. We will publish the code of HIGH feature for the community.
Gangqiang Zhao, Junsong Yuan 0001, Kang Dang
3DV3
2014 Location Constrained Pixel Classifiers for Image Parsing with Regular Spatial Layout
Kang Dang, Junsong Yuan 0001
BMVC1
2013 Voxel labelling in CT images with data-driven contextual features
abstract
Spatial contextual information is useful for voxel labelling and especially suitable for the images with relatively fixed scene structure such as CT images. For each voxel, the intensity values of nearby and far away positions are sampled as its contextual features and such contextual features have shown promising performance. However how to determine sampling position to construct good contextual features remains a critical problem since a good sampling could significantly improve the classification performance. In this paper we proposed a novel approach by discovering discriminative sampling pattern. We emphasize that the sampling pattern is not hand craft but data driven and can cater to a particular type of problem, such as kidneys labelling in contrast-enhanced CT images. After discriminative pattern is discovered it can be adapted for use in other datasets of the same problem. Experiments on kidney dataset showed considerable improvements over competing methods.
Kang Dang, Junsong Yuan 0001, Ho Yee Tiong
ICIP1