VLDB 2026 Research / reviewers in the wild / expert
Xun Wang 0007
dblp:82/1331-7
· DBLP profile ↗
92ranked-venue papers
8as first author
46since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 5 first-author · 24 since 2021Artificial intelligence and machine learning · 37 · 3 first-author · 22 since 2021Databases, data management, data science and information retrieval · 11 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | End-to-End Multi-Person Pose Estimation with Pose-Aware Video TransformerabstractExisting multi-person video pose estimation methods typically adopt a two-stage pipeline: detecting individuals in each frame, followed by temporal modeling for single-person pose estimation. This design relies on heuristic operations such as tracking, RoI cropping, and non-maximum suppression, limiting both accuracy and efficiency. In this paper, we present a fully end-to-end framework for multi-person 2D pose estimation in videos, effectively eliminating heuristic operations. A key challenge is to associate individuals across frames under complex and overlapping temporal trajectories. To address this, we introduce a novel Pose-Aware Video transformEr Network (PAVE-Net), which features a spatial encoder to model intra-frame relations and a spatiotemporal pose decoder to capture global dependencies across frames. To achieve accurate temporal association, we propose a pose-aware attention mechanism that enables each pose query to selectively aggregate features corresponding to the same individual across consecutive frames. Additionally, we explicitly model spatiotemporal dependencies among pose keypoints to improve accuracy. Notably, our approach is the first end-to-end method for multi-frame 2D human pose estimation. Extensive experiments show that PAVE-Net substantially outperforms prior image-based end-to-end methods, achieving a 6.0 mAP improvement on PoseTrack2017, and delivers accuracy competitive with state-of-the-art two-stage video-based approaches, while offering significant gains in efficiency. Yonghui Yu, Jiahang Cai, Xun Wang 0007, Wenwu Yang |
AAAI | 3 |
| 2026 | FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance CustomizationabstractGarment-centric fashion image generation aims to synthesize realistic and controllable human models dressing a given garment, which has attracted growing interest due to its practical applications in e-commerce. The key challenges of the task lie in two aspects: (1) faithfully preserving the garment details, and (2) gaining fine-grained controllability over the model's appearance. Existing methods typically require performing garment deformation in the generation process, which often leads to garment texture distortions. Also, they fail to control the fine-grained attributes of the generated models, due to the lack of specifically designed mechanisms. To address these issues, we propose FashionMAC, a novel diffusion-based deformation-free framework that achieves high-quality and controllable fashion showcase image generation. The core idea of our framework is to eliminate the need for performing garment deformation and directly outpaint the garment segmented from a dressed person, which enables faithful preservation of the intricate garment details. Moreover, we propose a novel region-adaptive decoupled attention (RADA) mechanism along with a chained mask injection strategy to achieve fine-grained appearance controllability over the synthesized human models. Specifically, RADA adaptively predicts the generated regions for each fine-grained text attribute and enforces the text attribute to focus on the predicted regions by a chained mask injection strategy, significantly enhancing the visual fidelity and the controllability. Extensive experiments validate the superior performance of our framework compared to existing state-of-the-art methods. Jinxiao Li, Jingnan Wang, Zhiwen Zuo, Jianfeng Dong, Wei Li 0111, Chi Wang 0004, Weiwei Xu 0003, Xun Wang 0007 |
AAAI | 9 |
| 2026 | PRVR: Partially Relevant Video RetrievalabstractIn current text-to-video retrieval (T2VR), videos to be retrieved have been properly trimmed so that a correspondence between the videos and ad-hoc textual queries naturally exists. Note in practice that videos circulated on the Internet and social media platforms, while being relatively short, are typically rich in their content. Often, multiple scenes / actions / events are shown in a single video, leading to a more challenging T2VR setting wherein only part of the video content is relevant w.r.t. a given query. This paper presents a first study on this setting which we term Partially Relevant Video Retrieval (PRVR). Considering that a video typically consists of multiple moments, a video is regarded as partially relevant w.r.t. to a given query if it contains a query-related moment. We formulate the PRVR task as a multiple instance learning problem, and propose a Multi-Scale Similarity Learning (MS-SL++) network that jointly learns both clip-scale and frame-scale similarities to determine the partial relevance between video-query pairs. Extensive experiments on three diverse video-text datasets (TVshow Retrieval, ActivityNet-Captions and Charades-STA) demonstrate the viability of the proposed method. Xianke Chen, Daizong Liu, Xun Yang 0001, Xirong Li 0001, Jianfeng Dong, Meng Wang 0001, Xun Wang 0007 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Innovative tooth segmentation using hierarchical features and bidirectional sequence modeling
Xinxin Zhao, Liqin Wu, Zhaocheng Xu, Wei-fa Yang, Yunuo Zou, Xun Wang 0007 |
Pattern Recognit. | 8 |
| 2026 | UCMIB-PNS: Balancing Sufficiency and Necessity With Probabilistic Causality and Cross-Modal Uncertainty in Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis aims to accurately identify sentiment orientations by integrating information from multiple modalities such as text, audio, and video. However, a key challenge in multimodal fusion is effectively balancing the sufficiency and necessity of information across modalities. Traditional models often fail to qualify and capture this balance due to the presence of noise and redundant information in multimodal data, leading to suboptimal performance in sentiment analysis. To address this issue, we propose a novel multimodal sentiment analysis method calledUCMIB-PNS, which is guided by information bottleneck and probabilistic causality. The method employs anUncertainCross-ModalInformationBottleneck(UCMIB)module to reduce redundant information within modalities and maximize discriminative information. The UCMIB utilizes codebooks to dynamically record the distributions of samples and employs random sampling to conduct uncertain modeling across different modalities. It integrates uncertainty-aware contrastive learning and KL divergence for dynamic comparison and compression of information from different modalities. Moreover, UCMIB-PNS uses differentiableProbability ofNecessity andSufficiency(PNS)estimators to estimate and re-weight the sufficiency and necessity of modalities by constructing several counterfactual scenarios through end-to-end learning. Experiments conducted on four publicly available multimodal sentiment analysis datasets demonstrate that UCMIB-PNS achieves optimal performance on both clean and noisy data. Extended experiments further validate the method's robustness under different types of noise. Jili Chen, Yihua Zhong, Qionghao Huang, Changqin Huang, Fan Jiang 0017, Xiaodi Huang 0001, Xun Wang 0007 |
IEEE Trans. Affect. Comput. | 7 |
| 2026 | Multi-Distribution Knowledge Distillation With Wasserstein Distance for Efficient 3D Human ReconstructionabstractRecently, transformer-based Human Mesh Recovery (HMR) from monocular images has achieved remarkable progress by effectively modeling long-range dependencies among body parts. However, the large-scale transformer architectures that drive these state-of-the-art results impose prohibitive computational and storage demands, severely limiting their deployment in real-time or resource-constrained scenarios. Existing knowledge distillation methods offer a potential solution but typically rely on a single distribution, either final outputs or intermediate features, thus failing to capture the diverse and complementary representations inherent in transformer-based HMR models, such as attention patterns. To address these limitations, we propose Multi-dIstribution kNowledge Distillation (MIND), a framework tailored for transformer-based HMR that transfers knowledge from multiple complementary distributions: final outputs for high-level task knowledge, intermediate features for visual representation learning, and internal attention heatmaps to preserve spatial focus patterns. Furthermore, motivated by the inherent conceptual similarity between Generative Adversarial Networks (GANs) and knowledge distillation, we introduce a Wasserstein distillation strategy that leverages the Wasserstein-GAN framework to robustly align these distributions between the teacher and student models. Extensive experiments on Human3.6M, 3DPW, and COCO demonstrate that MIND not only surpasses existing distillation baselines, reducing PA-MPJPE by 3.1mmon 3DPW, but also matches or exceeds the state-of-the-art teacher model HMR2.0 [1] while using only 17.8% of its parameters and 13.7% of its computation cost. Xun Wang 0007, Wenwu Yang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | TwinPose: Person-Specific Subspaces for Multi-View 3D Pose EstimationabstractFollowing the success of deep neural networks in 2D pose estimation, reconstruction-based approaches have significantly advanced multi-person 3D pose estimation from sparse multi-view images. These methods typically detect 2D poses independently in each view and then associate them for 3D reconstruction. However, despite strong progress, recent state-of-the-art methods still face critical limitations: 1) They often depend on global optimization over a large and complex set of multi-view 2D joints to jointly infer 3D poses for all individuals, making the process highly complex and prone to suboptimal solutions; 2) Their tight coupling with the bottom-up detector OpenPose hinders the use of more advanced top-down or single-stage 2D pose estimators and restricts the integration of richer instance-level cues learned by these models. To address these limitations, we propose TwinPose, a novel framework that alleviates the complexity of global pose inference by optimizing within person-specific 3D pose subspaces, while fully supporting diverse 2D pose detectors and effectively leveraging pose-instance cues. The key idea is to introduce a twin pose — a 3D counterpart of each 2D pose — that inherits its instance representation and aggregates geometrically consistent 2D joints from other views. All twin poses are unified in a common 3D space, where those belonging to the same individual naturally share a number of bones. This structural property enables association by counting shared bones, forming person-specific subspaces from which each individual's 3D pose can be inferred independently in an efficient and robust manner. Extensive experiments demonstrate that TwinPose achieves state-of-the-art performance in both accuracy and efficiency across multiple public and proprietary datasets. Importantly, it is fully detector-agnostic, allowing seamless integration with current and future advances in 2D pose estimation while remaining highly robust to noisy or imperfect 2D predictions. Project page with code and additional resources: https://github.com/zgspose/TwinPose Wenwu Yang, Tianyi He, Jiwei Ding, Xun Wang 0007, Kun Zhou 0001 |
ACM Trans. Graph. | 4 |
| 2025 | Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal RetrievalabstractExisting cross-modal retrieval methods typically rely on large-scale vision-language pair data. This makes it challenging to efficiently develop a cross-modal retrieval model for under-resourced languages of interest. Therefore, Cross-lingual Cross-modal Retrieval (CCR), which aims to align vision and the low-resource language (the target language) without using any human-labeled target-language data, has gained increasing attention. As a general parameter-efficient way, a common solution is to utilize adapter modules to transfer the vision-language alignment ability of Vision-Language Pretraining (VLP) models from a source language to a target language. However, these adapters are usually static once learned, making it difficult to adapt to target-language captions with varied expressions. To alleviate it, we propose Dynamic Adapter with Semantics Disentangling (DASD), whose parameters are dynamically generated conditioned on the characteristics of the input captions. Considering that the semantics and expression styles of the input caption largely influence how to encode it, we propose a semantic disentangling module to extract the semantic-related and semantic-agnostic features from the input, ensuring that generated adapters are well-suited to the characteristics of input caption. Extensive experiments on two image-text datasets and one video-text dataset demonstrate the effectiveness of our model for cross-lingual cross-modal retrieval, as well as its good compatibility with various VLP models. Zhiyu Dong, Jianfeng Dong, Xun Wang 0007 |
AAAI | 4 |
| 2025 | Towards Ship License Plate Recognition in the Wild: A Large Benchmark and Strong BaselineabstractThe paper targets the challenging task of Ship License Plate (SLP) recognition. Existing methods for SLP recognition are hampered by the scarcity of large and publicly available datasets, leading to evaluations on small and non-representative datasets. To alleviate it, we have built a large dataset, called SLP34K, which consists of 34,385 images collected by an intelligent traffic surveillance system. The dataset is carefully manually annotated with text labels and attributes, and presents high data diversity by multiple installation locations and long capturing period of the cameras. Additionally, we propose a simple yet effective SLP recognition baseline method. The baseline is equipped with a strong visual encoder that benefits from initial pre-training via self-supervised learning, followed by further refinement through our devised semantic enhancement module. Extensive experiments on SLP34K verify the effectiveness of our proposed baseline. Moreover, while our baseline is designed for SLP recognition, it can also be used for common scene text recognition and achieve state-of-the-art performance on seven mainstream scene text recognition datasets. Ruiqing Yang, Roukai Huang, Chuanhuang Li, Xun Wang 0007, Jianfeng Dong |
AAAI | 8 |
| 2025 | LLM-Assisted Entropy-Based Adaptive Distillation for Unsupervised Fine-Grained Visual Representation Learning
Jianfeng Dong, Daizong Liu, Jie Sun 0034, Xiaoye Qu, Xun Yang 0001, Dongsheng Liu 0003, Xun Wang 0007 |
ICCV | 8 |
| 2025 | Real-Time Misinformation Detection with Cyclic Evidence-Based FrameworkabstractExisting misinformation detection benchmark datasets (e.g., COVMIS and LIAR2) are limited by their reliance on fact-checking labels that are prone to factual inaccuracies due to cognitive constraints of fact-checkers and outdated labels. Prior misinformation detection tasks have been hindered by the dual problems of label redundancy and cold start. To this end, we propose a novel Cyclic Evidence-based Misinformation Detection (CEMD) framework, which incorporates two core mechanisms: (i) a Retrieval Augmented Generation (RAG) pipeline that leverages the latest external knowledge to augment insufficient prior knowledge; and (ii) a cyclic evidence-bootstrapping mechanism that mitigates label redundancy and cold start. We introduce an improved dataset, COVMIS2, built upon COVMIS, and conduct comprehensive experiments to evaluate the efficacy of our framework. Our results demonstrate that the CEMDo outperforms the prior state-of-the-art (SOTA) baseline on LIAR2 by 11.95% and surpasses the human baseline on COVMIS2 by 6.31%, leveraging the Llama-3-70B-Instruct model to augment prior knowledge and the DoRA fine-tuned Llama-3-8B-Instruct model for binary classification. Furthermore, we curate new benchmark datasets, COVMIS2024 and LIAR2024, by recategorizing the redundant labels of COVMIS2 and LIAR2 through the CEMDo. Zhiwen Hu, Lv Han, Haihua Jiang, Xi'ao Ma, Saihua Lei, Haojia Niu, Zehui Zhou, Xun Wang 0007 |
IJCNN | 9 |
| 2025 | Open-World Fine-Grained Fashion Retrieval with LLM-based Commonsense Knowledge InfusionabstractAttribute-Specific Fashion Retrieval (ASFR) focuses on retrieving images based on fine-grained, attribute-specific criteria rather than naive global visual similarity, enabling more precise and interpretable search results. Existing ASFR methods ideally assume that all attribute semantics are in-domain distributions of the training datasets. However, realistic scenarios are generally more complex and naturally contain unseen attribute information, often resulting in ungeneralizable retrieval outcomes. In this paper, we take the first step to address the new and challenging open-world ASFR setting, which involves handling diverse and practical attributes instead of relying solely on predefined attribute sets in closed-world scenarios. Specifically, to comprehend unseen attributes, we propose a novel LLM-based Commonsense Knowledge Infusion (CoKi) framework that integrates commonsense knowledge as complementary context into attribute representations using a Large Language Model (LLM). By infusing such LLM-based commonsense knowledge through descriptive contexts, our method enables robust semantic enrichment and effective generalization to unseen attributes. Additionally, we introduce a modality-switchable prompt and an imputation mechanism to ensure model robustness across diverse input configurations by dynamically adapting to missing modalities. Extensive experiments demonstrate that our approach not only achieves state-of-the-art in-domain retrieval performance but also significantly enhances adaptability to unseen attributes and cross-domain generalization, establishing a new benchmark for fine-grained fashion retrieval in open-world scenarios. Our source code is publicly available at https://github.com/HuiGuanLab/CoKi. Jianfeng Dong, Daizong Liu, Xiaoye Qu, Cuizhu Bao, Zhike Han, Jixiang Zhu, Xun Wang 0007 |
SIGIR | 8 |
| 2025 | Advancing Ship Re-Identification in the Wild: The ShipReID-2400 Benchmark Dataset and D2InterNet Baseline MethodabstractShip Re-Identification (ReID) aims to accurately identify ships with the same identity across different times and camera views, playing a crucial role in intelligent waterway transportation. However, compared to the widely researched pedestrian and vehicle ReID, Ship ReID has received much less attention, primarily due to the scarcity of large-scale and high-quality ship ReID datasets available for public access. Moreover, several unique challenges make ship ReID particularly difficult: ships are large objects that are hard to capture fully, and the visible area of ships vary significantly due to changes in cargo loading or water surface conditions. These challenges make it difficult to achieve ideal results by directly applying existing ReID methods. To address these challenges, in this paper, we introduce ShipReID-2400, a dataset for ship ReID compiled from a real-world intelligent waterway traffic monitoring system. It comprises 17,241 images of 2,400 distinct ship identities collected over 53 months, ensuring diversity and representativeness. Furthermore, we propose the Disentangle-to-Interact Network ( D2InterNet ), a simple but strong baseline for ship ReID designed to extract discriminative local features despite significant scale variations. Extensive experimental results show that D2InterNet achieves state-of-the-art performance on both the ShipReID-2400 and VesselReID datasets. In addition, despite being designed for ship ReID, D2InterNet also achieves competitive results on the MSMT17 pedestrian ReID dataset, showcasing its good generalization capability. Our dataset and code are publicly available at https://github.com/HuiGuanLab/ShipReID-2400. Roukai Huang, Chuanhuang Li, Jie Sun 0034, Jianfeng Dong, Xun Wang 0007 |
SIGIR | 7 |
| 2025 | Representation alignment contrastive regularisation for multi-object trackingabstractAbstract Achieving high‐performance in multi‐object tracking algorithms heavily relies on modelling spatial‐temporal relationships during the data association stage. Mainstream approaches encompass rule‐based and deep learning‐based methods for spatial‐temporal relationship modelling. While the former relies on physical motion laws, offering wider applicability but yielding suboptimal results for complex object movements, the latter, though achieving high‐performance, lacks interpretability and involves complex module designs. This work aims to simplify deep learning‐based spatial‐temporal relationship models and introduce interpretability into features for data association. Specifically, a lightweight single‐layer transformer encoder is utilised to model spatial‐temporal relationships. To make features more interpretative, two contrastive regularisation losses based on representation alignment are proposed, derived from spatial‐temporal consistency rules. By applying weighted summation to affinity matrices, the aligned features can seamlessly integrate into the data association stage of the original tracking workflow. Experimental results showcase that our model enhances the majority of existing tracking networks' performance without excessive complexity, with minimal increase in training overhead and nearly negligible computational and storage costs. Shujie Chen 0001, Zhonglin Liu, Jianfeng Dong, Xun Wang 0007 |
IET Comput. Vis. | 4 |
| 2025 | Sdreplay: diffusion model for continual semantic segmentation in traffic scenarios
Yongchuan Xu, Zhaocheng Xu, Xun Wang 0007 |
Multim. Syst. | 5 |
| 2025 | Efficient Retinex-Based Framework for Low-Light Image Enhancement Without Additional NetworksabstractImages captured in low-light environments often suffer from significant degradation. However, most existing Retinex-based methods require an additional decomposition network and overlook the degradation caused by the illumination adjustment process, which results in the consumption of significant computational resources to achieve only average performance. To address the above issues, this paper proposes a more efficient Retinex-based approach named RetinexMac that allows training without an additional decomposition network or regularization functions. RetinexMac first employs an illumination coefficient estimation network to estimate the transform map and light up the global illumination and the local contrast of input images, then a multiscale degradation estimation network is used to suppress the degradation amplified by the illumination adjustment. In order to accurately estimate the degradation, a convolution and attention mixed module integrates the global and local spatial information. This is shown to also significantly improve the performance of other previous Retinex-based methods. Extensive experiments on several representative datasets show that our RetinexMac achieves both current state-of-the-art (SOTA) performance and more ideal visual appearance in terms of illumination and detail, as well as computational efficiency. Philip Birch, Xun Wang 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Multi-threshold deep metric learning for facial expression recognitionabstractFeature representations generated through triplet-based deep metric learning offer significant advantages for facial expression recognition (FER). Each threshold in triplet loss inherently shapes a distinct distribution of inter-class variations, leading to unique representations of expression features. Nonetheless, pinpointing the optimal threshold for triplet loss presents a formidable challenge, as the ideal threshold varies not only across different datasets but also among classes within the same dataset. In this paper, we propose a novel multi-threshold deep metric learning approach that bypasses the complex process of threshold validation and markedly improves the effectiveness in creating expression feature representations. Instead of choosing a single optimal threshold from a valid range, we comprehensively sample thresholds throughout this range, which ensures that the representation characteristics exhibited by the thresholds within this spectrum are fully captured and utilized for enhancing FER. Specifically, we segment the embedding layer of the deep metric learning network into multiple slices, with each slice representing a specific threshold sample. We subsequently train these embedding slices in an end-to-end fashion, applying triplet loss at its associated threshold to each slice, which results in a collection of unique expression features corresponding to each embedding slice. Moreover, we identify the issue that the traditional triplet loss may struggle to converge when employing the widely-used Batch Hard strategy for mining informative triplets, and introduce a novel loss termed dual triplet loss to address it. Extensive evaluations demonstrate the superior performance of the proposed approach on both posed and spontaneous facial expression datasets. Wenwu Yang, Jinyi Yu, Tuo Chen, Zhenguang Liu, Xun Wang 0007, Jianbing Shen |
Pattern Recognit. | 5 |
| 2024 | Statistics Enhancement Generative Adversarial Networks for Diverse Conditional Image SynthesisabstractConditional generative adversarial networks (cGANs) aim to synthesize diverse images given the input conditions and the latent codes, but they are prone to map an input to a single output regardless of the variations in latent code, which is also well known as the mode collapse problem of cGANs. To alleviate the problem, in this paper, we investigate explicitly enhancing the statistical dependency between the latent code and the synthesized image in cGANs by utilizing mutual information neural estimators to estimate and maximize the conditional mutual information (CMI) between them given the input condition. The method provides a new perspective from information theory to improve diversity for cGANs and can facilitate many existing conditional image synthesis frameworks with a simple neural estimator extension. Moreover, our studies show that several key designs, including the neural estimator choice, the neural estimator’s network design, and the sampling strategy, are crucial to the success of the method. Extensive experiments on four popular conditional image synthesis tasks, including class-conditioned image generation, paired and unpaired image-to-image translation, and text-to-image generation, demonstrate the effectiveness and superiority of the proposed method. Zhiwen Zuo, Ailin Li, Zhizhong Wang, Lei Zhao 0011, Jianfeng Dong, Xun Wang 0007, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Dual-View Curricular Optimal Transport for Cross-Lingual Cross-Modal RetrievalabstractCurrent research on cross-modal retrieval is mostly English-oriented, as the availability of a large number of English-oriented human-labeled vision-language corpora. In order to break the limit of non-English labeled data, cross-lingual cross-modal retrieval (CCR) has attracted increasing attention. Most CCR methods construct pseudo-parallel vision-language corpora via Machine Translation (MT) to achieve cross-lingual transfer. However, the translated sentences from MT are generally imperfect in describing the corresponding visual contents. Improperly assuming the pseudo-parallel data are correctly correlated will make the networks overfit to the noisy correspondence. Therefore, we propose Dual-view Curricular Optimal Transport (DCOT) to learn with noisy correspondence in CCR. In particular, we quantify the confidence of the sample pair correlation with optimal transport theory from both the cross-lingual and cross-modal views, and design dual-view curriculum learning to dynamically model the transportation costs according to the learning stage of the two views. Extensive experiments are conducted on two multilingual image-text datasets and one video-text dataset, and the results demonstrate the effectiveness and robustness of the proposed method. Besides, our proposed method also shows a good expansibility to cross-lingual image-text baselines and a decent generalization on out-of-domain data. Shuhui Wang, Hao Luo 0004, Jianfeng Dong, Fan Wang 0019, Xun Wang 0007, Meng Wang 0001 |
IEEE Trans. Image Process. | 7 |
| 2024 | Cross-Lingual Cross-Modal Retrieval With Noise-Robust Fine-TuningabstractCross-lingual cross-modal retrieval aims at leveraging human-labeled annotations in a source language to construct cross-modal retrieval models for a new target language, due to the lack of manually-annotated dataset in low-resource languages (target languages). Contrary to the growing developments in the field of monolingual cross-modal retrieval, there has been less research focusing on cross-modal retrieval in the cross-lingual scenario. A straightforward method to obtain target-language labeled data is translating source-language datasets utilizing Machine Translations (MT). However, as MT is not perfect, it tends to introduce noise during translation, rendering textual embeddings corrupted and thereby compromising the retrieval performance. To alleviate this, we propose Noise-Robust Fine-tuning (NRF) which tries to extract clean textual information from a possibly noisy target-language input with the guidance of its source-language counterpart. Besides, contrastive learning involving different modalities are performed to strengthen the noise-robustness of our model. Different from traditional cross-modal retrieval methods which only employ image/video-text paired data for fine-tuning, in NRF, selected parallel data plays a key role in improving the noise-filtering ability of our model. Extensive experiments are conducted on three video-text and image-text retrieval benchmarks across different target languages, and the results demonstrate that our method significantly improves the overall performance without using any image/video-text paired data on target languages. Jianfeng Dong, Tianxiang Liang, Yonghui Liang, Xun Yang 0001, Xun Wang 0007, Meng Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2023 | Hierarchical Contrast for Unsupervised Skeleton-Based Action Representation LearningabstractThis paper targets unsupervised skeleton-based action representation learning and proposes a new Hierarchical Contrast (HiCo) framework. Different from the existing contrastive-based solutions that typically represent an input skeleton sequence into instance-level features and perform contrast holistically, our proposed HiCo represents the input into multiple-level features and performs contrast in a hierarchical manner. Specifically, given a human skeleton sequence, we represent it into multiple feature vectors of different granularities from both temporal and spatial domains via sequence-to-sequence (S2S) encoders and unified downsampling modules. Besides, the hierarchical contrast is conducted in terms of four levels: instance level, domain level, clip level, and part level. Moreover, HiCo is orthogonal to the S2S encoder, which allows us to flexibly embrace state-of-the-art S2S encoders. Extensive experiments on four datasets, i.e., NTU-60, NTU-120, PKU-I and PKU-II, show that HiCo achieves a new state-of-the-art for unsupervised skeleton-based action representation learning in two downstream tasks including action recognition and retrieval, and its learned action representation is of good transferability. Besides, we also show that our framework is effective for semi-supervised skeleton-based action recognition. Our code is available at https://github.com/HuiGuanLab/HiCo. Jianfeng Dong, Shengkai Sun, Zhonglin Liu, Shujie Chen 0001, Xun Wang 0007 |
AAAI | 6 |
| 2023 | Weakly Supervised Method for Domain Adaptation in Instance Segmentation
Jie Sun 0034, Zhaocheng Xu, Zhaoyi Jiang, Xun Wang 0007 |
CGI (1) | 7 |
| 2023 | Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video RetrievalabstractAlmost all previous text-to-video retrieval works assume that videos are pre-trimmed with short durations. However, in practice, videos are generally untrimmed containing much background content. In this work, we investigate the more practical but challenging Partially Relevant Video Retrieval (PRVR) task, which aims to retrieve partially relevant untrimmed videos with the query input. Particularly, we propose to address PRVR from a new perspective, i.e., distilling the generalization knowledge from the large-scale vision-language pre-trained model and transferring it to a task-specific PRVR network. To be specific, we introduce a Dual Learning framework with Dynamic Knowledge Distillation (DL-DKD), which exploits the knowledge of a large vision-language model as the teacher to guide a student model. During the knowledge distillation, an inheritance student branch is devised to absorb the knowledge from the teacher model. Considering that the large model may be of mediocre performance due to the domain gaps, we further develop an exploration student branch to take the benefits of task-specific information. In addition, a dynamical knowledge distillation strategy is further devised to adjust the effect of each student branch learning during the training. Experiment results demonstrate that our proposed model achieves state-of-the-art performance on ActivityNet and TVR datasets for PRVR. Jianfeng Dong, Minsong Zhang, Xianke Chen, Daizong Liu, Xiaoye Qu, Xun Wang 0007 |
ICCV | 7 |
| 2023 | Lite-MKD: A Multi-modal Knowledge Distillation Framework for Lightweight Few-shot Action RecognitionabstractExisting few-shot action recognition methods have placed primary focus on improving the recognition accuracy while neglecting another important indicator in practical scenarios, i.e., model efficiency. In this paper, we make the first attempt and propose a Lightweight Multi-modal Knowledge Distillation framework (Lite-MKD) for few-shot action recognition. In this framework, the teacher model conducts multi-modal learning to achieve a comprehensive fusion of the optical flow, depth, and appearance features of human movements, thus achieving a more robust representation of actions. The student model is utilized to learn to recognize actions from the single RGB modality at a lower computational cost under the guidance of the teacher. To fully explore and integrate multi-modal information, a hierarchical Multi-modal Fusion Module (MFM) is introduced in the teacher model. Besides, a multi-level Distinguish-to-Mimic (D2M) knowledge distillation component is proposed for the student model. D2M improves the ability of the student model to mimic the action classification probabilities of the teacher model by enhancing the distinguishability of the student model for different video categories in the support set. Extensive experiments on three action recognition datasets Kinetics, HMDB51, and UCF101 demonstrate our framework's effectiveness and stable generalization ability. With a much more lightweight network for inference, we achieve comparable performance to previous state-of-the-art methods. Our source code is available at https://github.com/HuiGuanLab/Lite-MKD Daizong Liu, Xiaoye Qu, Junyu Gao 0002, Jianfeng Dong, Xun Wang 0007 |
ACM Multimedia | 8 |
| 2023 | Unified Multi-modal Unsupervised Representation Learning for Skeleton-based Action UnderstandingabstractUnsupervised pre-training has shown great success in skeleton-based action understanding recently. Existing works typically train separate modality-specific models (i.e., joint, bone, and motion), then integrate the multi-modal information for action understanding by a late-fusion strategy. Although these approaches have achieved significant performance, they suffer from the complex yet redundant multi-stream model designs, each of which is also limited to the fixed input skeleton modality. To alleviate these issues, in this paper, we propose a Unified Multimodal Unsupervised Representation Learning framework, called UmURL, which exploits an efficient early-fusion strategy to jointly encode the multi-modal features in a single-stream manner. Specifically, instead of designing separate modality-specific optimization processes for uni-modal unsupervised learning, we feed different modality inputs into the same stream with an early-fusion strategy to learn their multi-modal features for reducing model complexity. To ensure that the fused multi-modal features do not exhibit modality bias, i.e., being dominated by a certain modality input, we further propose both intra- and inter-modal consistency learning to guarantee that the multi-modal features contain the complete semantics of each modal via feature decomposition and distinct alignment. In this manner, our framework is able to learn the unified representations of uni-modal or multi-modal skeleton input, which is flexible to different kinds of modality input for robust action understanding in practical cases. Extensive experiments conducted on three large-scale datasets, i.e., NTU-60, NTU-120, and PKU-MMD II, demonstrate that UmURL is highly efficient, possessing the approximate complexity with the uni-modal methods, while achieving new state-of-the-art performance across various downstream task scenarios in skeleton-based action representation learning. Our source code is available at https://github.com/HuiGuanLab/UmURL. Shengkai Sun, Daizong Liu, Jianfeng Dong, Xiaoye Qu, Junyu Gao 0002, Xun Yang 0001, Xun Wang 0007, Meng Wang 0001 |
ACM Multimedia | 7 |
| 2023 | RLP-VIO: Robust and lightweight plane-based visual-inertial odometry for augmented realityabstractAbstract We propose RLP‐VIO—a robust and lightweight monocular visual‐inertial odometry system using multiplane priors. With planes extracted from the point cloud, visual‐inertial‐plane PnP uses the plane information for fast localization. Depth estimation is susceptible to degenerated motion, so the planes are expanded in a reprojection consensus‐based way robust to depth errors. For sensor fusion, our sliding‐window optimization uses a novel structureless plane‐distance error cost, which prevents the fill‐in effect that poisons the BA problem's sparsity and permits the use of a smaller sliding window while maintaining good accuracy. The total computational cost is further reduced with our modified marginalization strategy. To further improve the tracking robustness, the landmark depths are constrained using the planes during degenerated motion. The whole system is parallelized with a three‐stage pipeline. Under controlled environments, this parallelization runs deterministically and produces consistent results. The resulting VIO system is tested on widely used datasets and compared with several state‐of‐the‐art systems. Our system achieves competitive accuracy and works robustly even on long and challenging sequences. To demonstrate the effectiveness of the proposed system, we also show the AR application running on mobile devices in real‐time. Jinyu Li 0002, Bangbang Yang, Guofeng Zhang 0001, Xun Wang 0007, Hujun Bao |
Comput. Animat. Virtual Worlds | 5 |
| 2023 | Combining Graph Neural Networks With Expert Knowledge for Smart Contract Vulnerability DetectionabstractSmart contract vulnerability detection draws extensive attention in recent years due to the substantial losses caused by hacker-attacks. Existing efforts for contract security analysis heavily rely on rigid rules defined by experts, which is labor-intensive and non-scalable. More importantly, expert-defined rules tend to be error-prone and suffer the inherent risk of being cheated by crafty attackers. Recent researches focus on the symbolic execution and formal analysis of smart contract for vulnerability detection, yet to achieve a precise and scalable solution. Although several methods have been proposed to detect vulnerabilities in smart contracts, there is still a lack of effort that considers combining expert-defined security patterns with deep neural networks. In this paper, we explore using graph neural networks and expert knowledge for smart contract vulnerability detection. Specifically, we cast the rich control- and data- flow semantics of the source code into a contract graph. Then, we propose a novel temporal message propagation network to extract graph feature from the normalized graph, and combine the graph feature with expert patterns to yield a final detection system. Extensive experiments are conducted on all the smart contracts that have source code in two platforms. Empirical results show significant accuracy improvements over state-of-the-art methods. Zhenguang Liu, Xiaoyang Wang 0002, Xun Wang 0007 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | Progressive Localization Networks for Language-Based Moment LocalizationabstractThis article targets the task of language-based video moment localization. The language-based setting of this task allows for an open set of target activities, resulting in a large variation of the temporal lengths of video moments. Most existing methods prefer to first sample sufficient candidate moments with various temporal lengths, then match them with the given query to determine the target moment. However, candidate moments generated with a fixed temporal granularity may be suboptimal to handle the large variation in moment lengths. To this end, we propose a novel multi-stage Progressive Localization Network (PLN) that progressively localizes the target moment in a coarse-to-fine manner. Specifically, each stage of PLN has a localization branch and focuses on candidate moments that are generated with a specific temporal granularity. The temporal granularities of candidate moments are different across the stages. Moreover, we devise a conditional feature manipulation module and an upsampling connection to bridge the multiple localization branches. In this fashion, the later stages are able to absorb the previously learned information, thus facilitating the more fine-grained localization. Extensive experiments on three public datasets demonstrate the effectiveness of our proposed PLN for language-based moment localization, especially for localizing short moments in long videos. Jianfeng Dong, Xiaoye Qu, Xun Yang 0001, Pan Zhou 0001, Xun Wang 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2022 | Stable Community Detection in Signed Social Networks (Extended abstract)abstractCommunity detection is a fundamental problem in graph analysis, while most existing research focuses on unsigned graphs. In many applications, networks involve both positive and negative connections. It is important to exploit the signed information to identify more stable communities. In this paper, we propose a novel model, named stable k-core, to measure the stability of a community in signed graphs by leveraging the concept of balance theory. We show that the problem of finding the maximum stable k-core is NP-hard. Advanced approaches are proposed to accelerate the processing. Experiments on 6 signed networks are conducted to verify the efficiency and effectiveness of proposed model and techniques. Renjie Sun, Chen Chen 0017, Xiaoyang Wang 0002, Xun Wang 0007 |
ICDE | 4 |
| 2022 | Partially Relevant Video RetrievalabstractCurrent methods for text-to-video retrieval (T2VR) are trained and tested on video-captioning oriented datasets such as MSVD, MSR-VTT and VATEX. A key property of these datasets is that videos are assumed to be temporally pre-trimmed with short duration, whilst the provided captions well describe the gist of the video content. Consequently, for a given paired video and caption, the video is supposed to be fully relevant to the caption. In reality, however, as queries are not known a priori, pre-trimmed video clips may not contain sufficient content to fully meet the query. This suggests a gap between the literature and the real world. To fill the gap, we propose in this paper a novel T2VR subtask termed Partially Relevant Video Retrieval (PRVR). An untrimmed video is considered to be partially relevant w.r.t. a given textual query if it contains a moment relevant to the query. PRVR aims to retrieve such partially relevant videos from a large collection of untrimmed videos. PRVR differs from single video moment retrieval and video corpus moment retrieval, as the latter two are to retrieve moments rather than untrimmed videos. We formulate PRVR as a multiple instance learning (MIL) problem, where a video is simultaneously viewed as a bag of video clips and a bag of video frames. Clips and frames represent video content at different time scales. We propose a Multi-Scale Similarity Learning (MS-SL) network that jointly learns clip-scale and frame-scale similarities for PRVR. Extensive experiments on three datasets (TVR, ActivityNet Captions, and Charades-STA) demonstrate the viability of the proposed method. We also show that our method can be used for improving video corpus moment retrieval. Jianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang 0001, Shujie Chen 0001, Xirong Li 0001, Xun Wang 0007 |
ACM Multimedia | 7 |
| 2022 | Cross-Lingual Cross-Modal Retrieval with Noise-Robust LearningabstractDespite the recent developments in the field of cross-modal retrieval, there has been less research focusing on low-resource languages due to the lack of manually annotated datasets. In this paper, we propose a noise-robust cross-lingual cross-modal retrieval method for low-resource languages. To this end, we use Machine Translation (MT) to construct pseudo-parallel sentence pairs for low-resource languages. However, as MT is not perfect, it tends to introduce noise during translation, rendering textual embeddings corrupted and thereby compromising the retrieval performance. To alleviate this, we introduce a multi-view self-distillation method to learn noise-robust target-language representations, which employs a cross-attention module to generate soft pseudo-targets to provide direct supervision from the similarity-based view and feature-based view. Besides, inspired by the back-translation in unsupervised MT, we minimize the semantic discrepancies between origin sentences and back-translated sentences to further improve the noise robustness of the textual encoder. Extensive experiments are conducted on three video-text and image-text cross-modal retrieval benchmarks across different languages, and the results demonstrate that our method significantly improves the overall performance without using extra human-labeled data. In addition, equipped with a pre-trained visual encoder from a recent vision and language pre-training framework, i.e., CLIP, our model achieves a significant performance gain, showing that our method is compatible with popular pre-training models. Code and data are available at https://github.com/HuiGuanLab/nrccr. Jianfeng Dong, Tianxiang Liang, Minsong Zhang, Xun Wang 0007 |
ACM Multimedia | 6 |
| 2022 | Non-homogeneous haze data synthesis based real-world image dehazing with enhancement-and-restoration fused CNNs
Shuangshuang Ye, Lideng Zhang, Haiyong Bao, Xun Wang 0007, Fanding Wu |
Comput. Graph. | 5 |
| 2022 | Geometrically interpretable Variance Hyper Rectangle learning for pattern classification
Jie Sun 0034, Huamao Gu, Haoyu Peng, Yili Fang, Xun Wang 0007 |
Eng. Appl. Artif. Intell. | 5 |
| 2022 | Harshness-aware sentiment mining framework for product review
Xun Wang 0007, Xiaoyang Wang 0002, Yili Fang |
Expert Syst. Appl. | 1 |
| 2022 | A tooth surface design method combining semantic guidance, confidence, and structural coherenceabstractAbstract Research on tooth surface design based on deep neural networks has recently achieved progress in terms of both accuracy and execution efficiency. However, unrealistic outputs are still a challenging issue, partially because of (1) the lack of semantic guidance, (2) the inability to discover and rectify false results, and (3) the lack of exploration of structural coherence in intermediate layers. In this paper, we present an approach to predict depth images for designed teeth based on a conditional generative adversarial network (CGAN) by incorporating semantic guidance. Moreover, the uncertainty of semantic inference is employed to improve the model outputs, and a structural coherence loss is proposed for adversarial learning to enhance the discrimination capability of the network in intermediate layers. We evaluate the performance of our approach with the Shining3D tooth dataset. The experimental results show that our method produces better results than the other available approaches in terms of accuracy. Nali Liu, Shuangming Chai, Xun Wang 0007, Ruili Wang 0001 |
IET Comput. Vis. | 6 |
| 2022 | FeatInter: Exploring fine-grained object features for video-text retrieval
Minsong Zhang, Jianfeng Dong, Xun Wang 0007 |
Neurocomputing | 6 |
| 2022 | Heterogeneous data fusion and loss function design for tooth point cloud segmentation
Dongsheng Liu 0003, Judith Gelernter, Xun Wang 0007 |
Neural Comput. Appl. | 5 |
| 2022 | Privacy Enhanced Cloud-Based Facial Recognition
Jie Sun 0034, Xun Wang 0007 |
Neural Process. Lett. | 4 |
| 2022 | Dual Encoding for Video Retrieval by TextabstractThis paper attacks the challenging problem of video retrieval by text. In such a retrieval paradigm, an end user searches for unlabeled videos by ad-hoc queries described exclusively in the form of a natural-language sentence, with no visual example provided. Given videos as sequences of frames and queries as sequences of words, an effective sequence-to-sequence cross-modal matching is crucial. To that end, the two modalities need to be first encoded into real-valued vectors and then projected into a common space. In this paper we achieve this by proposing a dual deep encoding network that encodes videos and queries into powerful dense representations of their own. Our novelty is two-fold. First, different from prior art that resorts to a specific single-level encoder, the proposed network performs multi-level encoding that represents the rich content of both modalities in a coarse-to-fine fashion. Second, different from a conventional common space learning algorithm which is either concept based or latent space based, we introduce hybrid space learning which combines the high performance of the latent space and the good interpretability of the concept space. Dual encoding is conceptually simple, practically effective and end-to-end trained with hybrid space learning. Extensive experiments on four challenging video datasets show the viability of the new method. Code and data are available at https://github.com/danieljf24/hybrid_space. Jianfeng Dong, Xirong Li 0001, Chaoxi Xu, Xun Yang 0001, Gang Yang 0001, Xun Wang 0007, Meng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Reading-Strategy Inspired Visual Representation Learning for Text-to-Video RetrievalabstractThis paper aims for the task of text-to-video retrieval, where given a query in the form of a natural-language sentence, it is asked to retrieve videos which are semantically relevant to the given query, from a great number of unlabeled videos. The success of this task depends on cross-modal representation learning that projects both videos and sentences into common spaces for semantic similarity computation. In this work, we concentrate on video representation learning, an essential component for text-to-video retrieval. Inspired by the reading strategy of humans, we propose a Reading-strategy Inspired Visual Representation Learning (RIVRL) to represent videos, which consists of two branches: a previewing branch and an intensive-reading branch. The previewing branch is designed to briefly capture the overview information of videos, while the intensive-reading branch is designed to obtain more in-depth information. Moreover, the intensive-reading branch is aware of the video overview captured by the previewing branch. Such holistic information is found to be useful for the intensive-reading branch to extract more fine-grained features. Extensive experiments on three datasets are conducted, where our model RIVRL achieves a new state-of-the-art on TGIF and VATEX. Moreover, on MSR-VTT, our model using two video features shows comparable performance to the state-of-the-art using seven video features and even outperforms models pre-trained on the large-scale HowTo100M dataset. Code is available athttps://github.com/LiJiaBei-7/rivrl. Jianfeng Dong, Xianke Chen, Xiaoye Qu, Xirong Li 0001, Yuan He 0011, Xun Wang 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | EFINet: Restoration for Low-Light Images via Enhancement-Fusion Iterative NetworkabstractThe lighting environment in the real world is so complex that most existing low-light image restoration methods suffer from color cast and local over-exposure. In order to solve these problems, this paper proposes the enhancement-fusion iterative network (EFINet) for low-light image enhancement. Within each iteration of EFINet, a stretching coefficient estimation network based enhancement module is designed to adjust the input image pixel-wisely to obtain the initial enhancement result with the estimated coefficient maps. Then, an encoder-decoder based fusion network is devised to extract the deep features and combine the well-exposed local areas in both the input image and the initially enhanced image, to obtain a visually pleasing, high-quality image enhancement result. The coefficient estimation network and the fusion network are weight shared among all iterations. What’s more, most of the low-light image datasets are generated through illumination reduction in a global way. To better simulate the diverse illumination distribution in the real world, we put forward a new low-light image synthesis method to produce the low-light images with non-uniform illumination for the network training purpose. After conducting extensive experiments on both synthetic and real-world low-light images, the results verify the superiority of our algorithm over the state-of-the-art (SOTA) methods, especially in balancing the brightness difference and preventing over-enhancement. Fanding Wu, Xun Wang 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Stable Community Detection in Signed Social NetworksabstractCommunity detection is one of the most fundamental problems in social network analysis, while most existing research focuses on unsigned graphs. In real applications, social networks involve not only positive relationships but also negative ones. It is important to exploit the signed information to identify more stable communities. In this paper, we propose a novel model, named stable$k$-core, to measure the stability of a community in signed graphs. The stable$k$-core model not only emphasizes user engagement, but also eliminates unstable structures. We show that the problem of finding the maximum stable$k$-core is NP-hard. To scale for large graphs, novel pruning strategies and searching methods are proposed. We conduct extensive experiments on 6 real-world signed networks to verify the efficiency and effectiveness of proposed model and techniques. Renjie Sun, Chen Chen 0017, Xiaoyang Wang 0002, Ying Zhang 0001, Xun Wang 0007 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Locating pivotal connections: the K-Truss minimization and maximization problems
Chen Chen 0017, Renjie Sun, Xiaoyang Wang 0002, Xun Wang 0007 |
World Wide Web | 6 |
| 2021 | Deep Dual Consecutive Network for Human Pose EstimationabstractMulti-frame human pose estimation in complicated situations is challenging. Although state-of-the-art human joints detectors have demonstrated remarkable results for static images, their performances come short when we apply these models to video sequences. Prevalent shortcomings include the failure to handle motion blur, video defocus, or pose occlusions, arising from the inability in capturing the temporal dependency among video frames. On the other hand, directly employing conventional recurrent neural networks incurs empirical difficulties in modeling spatial contexts, especially for dealing with pose occlusions. In this paper, we propose a novel multi-frame human pose estimation framework, leveraging abundant temporal cues between video frames to facilitate keypoint detection. Three modular components are designed in our framework. A Pose Temporal Merger encodes keypoint spatiotemporal context to generate effective searching scopes while a Pose Residual Fusion module computes weighted pose residuals in dual directions. These are then processed via our Pose Correction Network for efficient refining of pose estimations. Our method ranks No.1 in the Multi-frame Person Pose Estimation Challenge on the large-scale benchmark datasets PoseTrack2017 and PoseTrack2018. We have released our code, hoping to inspire future research. Zhenguang Liu, Haoming Chen, Runyang Feng, Shuang Wu 0002, Shouling Ji, Bailin Yang, Xun Wang 0007 |
CVPR | 7 |
| 2021 | Multi-cue based 3D residual network for action recognition
Ming Zong, Ruili Wang 0001, Zhe Chen 0004, Maoli Wang, Xun Wang 0007, Johan Potgieter |
Neural Comput. Appl. | 5 |
| 2021 | Feature Re-Learning with Data Augmentation for Video Relevance PredictionabstractPredicting the relevance between two given videos with respect to their visual content is a key component for content-based video recommendation and retrieval. Thanks to the increasing availability of pre-trained image and video convolutional neural network models, deep visual features are widely used for video content representation. However, as how two videos are relevant is task-dependent, such off-the-shelf features are not always optimal for all tasks. Moreover, due to varied concerns including copyright, privacy and security, one might have access to only pre-computed video features rather than original videos. We propose in this paper feature re-learning for improving video relevance prediction, with no need of revisiting the original video content. In particular, re-learning is realized by projecting a given deep feature into a new space by an affine transformation. We optimize the re-learning process by a novel negative-enhanced triplet ranking loss. In order to generate more training data, we propose a new data augmentation strategy which works directly on frame-level and video-level features. Extensive experiments in the context of the Hulu Content-based Video Relevance Prediction Challenge 2018 justify the effectiveness of the proposed method and its state-of-the-art performance for content-based video relevance prediction. Jianfeng Dong, Xun Wang 0007, Leimin Zhang, Chaoxi Xu, Gang Yang 0001, Xirong Li 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Discovering Cliques in Signed Networks Based on Balance Theory
Renjie Sun, Qiuyu Zhu 0002, Chen Chen 0017, Xiaoyang Wang 0002, Ying Zhang 0001, Xun Wang 0007 |
DASFAA (2) | 6 |
| 2020 | Cross-View Attention Network for Breast Cancer Screening from Multi-View MammogramsabstractIn this paper, we address the problem of breast caner detection from multi-view mammograms. We present a novel cross-view attention module (CvAM) which implicitly learns to focus on the cancer-related local abnormal regions and highlighting salient features by exploring cross-view information among four views of a screening mammography exam, e.g. asymmetries between left and right breasts and lesion correspondence between two views of the same breast. More specifically, the proposed CvAM calculates spatial attention maps based on the same view of different breasts to enhance bilateral asymmetric regions, and channel attention maps based on two different views of the same breast to enhance the feature channels corresponding to the same lesion in a single breast. CvAMs can be easily integrated into standard convolutional neural networks (CNN) architectures such as ResNet to form a multi-view classification model. Experiments are conducted on DDSM dataset, and results show that CvAMs can not only provide better classification accuracy over non-attention and single-view attention models, but also demonstrate better abnormality localization power using CNN visualization tools. Xuran Zhao, Luyang Yu, Xun Wang 0007 |
ICASSP | 3 |
| 2020 | Tree-Augmented Cross-Modal Encoding for Complex-Query Video RetrievalabstractThe rapid growth of user-generated videos on the Internet has intensified the need for text-based video retrieval systems. Traditional methods mainly favor the concept-based paradigm on retrieval with simple queries, which are usually ineffective for complex queries that carry far more complex semantics. Recently, embedding-based paradigm has emerged as a popular approach. It aims to map the queries and videos into a shared embedding space where semantically-similar texts and videos are much closer to each other. Despite its simplicity, it forgoes the exploitation of the syntactic structure of text queries, making it suboptimal to model the complex queries. Xun Yang 0001, Jianfeng Dong, Yixin Cao 0002, Xun Wang 0007, Meng Wang 0001, Tat-Seng Chua |
SIGIR | 4 |
| 2020 | Underwater salient object detection by combining 2D and 3D visual features
Zhe Chen 0004, Hongmin Gao 0001, Zhen Zhang 0019, Helen Zhou, Xun Wang 0007 |
Neurocomputing | 5 |
| 2020 | Rapido: Scaling blockchain with multi-path payment channels
Changting Lin, Xun Wang 0007, Jianhai Chen |
Neurocomputing | 3 |
| 2020 | Evaluating facial recognition web services with adversarial and synthetic samples
Xuran Zhao, Xun Wang 0007, Hexin Lv |
Neurocomputing | 3 |
| 2020 | Fast and parameter-light rare behavior detection in maritime trajectoriesabstractRare behaviors indicate important events and situations in maritime surveillance applications. State-of-the-art methods provide many effective solutions to detect anomalous behaviors. Meanwhile, most solutions are parameter-laden and too costly to identify useful rare behaviors with human knowledge in a visual analytics manner. This paper is concerned with a scheme cross trajectories, vessel attributes and the movement context for detecting rare behaviors through preprocessing, kNN-based clustering, and verification. Although the scheme involves several parameters, we demonstrate that they are able to be tackled in thresholds. As a result, a rare behavior factor is the single parameter that affect the detecting results. The proposed scheme is evaluated via a simulated data set for performance and a real life AIS data for effectiveness. Results show that high accuracy to labelled anomalies and useful rare behaviors can be achieved. Yifan Lei, Zhenguang Liu, Xun Wang 0007, Shouling Ji, Anthony K. H. Tung |
Inf. Process. Manag. | 4 |
| 2020 | Example-based image recoloring in an indoor environmentabstractAbstract Color structure of a home scene image closely relates to the material properties of its local regions. Existing color migration methods typically fail to fully infer the correlation between the coloring of local home scene regions, leading to a local blur problem. In this paper, we propose a color migration framework for home scene images. It picks the coloring from a template image and transforms such coloring to a home scene image through a simple interaction. Our framework comprises three main parts. First, we carry out an interactive segmentation to divide an image into local regions and extract their corresponding colors. Second, we generate a matching color table by sampling the template image according to the color structure of the original home scene image. Finally, we transform colors from the matching color table to the target home scene image with the boundary transition maintained. Experimental results show that our method can effectively transform the coloring of a scene matching with the color composition of a given natural or interior scenery. Xianxuan Lin, Xun Wang 0007, Frederick W. B. Li, Jinyu Li 0002, Bailin Yang, Tianxiang Wei |
Comput. Animat. Virtual Worlds | 2 |
| 2020 | Deep multi-person kinship matching and recognition for family photos
Mengyin Wang, Xiangbo Shu, Jiashi Feng, Xun Wang 0007, Jinhui Tang 0001 |
Pattern Recognit. | 4 |
| 2020 | Deep supervised feature selection for social relationship recognition
Mengyin Wang, Xiaoyu Du 0002, Xiangbo Shu, Xun Wang 0007, Jinhui Tang 0001 |
Pattern Recognit. Lett. | 4 |
| 2020 | Improving Multiperson Pose Estimation by Mask-aware Deep Reinforcement LearningabstractResearch on single-person pose estimation based on deep neural networks has recently witnessed progress in both accuracy and execution efficiency. However, multiperson pose estimation is still a challenging topic, partially because the object regions are selected greedily from proposals via class-agnostic nonmaximum suppression (NMS), and the misalignment in the redundant detection yields inaccurate human poses. Therefore, we consider how to obtain the optimal input in human pose estimation under conditions in which intermediate label information is not available. As supervised learning–based alignment does not generalize well to unseen samples in the human pose space, in this article, we present a mask-aware deep reinforcement learning approach to modify the detection result. We use mask information to remove the adverse effects from the cluttered background and to select the optimal action according to the revised reward function. We also propose a new regularization term to punish joints that are outside of the silhouette region in the human pose estimation stage. We evaluate our approach on the MPII Multiperson dataset and the MS-COCO Keypoints Challenge. The results show that our approach yields competing inference results when it is compared to the other state-of-the-art approaches. Xun Wang 0007, Xuran Zhao, Judith Gelernter, Guohua Cheng, Wei Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Dual Encoding for Zero-Example Video RetrievalabstractThis paper attacks the challenging problem of zero-example video retrieval. In such a retrieval paradigm, an end user searches for unlabeled videos by ad-hoc queries described in natural language text with no visual example provided. Given videos as sequences of frames and queries as sequences of words, an effective sequence-to-sequence cross-modal matching is required. The majority of existing methods are concept based, extracting relevant concepts from queries and videos and accordingly establishing associations between the two modalities. In contrast, this paper takes a concept-free approach, proposing a dual deep encoding network that encodes videos and queries into powerful dense representations of their own. Dual encoding is conceptually simple, practically effective and end-to-end. As experiments on three benchmarks, i.e. MSR-VTT, TRECVID 2016 and 2017 Ad-hoc Video Search show, the proposed solution establishes a new state-of-the-art for zero-example video retrieval. Jianfeng Dong, Xirong Li 0001, Chaoxi Xu, Shouling Ji, Yuan He 0011, Gang Yang 0001, Xun Wang 0007 |
CVPR | 7 |
| 2019 | Exploring Content-based Video Relevance for Video Click-Through Rate PredictionabstractThis paper describes our solution for the Hulu Challenge. To answer the challenge, we introduce two content-based models, namely, Cascading Mapping Network (CMN) and Relevant-Enhanced Deep Interest Network (REDIN). CMN predicts video Click-Through Rate (CTR) by predicting content-based video relevance. REDIN mainly improves the popular Deep Interest Network by adding explicit video relevance constraint, which provides guidance for low-level video feature learning thus helpful for CTR prediction. Based on the two models, our solution obtains Area Under Curve (AUC) score of 0.6022 and 0.6155 on the TV-shows and Movie track respectively. What is more, we are one of the only two teams giving scores of over 0.6 on both tracks. The results justify the effectiveness and stability of our proposed solution. Xun Wang 0007, Yali Du 0001, Leimin Zhang, Xirong Li 0001, Jianfeng Dong |
ACM Multimedia | 1 |
| 2019 | A Color-Pair Based Approach for Accurate Color Harmony EstimationabstractAbstract Harmonious color combinations can stimulate positive user emotional responses. However, a widely open research question is: how can we establish a robust and accurate color harmony measure for the public and professional designers to identify the harmony level of a color theme or color set. Building upon the key discovery that color pairs play an important role in harmony estimation, in this paper we present a novel color‐pair based estimation model to accurately measure the color harmony. It first takes a two‐layer maximum likelihood estimation (MLE) based method to compute an initial prediction of color harmony by statistically modeling the pair‐wise color preferences from existing datasets. Then, the initial scores are refined through a back‐propagation neural network (BPNN) with a variety of color features extracted in different color spaces, so that an accurate harmony estimation can be obtained at the end. Our extensive experiments, including performance comparisons of harmony estimation applications, show the advantages of our method in comparison with the state of the art methods. Bailin Yang, Tianxiang Wei, Xianyong Fang, Zhigang Deng 0001, Frederick W. B. Li, Xun Wang 0007 |
Comput. Graph. Forum | 7 |
| 2019 | A preliminary geometric structure simplification for Principal Component Analysis
Huamao Gu, Tong Lin 0002, Xun Wang 0007 |
Neurocomputing | 3 |
| 2019 | Multi-scale Hierarchical Residual Network for Dense CaptioningabstractRecent research on dense captioning based on the recurrent neural network and the convolutional neural network has made a great progress. However, mapping from an image feature space to a description space is a nonlinear and multimodel task, which makes it difficult for the current methods to get accurate results. In this paper, we put forward a novel approach for dense captioning based on hourglass-structured residual learning. Discriminant feature maps are obtained by incorporating dense connected networks and residual learning in our model. Finally, the performance of the approach on the Visual Genome V1.0 dataset and the region labelled MS-COCO (Microsoft Common Objects in Context) dataset are demonstrated. The experimental results have shown that our approach outperforms most current methods. Xun Wang 0007, Jiachen Wu, Ruili Wang 0001, Bailin Yang |
J. Artif. Intell. Res. | 2 |
| 2019 | Traffic Sign Detection Using a Multi-Scale Recurrent Attention NetworkabstractTraffic sign detection plays an important role in intelligent transportation systems. But traffic signs are still not well-detected by deep convolution neural network-based methods because the sizes of their feature maps are constrained, and the environmental context information has not been fully exploited by other researchers. What we need is a way to incorporate relevant context detail from the neighboring layers into the detection architecture. We have developed a novel traffic sign detection approach based on recurrent attention for multi-scale analysis and use of local context in the image. Experiments on the German traffic sign detection benchmark and the Tsinghua-Tencent 100K data set demonstrated that our approach obtained an accuracy comparable to the state-of-the-art approaches in traffic sign detection. Judith Gelernter, Xun Wang 0007, Jianyuan Li, Yizhou Yu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Emotion information visualization through learning of 3D morphable face model
Hai Jin 0002, Xun Wang 0007, Yuanfeng Lian, Jing Hua 0001 |
Vis. Comput. | 2 |
| 2018 | Feature Re-Learning with Data Augmentation for Content-based Video RecommendationabstractThis paper describes our solution for the Hulu Content-based Video Relevance Prediction Challenge. Noting the deficiency of the original features, we propose feature re-learning to improve video relevance prediction. To generate more training instances for supervised learning, we develop two data augmentation strategies, one for frame-level features and the other for video-level features. In addition, late fusion of multiple models is employed to further boost the performance. Evaluation conducted by the organizers shows that our best run outperforms the Hulu baseline, obtaining relative improvements of 26.2% and 30.2% on the TV-shows track and the Movies track, respectively, in terms of [email protected] The results clearly justify the effectiveness of the proposed solution. Jianfeng Dong, Xirong Li 0001, Chaoxi Xu, Gang Yang 0001, Xun Wang 0007 |
ACM Multimedia | 5 |
| 2018 | Lane marking detection via deep convolutional neural network
Judith Gelernter, Xun Wang 0007, Weigang Chen, Junxiang Gao |
Neurocomputing | 3 |
| 2018 | Review on mining data from multiple data sources
Ruili Wang 0001, Wanting Ji, Mingzhe Liu 0001, Xun Wang 0007, Jian Weng 0001, Song Deng, Suying Gao, Chang-an Yuan 0001 |
Pattern Recognit. Lett. | 4 |
| 2017 | Robust 3D face modeling and reconstruction from frontal and side images
Hai Jin 0002, Xun Wang 0007, Zichun Zhong, Jing Hua 0001 |
Comput. Aided Geom. Des. | 2 |
| 2017 | Object localization via evaluation multi-task learning
Huiyan Wang 0002, Xun Wang 0007 |
Neurocomputing | 3 |
| 2017 | Temporally Consistent Depth Map Prediction Using Deep Convolutional Neural Network and Spatial-Temporal Conditional Random Field
Xuran Zhao, Xun Wang 0007, Qi-Chao Chen |
J. Comput. Sci. Technol. | 2 |
| 2017 | Sky detection- and texture smoothing-based high-visibility haze removal from images and videosabstractAbstract To address the gloomy sky and the low contrast caused by the left fog in the existing image dehazing methods, we propose a robust haze removal algorithm for images and videos. First, a sky detection‐based adaptive atmospheric light estimation method is designed for brighter and cleaner restoration results for the sky regions. Second, in order to reconstruct a transmission map in line with the depth variation, we preprocess the input image with texture smoothing to keep the color consistency inside the same planar object and devise a texture smoothing‐based robust transmission estimation method, with which the contrast and color saturation of fog‐free image are greatly promoted. Finally, the restored results are post‐processed with the joint bilateral filter for the purpose of noise removal. What's more, a guided filter‐based temporally coherent atmospheric light smoothing strategy and a Gaussian filter‐based spatial‐temporally coherent transmission smoothing strategy are put forward for video dehazing, which can ensure the spatial as well as temporal continuity of the haze‐free videos. Experimental results show that the recovered haze‐free images and videos have high contrast and color saturation with cleaner sky regions, and the haze‐free videos are free of jittering and flickering phenomena. Yiyun Shen, Yaqi Shao, Jinwei Zhao, Xun Wang 0007 |
Comput. Animat. Virtual Worlds | 5 |
| 2017 | Diffusion map based interactive image segmentation
Xun Wang 0007, Jianqiu Jin, Bailin Yang |
Multim. Tools Appl. | 1 |
| 2017 | Multi-view dimensionality reduction via subspace structure agreement
Xuran Zhao, Xun Wang 0007, Huiyan Wang 0002 |
Multim. Tools Appl. | 2 |
| 2017 | Pedestrian recognition in multi-camera networks using multilevel important salient feature and multicategory incremental learning
Huiyan Wang 0002, Yixiang Yan, Jing Hua 0001, Yutao Yang, Xun Wang 0007, John R. Deller Jr., Guofeng Zhang 0001, Hujun Bao |
Pattern Recognit. | 5 |
| 2017 | Semantic annotation for complex video street views based on 2D-3D multi-feature fusion and aggregated boosting decision forests
Xun Wang 0007, Guoli Yan, Huiyan Wang 0002, Jianhai Fu, Jing Hua 0001, Yutao Yang, Guofeng Zhang 0001, Hujun Bao |
Pattern Recognit. | 1 |
| 2017 | Manifold Learning by Curved Cosine MappingabstractIn the field of pattern recognition, data analysis, and machine learning, data points are usually modeled as high-dimensional vectors. Due to the curse-of-dimensionality, it is non-trivial to efficiently process the orginal data directly. Given the unique properties of nonlinear dimensionality reduction techniques, nonlinear learning methods are widely adopted to reduce the dimension of data. However, existing nonlinear learning methods fail in many real applications because of the too-strict requirements (for real data) or the difficulty in parameters tuning. Therefore, in this paper, we investigate the manifold learning methods which belong to the family of nonlinear dimensionality reduction methods. Specifically, we proposed a new manifold learning principle for dimensionality reduction named Curved Cosine Mapping (CCM). Based on the law of cosines in Euclidean space, CCM applies a brand new mapping pattern to manifold learning. In CCM, the nonlinear geometric relationships are obtained by utlizing the law of cosines, and then quantified as the dimensionality-reduced features. Compared with the existing approaches, the model has weaker theoretical assumptions over the input data. Moreover, to further reduce the computation cost, an optimized version of CCM is developed. Finally, we conduct extensive experiments over both artificial and real-world datasets to demonstrate the performance of proposed techniques. Huamao Gu, Xun Wang 0007, Xue-wen Chen 0001, Shaoping Deng, Jin-Qin Shi |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | Multi-scale inherent variation features-based texture filtering
Huan Shao 0003, Yanggang Zhou, Yaqi Shao, Xun Wang 0007 |
Vis. Comput. | 6 |
| 2016 | Topology-aware moving least square deformation for 2D charactersabstractAbstract Deformation method based on moving least squares (MLS) allows the user to manipulate 2D characters using either sets of points or line segments in real time. However, the traditional MLS deformation spreads the deformation of the controls with respect to the spatial distance, but oblivious to the shape topology, which would possibly lead to distortion. In this paper, we present a topology‐aware MLS deformation approach for 2D characters. First, a Laplace equation is solved to obtain a set of weights, which are called harmonic weights. Then, the MLS deformation is performed by using the harmonic weights as the deformation influence of the user‐specified controls. Finally, the possible distortion in the traditional MLS deformation can be effectively avoided, as the harmonic weights spread the deformation of the controls in a localized and topology‐aware way. In addition, a simple but effective area‐preserving variant of MLS deformation is proposed, which is suitable for the editing of incompressible objects. Copyright © 2015 John Wiley & Sons, Ltd. Xun Wang 0007, Wenwu Yang, Wangbin Kou, Bailin Yang |
Comput. Animat. Virtual Worlds | 1 |
| 2016 | Texture filtering based physically plausible image dehazing
Jinwei Zhao, Yiyun Shen, Yanggang Zhou, Xun Wang 0007 |
Vis. Comput. | 5 |
| 2016 | Visual saliency guided textured model simplification
Bailin Yang, Frederick W. B. Li, Xun Wang 0007, Mingliang Xu 0001, Xiaohui Liang 0001, Zhaoyi Jiang, Yanhui Jiang |
Vis. Comput. | 3 |
| 2015 | Parallel multi-level 2D-DWT on CUDA GPUs and its application in ring artifact removalabstractSummary This paper presented two schemes of parallel 2D discrete wavelet transform (DWT) on Compute Unified Device Architecture graphics processing units. For the first scheme, the image and filter are transformed to spectral domain by using Fast Fourier Transformation (FFT), multiplied and then transformed back to space domain by using inverse FFT. For the second scheme, the image pixels are convolved directly with filters. Because there is no data relevance, the convolution for data points on different positions could be executed concurrently. To reduce data transfer, the boundary extension and down‐sampling are processed during data loading stage, and transposing is completed implicitly during data storage. A similar skill is adopted when parallelizing inverse 2D DWT. To further speed up the data access, the filter coefficients are stored in the constant memory. We have parallelized the 2D DWT for dozens of wavelet types and achieved a speedup factor of over 380 times compared with that of its CPU version. We applied the parallel 2D DWT in a ring artifact removal procedure; the executing speed was accelerated near 200 times compared with its CPU version. The experimental results showed that the proposed parallel 2D DWT on graphics processing units can significantly improve the performance for a wide variety of wavelet types and is promising for various applications. Copyright © 2015 John Wiley & Sons, Ltd. Leqing Zhu, Daxing Zhang, Dadong Wang, Huiyan Wang 0002, Xun Wang 0007 |
Concurr. Comput. Pract. Exp. | 6 |
| 2015 | Recognition of Low-Resolution Logos in Vehicle Images Based on Statistical Random Sparse DistributionabstractTraditional image recognition approaches can achieve high performance only when the images have high resolution and superior quality. A new vehicle logo recognition (VLR) method is proposed to treat low-resolution and poor-quality images captured from urban crossings in intelligent transport system, and the proposed approach is based on statistical random sparse distribution (SRSD) feature and multiscale scanning. The SRSD feature is a novel feature representation strategy that uses the correlation between random sparsely sampled pixel pairs as an image feature and describes the distribution of a grayscale image statistically. Multiscale scanning is a creative classification algorithm that locates and classifies a logo integrally, which alleviates the effect of propagation errors in traditional methods by processing the location and classification separately. Experiments show an overall recognition rate of 97.21% for a set of 3370 vehicle images, which showed that the proposed algorithm outperforms classical VLR methods for low-resolution and inferior quality images and is very suitable for on-site supervision in ITSs. Haoyu Peng, Xun Wang 0007, Huiyan Wang 0002, Wenwu Yang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2014 | Fast corotational simulation for example-driven deformation
Chao Song 0001, Hongxin Zhang 0001, Xun Wang 0007, Jianwei Han, Huiyan Wang 0002 |
Comput. Graph. | 3 |
| 2014 | 3D geometry-dependent texture map compression with a hybrid ROI coding
Bailin Yang, Jianqiu Jing, Xun Wang 0007, Jianwei Han |
Sci. China Inf. Sci. | 3 |
| 2014 | Video object matching across multiple non-overlapping camera views based on multi-feature fusion and incremental learning
Huiyan Wang 0002, Xun Wang 0007, Jia Zheng 0008, John R. Deller Jr., Haoyu Peng, Leqing Zhu, Weigang Chen, Riji Liu, Hujun Bao |
Pattern Recognit. | 2 |
| 2014 | Part-to-part morphing for planar curves
Wenwu Yang, Xun Wang 0007 |
Vis. Comput. | 2 |
| 2013 | Relighting abstracted image via salient edge-guided luminance field optimizationabstractABSTRACT Because existing image abstraction systems can hardly incorporate with the changing light, we present an integrated image abstraction and relighting rendering system, which is based on a salient edge‐guided luminance field optimization approach. For an input image, we first adopt a sparsity prior‐based illumination decomposition method to remove its original illumination and have an intrinsic image. Meanwhile, we iteratively extract the salient edge inside by employing a message‐passing strategy. Then, we simplify this image with a proposed salient edge‐guided image abstraction optimization algorithm in the luminance field. Finally, we put forward a salient edge‐guided image relighting optimization method to simulate the effect of dynamic lighting along different directions. Experiment results show that our system can artistically adjust the illumination of the abstracted image and makes it more vivid. Copyright © 2013 John Wiley & Sons, Ltd. Qunsheng Peng 0001, Xun Wang 0007, Enhua Wu |
Comput. Animat. Virtual Worlds | 4 |
| 2013 | Shape-aware skeletal deformation for 2D characters
Xun Wang 0007, Wenwu Yang, Haoyu Peng |
Vis. Comput. | 1 |
| 2012 | Structure Preserving Manipulation and Interpolation for Multi-element 2D ShapesabstractAbstract This paper presents a method that generates natural and intuitive deformations via direct manipulation and smooth interpolation for multi‐element 2D shapes. Observing that the structural relationships between different parts of a multi‐element 2D shape are important for capturing its feature semantics, we introduce a simple structure called a feature frame to represent such relationships. A constrained optimization is solved for shape manipulation to find optimal deformed shapes under user‐specified handle constraints. Based on the feature frame, local feature preservation and structural relationship maintenance are directly encoded into the objective function. Beyond deforming a given multi‐element 2D shape into a new one at each key frame, our method can automatically generate a sequence of natural intermediate deformations by interpolating the shapes between the key frames. The method is computationally efficient, allowing real‐time manipulation and interpolation, as well as generating natural and visually plausible results. Wenwu Yang, Jieqing Feng, Xun Wang 0007 |
Comput. Graph. Forum | 3 |
| 2012 | Adaptive tone-preserved image detail enhancement
Caiping Yan, Xun Wang 0007 |
Vis. Comput. | 4 |
| 2008 | An Effective Error Resilient Packetization Scheme for Progressive Mesh Transmission over Unreliable Networks
Bailin Yang, Frederick W. B. Li, Xun Wang 0007 |
J. Comput. Sci. Technol. | 4 |
| 2007 | An Edge Detection Algorithm Based on Improved CANNY OperatorabstractEdge detection is an important aspect for image processing, and primary step of spatial data extraction in geography information system. For contour detection, this paper proposed the improved template algorithm, which is not only including the gradient directions of X and Y, but also the first order partial finite differences of directions 45 and 135 degree in calculating the amplitude values. These mostly improved the calculation accuracy of the amplitude values. In the non-maxima suppression process, the factor ratio of four quadrants of linear interpolation is improved to achieve better detection results. Experiments showed that this improved CANNY algorithm has better noise suppression and edge continuity. Xun Wang 0007, Jianqiu Jin |
ISDA | 1 |