Lin Wu 0001

dblp:65/6292-1 · also Lin Yuanbo Wu · DBLP profile ↗
← Back
92ranked-venue papers
26as first author
43since 2021 · last 2026
0000-0001-6119-058XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 14 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 11 first-author · 25 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
abstract
Multimodal Misinformation Detection (MMD) refers to the task of detecting social media posts involving misinformation, where the post often contains text and image modalities. However, by observing the MMD posts, we hold that the text modality may be much more informative than the image modality because the text generally describes the whole event/story of the current post but the image often presents partial scenes only. Our preliminary empirical results indicate that the image modality exactly contributes less to MMD. Upon this idea, we propose a new MMD method named RETSIMD. Specifically, we suppose that each text can be divided into several segments, and each text segment describes a partial scene that can be presented by an image. Accordingly, we split the text into a sequence of segments, and feed these segments into a pre-trained text-to-image generator to augment a sequence of images. We further incorporate two auxiliary objectives concerning text-image and image-label mutual information, and further post-train the generator over an auxiliary text-to-image generation benchmark dataset. Additionally, we propose a graph structure by defining three heuristic relationships between images, and use a graph neural network to generate the fused features. Extensive empirical results validate the effectiveness of RETSIMD.
Bing Wang 0018, Ximing Li 0002, Changchun Li, Lin Wu 0001, Buyu Wang, Sheng-Sheng Wang 0001
AAAI5
2026 Geometry-Aware Noisy Correspondence Mitigation for Cross-Modal Text-Based Person Retrieval
abstract
Text-Based Person Retrieval (TBPR) aims to accurately retrieve target individuals from large-scale image databases using only textual descriptions. Existing methods typically assume a ground-truth correspondence between text and images (i.e., strongly correlated). However, in real-world scenarios, this assumption may not be able to hold for the cross-modal matching due to weak or even corrupted correlations between textual descriptions and visual content, referred to as noisy correspondence (NC). Such NC largely disrupts the correspondence learning between visual and semantic modalities. Though prior works have improved single-modal robustness against noisy labels, systematic modeling of both cross-modal and intra-modal geometric structures in TBPR remains limited attention. In this paper, we propose Geometric Structure Consistency Alignment (GSCA) to TBPR, which leverages cross-modal cosine similarity and intra-modal nearest-neighbor affinity to learn visual-semantic consistency under noisy correspondence. To mitigate the structural corruption caused by noisy pairs, we introduce the Structure Refinement and Mining (SRAM) module. By partitioning training data into clean, ambiguous, and noisy subsets, SRAM enables the model to strategically refine the cross-modal correspondence by mining reliable pairs, thus enhancing the reliability of positive or negative samples discrimination and preserving structural consistency across modalities. Extensive experiments demonstrate that our method achieves state-of-the-art performance across three public datasets. On CUHK-PEDES, it boosts Rank-1 by 1.42% in noise-free conditions, sustaining a robust 74.25% Rank-1 under a 50% noise ratio.
Xinpan Yuan, Shaomin Xie, Liujie Hua, Chengyuan Zhang 0001, Guihu Zhao, Lin Wu 0001
AAAI6
2026 Cooperative meta-learning for incremental few-shot object detection in open urban environments
Chengyuan Zhang 0001, Lin Wu 0001
Pattern Recognit.5
2026 Beyond consistency: Preserving temporal structure in zero-shot video editing
Deyin Liu, Yisheng Ding, Zhe Jin 0001, Xiatian Zhu, Anjan Dutta 0001, Lin Wu 0001
Pattern Recognit.6
2026 HAViG: Hierarchical adaptive visual grounding framework for video question answering
Lei Zhu 0005, Lingmin Pan, Siqiao Tan, Chengyuan Zhang 0001, Deyin Liu, Lin Wu 0001, Farid Boussaïd, Mohammed Bennamoun
Pattern Recognit.6
2026 Towards efficient pixel labeling for industrial anomaly detection and localization
Jingqi Wu, Lin Wu 0001, Hao Chen 0041, Deyin Liu, Haiqiang Jin
Pattern Recognit. Lett.3
2026 Accurate Industrial Anomaly Detection and Localization Using Weakly-Supervised Residual Transformers
abstract
Recent advancements in industrial anomaly detection (AD) have demonstrated that incorporating a small number of anomalous samples during training can significantly enhance accuracy. However, this improvement often comes at the cost of extensive annotation efforts, which are impractical for many real-world applications. In this paper, we introduce a novel framework, "Weakly-supervised RESidual $T$ ransformer" (WeakREST), designed to achieve high anomaly detection accuracy while minimizing the reliance on manual annotations. First, we reformulate the pixel-wise anomaly localization task into a block-wise classification problem. Second, we introduce a residual-based feature representation called "Positional $F$ ast $A$ nomaly $R$ esiduals" (PosFAR) which captures anomalous patterns more effectively. To leverage this feature, we adapt the Swin Transformer for enhanced anomaly detection and localization. Additionally, we propose a weak annotation approach utilizing bounding boxes and image tags to define anomalous regions. This approach establishes a semi-supervised learning context that reduces the dependency on precise pixel-level labels. To further improve the learning process, we develop a novel ResMixMatch algorithm, capable of handling the interplay between weak labels and residual-based representations. On the benchmark dataset MVTec-AD, our method achieves an Average Precision (AP) of 83.0%, surpassing the previous best result of 82.7% in the unsupervised setting. In the supervised AD setting, WeakREST attains an AP of 87.6%, outperforming the previous best of 86.0%. Notably, even when using weaker annotations such as bounding boxes, WeakREST exceeds the performance of leading methods relying on pixel-wise supervision, achieving an AP of 87.1% compared to the prior best of 86.0% on MVTec-AD. This superior performance is consistently replicated across other well-established AD datasets, including MVTec 3D, KSDD2 and Real-IAD. Code is available at: https://github.com/BeJane/Semi_REST.
Jingqi Wu, Deyin Liu, Lin Wu 0001, Hao Chen 0041, Chunhua Shen
IEEE Trans. Image Process.4
2026 Detecting Misinformation by Uncovering Commonsense Conflicts With LLM Workflows
abstract
The advancement of Internet technology has spurred a rise in the dissemination of misinformation, which has had profoundly negative impacts across a wide array of fields. To address this issue, the field of Misinformation Detection (MD), which focuses on the automated identification of online misinformation, has gained significant traction among researchers. In our study, we introduce an innovative plugand- play augmentation technique for MD, termed DEtecting Misinformation by Uncovering Commonsense Conflict (DEMUC). Our approach is grounded in previous psychological research that suggests that fake content often contains commonsense. Accordingly, we develop commonsense expressions for articles to highlight potential conflicts between the inferred commonsense triplets and the established ones derived from reliable commonsense reasoning tools. According to the used tools, we induce two variants DEMUC-KLM using the knowledge language model COMET and DEMUC-LLM using the large language models. These generated expressions are then applied as augmentations to each article, enabling any MD method to be trained on these augmented datasets. Additionally, we have manually compiled a new dataset CoMis, which consists exclusively of fake articles characterized by commonsense conflicts. By integrating DEMUC with various existing MD frameworks and evaluating them on four public benchmark datasets and CoMis, our empirical findings show that both DEMUC-KLM and DEMUC-LLM consistently and significantly outperform current MD baselines, while also generating precise commonsense expressions.
Bing Wang 0018, Ximing Li 0002, Changchun Li, Bingrui Zhao 0001, Renchu Guan, Lin Wu 0001, Jungong Han
IEEE Trans. Knowl. Data Eng.6
2026 Multi-View Aligned Clustering via Sample-Bundled Optimization: Anchor Graph Enhancement and Contrastive Propagation
abstract
Multi-view representation is powerful to capture the complex characteristics of real-world data by integrating complementary information from various modalities. However, in many use cases, such as boiler combustion monitoring, factors including sensor sampling frequency, equipment malfunctions, and network delays can lead to temporal asynchrony in data collection. This asynchrony leads misaligned multi-modal data, furthering the difficulty of learning optimal fused representation. To address this misalignment in multi-view data, a body of methods based on autoencoders and non-negative matrix factorization (NMF) have been presented. However, those methods are incapable of jointly exploring the underlying structure inherent in each view as well as the semantic consistency and structural similarity across views. To these ends, we propose a novel sample-bundled optimization for multi-view aligned clustering, which is based on Enhanced Anchor Graph and Contrastive Propagation (termed asEAGCP). We state that there is semantic consistency among intra-class samples (with the same and cross views) and that the global structures across different views demonstrate similarity. By leveraging these associations, we introduce an enhanced anchor graph with learnable sample correlation and a contrastive graph with feature propagation. Specifically, each anchor graphs preserves the semantic relationship among same-view samples, while the contrastive graph propagates feature information across multi-view samples. Experimental results demonstrate the superiority of our method on benchmark datasets by validating its effectiveness in aligning and clustering multi-view data.
Shubin Ma, Zhikui Chen, Lin Wu 0001, Liang Zhao 0005
IEEE Trans. Multim.3
2026 Multi-Modal Refined Prompting for Advancing Knowledge-Based Visual Question Answering
abstract
Knowledge-based Visual Question Answering (KB-VQA) has surfaced as a critical task in advancing AI capabilities. Despite significant progress enabled by large language models (LLMs), there are still three major challenges: (1) flawed image captions cause unreliable reasoning; (2) noisy explicit knowledge can disrupts answering; and (3) massive LLMs scale is irreplaceable to robustness. To overcome these challenges, we develop a novel approach, Multi-Modal Refined Prompting (MMRP), which generates high-quality prompts tailored for LLMs. To tackle the first challenge, a multi-faceted image captioning strategy is employed to generate detailed, contextually relevant visual descriptions. In addition, we introduce a complementary knowledge retrieval and refinement strategy to deliver concise, contextually relevant knowledge, effectively overcoming the second challenge. These enhanced image captions and explicit knowledge are then integrated into a knowledge-infused in-context prompt, effectively activating the reasoning capabilities of LLMs. Importantly, MMRP eliminates reliance on massive LLMs and avoids the need for model fine-tuning, while achieving significant improvements in answer accuracy. Extensive evaluations on the widely-used OK-VQA benchmark against 22 baselines prove the superiority of MMRP, establishing a new state-of-the-art in KB-VQA.
Lei Zhu 0005, Mengxi Ying, Chengyuan Zhang 0001, Deyin Liu, Lin Wu 0001, Shichao Zhang 0001, Xuelong Li 0001
IEEE Trans. Multim.5
2025 Utterance-level Emotion Recognition in Conversation with Conversation-level Supervision
abstract
Emotion Recognition in Conversations (ERC) involves automatically identifying the emotion of each utterance in conversations. The emotion of an utterance is contingent to the conversation context, and thus, annotating each utterance in ERC entails repetitive screening the whole conversation from annotators. Such a requirement leads to prohibitive cost in fine-grained labeling on utterance. In this paper, we propose an efficient coarse-grained labeling strategy for ERC, which assigns a set of emotions for each conversation. In specific, we reformulate the ERC predictors with conversation-level emotion sets as weakly-supervised learning to optimise a potential candidate for ERC, which is termed as Dataless ERC (DERC). To validate this, we propose a simple-yet-flexible DERC framework with Progressive Learning (DERC-PL). We jointly update pseudo-utterance-level emotions and the ERC predictor in a self-training manner, where we progressively update the ERC predictor from training subsets with lower noise densities to the ones with higher noise densities. We implemented several versions of \baby by incorporating various off-the-shelf ERC methods. Extensive experimental results demonstrate that the proposed \baby can be on par with existing weakly-supervised learning baselines and supervised learning ERC methods.
Ximing Li 0002, Yuanchao Dai, Zhiyao Yang, Jinjin Chi, Wanfu Gao, Lin Wu 0001
AAAI6
2025 PBD: A Manually Curated Full-Chain Benchmark Dataset for Evaluating LLMs on ACMG PS3/BS3 Functional Evidence Acquisition
abstract
The PS3 (Pathogenic Functional Evidence) and BS3 (Benign Functional Evidence) criteria in the American College of Medical Genetics and Genomics (ACMG) guidelines are critical for genetic variant classification. However, manual evaluation is time-consuming and prone to inter-laboratory inconsistencies, limiting the clinical interpretation of Variants of Uncertain Significance (VUS). Large Language Models (LLMs) offer potential for automated assessment, but their performance validation is hindered by the lack of standardized, high-quality datasets. This study introduces the PS3/BS3 Full-Chain Evidence Benchmark Dataset (PBD), the first manually curated dataset comprising 77 peer-reviewed publications, covering 266 cDNA variants (including duplicates) and their PS3/BS3 rating results. Adhering to ClinGen Sequence Variant Interpretation (SVI) standards, PBD includes structured, comprehensive evidence chains spanning genes, diseases, variants, experiments, and evidence ratings, designed to evaluate LLM capabilities in functional evidence extraction. We detail the dataset construction process, including literature screening, data extraction, standardization, and quality control, and developed a Python-based automated evaluation pipeline for reproducible, standardized analysis. Experiments using DeepSeek models$(1.5 \mathrm{b} / 7 \mathrm{b} / 14 \mathrm{b})$demonstrate PBD's potential in supporting automated variant interpretation. PBD provides a vital resource for bioinformatics and precision medicine, facilitating the development and standardization of variant classification tools. Data examples are available at https://github.com/User8588/PBD, with full data and code to be released upon paper acceptance.
Xinpan Yuan, Bozhao Li, Chenbin Liu, Xinxue Li, Liujie Hua, Jinchen Li, Lin Wu 0001, Guihu Zhao
BIBM7
2025 Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding
abstract
Large Vision-Language Models (LVLMs) have obtained impressive performance in visual content understanding and multi-modal reasoning. Unfortunately, these large models suffer from serious hallucination problems and tend to generate fabricated responses. Recently, several Contrastive Decoding (CD) strategies have been proposed to alleviate hallucination by introducing disturbed inputs. Although great progress has been made, these CD strategies mostly apply a one-size-fits-all approach for all input conditions. In this paper, we revisit this process through extensive experiments. Related results show that hallucination causes are hybrid and each generative step faces a unique hallucination challenge. Leveraging these meaningful insights, we introduce a simple yet effective Octopus-Like framework that enables the model to adaptively identify hallucination types and create a dynamic CD workflow. Our Octopus framework not only outperforms existing methods across four benchmarks but also demonstrates excellent deployability and expansibility. Code is available at https://github.com/LijunZhang01/Octopus.
Wei Suo, Mengyang Sun, Lin Wu 0001, Peng Wang 0015
CVPR4
2025 Unlocking Generalization Power in LiDAR Point Cloud Registration
abstract
In real-world environments, a LiDAR point cloud registration method with robust generalization capabilities (across varying distances and datasets) is crucial for ensuring safety in autonomous driving and other LiDAR-based applications. However, current methods fall short in achieving this level of generalization. To address these limitations, we propose UGP, a pruned framework designed to enhance generalization power for LiDAR point cloud registration. The core insight in UGP is the elimination of cross-attention mechanisms to improve generalization, allowing the network to concentrate on intra-frame feature extraction. Additionally, we introduce a progressive self-attention module to reduce ambiguity in large-scale scenes and integrate Bird’s Eye View (BEV) features to incorporate semantic information about scene elements. Together, these enhancements significantly boost the network’s generalization performance. We validated our approach through various generalization experiments in multiple outdoor scenes. In cross-distance generalization experiments on KITTI and nuScenes, UGP achieved state-of-the-art mean Registration Recall rates of 94.5% and 91.4%, respectively. In cross-dataset generalization from nuScenes to KITTI, UGP achieved a state-of-the-art mean Registration Recall of 90.9%. Code will be available at https://github.com/peakpang/UGP
Zhenxuan Zeng, Qiao Wu, Xiyu Zhang 0001, Lin Wu 0001, Pei An, Jiaqi Yang 0002, Peng Wang 0015
CVPR4
2025 Towards Robust Category-level Articulation Pose Estimation via Integrated Differentiable Rendering
abstract
Accurate object pose estimation is crucial for embodied intelligence tasks such as manipulation, grasping, and human-robot interaction. However, due to the inherent characteristics of articulated objects, such as kinematic constraints and self-occlusion, pose estimation for articulated objects has remained a significant challenge. To address these issues, this paper proposes CAPED, an end-to-end robust Category-level Articulated object Pose Estimator integrated differentiable rendering. Given partial point cloud as input, CAPED outputs the per-part 6D pose for articulation. Specifically, with the proposed joint-centric modeling manner, CAPED firstly estimates the pose for the free part. Afterward, we canonicalize the input point cloud to estimate constrained parts’ poses by predicting the joint parameters and states as replacements. For further refinement, we propose a differentiable rendering scheme for pose optimization. Evaluations of the ArtImage and RobotArm datasets demonstrate that CAPED exhibits outstanding effectiveness and generalization in tasks ranging from synthetic data to real-world scenarios. We will publicly release the code.
Li Zhang 0104, Yukang Huo, Lin Wu 0001, Yanyan Wei, Harshal Suresh Shende, Liu Liu 0012, Linlin Ou
ICASSP5
2025 Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language Models
abstract
Although Large Vision-Language Models (LVLMs) have achieved impressive results, their high computational costs pose a significant barrier to wide application. To enhance inference efficiency, most existing approaches can be categorized as parameter-dependent or token-dependent strategies to reduce computational demands. However, parameter-dependent methods require retraining LVLMs to recover performance while token-dependent strategies struggle to consistently select the most relevant tokens. In this paper, we systematically analyze the above challenges and provide a series of valuable insights for inference acceleration. Based on these findings, we propose a novel framework, the Pruning All-Rounder (PAR). Different from previous works, PAR develops a meta-router to adaptively organize pruning flows across both tokens and layers. With a self-supervised learning manner, our method achieves a superior balance between performance and efficiency. Notably, PAR is highly flexible, offering multiple pruning versions to address a range of acceleration scenarios. The code for this work is publicly available at https://github.com/ASGO-MM/Pruning-All-Rounder.
Wei Suo, Ji Ma 0008, Mengyang Sun, Lin Wu 0001, Peng Wang 0015, Yanning Zhang 0001
ICCV4
2025 Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
abstract
Nowadays, misinformation articles, especially multimodal ones, are widely spread on social media platforms and cause serious negative effects. To control their propagation, Multimodal Misinformation Detection (MMD) becomes an active topic in the community to automatically identify misinformation. Previous MMD methods focus on supervising detectors by collecting offline data. However, in real-world scenarios, new events always continually emerge, making MMD models trained on offline data consistently outdated and ineffective. To address this issue, training MMD models under online data streams is an alternative, inducing an emerging task named continual MMD. Unfortunately, it is hindered by two major challenges. First, training on new data consistently decreases the detection performance on past data, named past knowledge forgetting. Second, the social environment constantly evolves over time, affecting the generalization on future data. To alleviate these challenges, we propose to remember past knowledge by isolating interference between event-specific parameters with a Dirichlet process-based mixture-of-expert structure, and anticipate future environmental distributions by learning a continuous-time dynamics model. Accordingly, we induce a new continual MMD method DAEDCMD. Extensive experiments demonstrate that DAEDCMD can consistently and significantly outperform the compared methods, including six MMD baselines and three continual learning methods.
Bing Wang 0018, Ximing Li 0002, Mengzhe Ye, Changchun Li, Bo Fu 0001, Jianfeng Qu, Lin Wu 0001
ACM Multimedia7
2025 Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models
Mingyu Fu, Wei Suo, Ji Ma 0008, Lin Wu 0001, Peng Wang 0015, Yanning Zhang 0001
ACM Multimedia4
2025 EntroFormer: An entropy-based sparse vision transformer for real-time semantic segmentation
Song Wang 0008, Lin Wu 0001, Deyin Liu, Lei Gao 0001, Lin Qi 0001, Guanghui Wang 0001
Comput. Vis. Image Underst.3
2025 T-Person-GAN: Text-to-Person image generation with identity-consistency and manifold mix-up
Deyin Liu, Lin Wu 0001, Bo Li 0090, Ye Zhao 0001, ZongYuan Ge
Expert Syst. Appl.2
2025 Textual semantics enhancement adversarial hashing for cross-modal retrieval
Lei Zhu 0005, Runbing Wu, Deyin Liu, Chengyuan Zhang 0001, Lin Wu 0001, Ying Zhang 0001, Shichao Zhang 0001
Knowl. Based Syst.5
2025 Information bottleneck-guided KNN contrastive hashing for unsupervised cross-modal retrieval
abstract
Unsupervised cross-modal hashing (UCMH) has emerged as a promising solution for scalable multi-modal retrieval without costly annotations. However, existing methods often rely on rigid pairwise contrastive learning and fixed-size neighborhood selection, which suffer from false negatives and semantic noise, respectively—limiting their ability to model complex semantic structures in open-world scenarios. In this paper, we propose a novel framework, I nformation B ottleneck-guided K NN C ontrastive H ashing ( IBKCH ), which introduces a flexible and semantically adaptive contrastive paradigm for UCMH. Specifically, we design an information-aware neighbor sampling strategy that integrates: (1) a Hard-negative and Soft-positive (HN-SP) mechanism to adaptively distinguish informative negatives and softly aggregate latent positives; (2) an information bottleneck loss to retain task-relevant semantics while suppressing redundancy; and (3) an entropy sparsity regularizer to mitigate noisy neighbor interference. Furthermore, we develop an adaptive KNN contrastive learning scheme that unifies intra-modal and inter-modal alignment, enabling robust and discriminative hash code learning. Extensive experiments on three benchmark datasets demonstrate that IBKCH consistently outperforms state-of-the-art methods, especially under noisy or semantically diverse conditions—highlighting its effectiveness and generalizability in real-world UCMH applications.
Lei Zhu 0005, Zhengchang Yuan, Zeqian Yi, Chengyuan Zhang 0001, Lin Wu 0001, Ying Zhang 0001, Farid Boussaïd, Mohammed Bennamoun, Shichao Zhang 0001
Knowl. Based Syst.5
2025 Bi-Direction Label-Guided Semantic Enhancement for Cross-Modal Hashing
abstract
Supervised cross-modal hashing has gained significant attention due to its efficiency in reducing storage and computation costs while maintaining rich semantic information. Despite substantial progress in generating compact binary codes, two key challenges remain: (1) insufficient utilization of labels to mine and fuse multi-grained semantic information, and (2) unreliable cross-modal interaction, which does not fully leverage multi-grained semantics or accurately capture sample relationships. To address these limitations, we propose a novel method called Bi-direction Label-Guided Semantic Enhancement for cross-modal Hashing (BiLGSEH). To tackle the first challenge, we introduce a label-guided semantic fusion strategy that extracts and integrates multi-grained semantic features guided by multi-labels. For the second challenge, we propose a semantic-enhanced relation aggregation strategy that constructs and aggregates multi-modal relational information through bi-directional similarity. Additionally, we incorporate CLIP features to improve the alignment between multi-modal content and complex semantics. In summary, BiLGSEH generates discriminative hash codes by effectively aligning semantic distribution and relational structure across modalities. Extensive performance evaluations against 18 competitive methods demonstrate the superiority of our approach. The source code for our method is publicly available at:https://github.com/yileicc/BiLGSEH.
Lei Zhu 0005, Runbing Wu, Chengyuan Zhang 0001, Lin Wu 0001, Shichao Zhang 0001, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 A Deep Semantic Segmentation Network With Semantic and Contextual Refinements
abstract
Semantic segmentation is a fundamental task in multimedia processing, which can be used for analyzing, understanding, editing contents of images and videos, among others. To accelerate the analysis of multimedia data, existing segmentation researches tend to extract semantic information by progressively reducing the spatial resolutions of feature maps. However, this approach introduces a misalignment problem when restoring the resolution of high-level feature maps. In this paper, we design a Semantic Refinement Module (SRM) to address this issue within the segmentation network. Specifically, SRM is designed to learn a transformation offset for each pixel in the upsampled feature maps, guided by high-resolution feature maps and neighboring offsets. By applying these offsets to the upsampled feature maps, SRM enhances the semantic representation of the segmentation network, particularly for pixels around object boundaries. Furthermore, a Contextual Refinement Module (CRM) is presented to capture global context information across both spatial and channel dimensions. To balance dimensions between channel and space, we aggregate the semantic maps from all four stages of the backbone to enrich channel context information. The efficacy of these proposed modules is validated on three widely used datasets—Cityscapes, Bdd100 K, and ADE20K—demonstrating superior performance compared to state-of-the-art methods. Additionally, this paper extends these modules to a lightweight segmentation network, achieving an mIoU of 82.5% on the Cityscapes validation set with only 137.9 GFLOPs.
Deyin Liu, Lin Wu 0001, Song Wang 0008, Xin Guo 0005, Lin Qi 0001
IEEE Trans. Multim.3
2025 Deformation Field Fusion for Medical Image Registration
abstract
Deformable medical image registration is to find a series of non-linear spatial transformations to align a pair of fixed and moving voxel images. Deep learning based registration models are effective in learning differences between such image pair to obtain the deformation field which is specialized in describing non-rigid deformations in the 3D voxel context. However, existing models tend to learn either one single deformation field only or multi-stage (multi-level) deformation fields progressively arriving at a final optimal field. Actually, deformation fields resulting from different architectures or losses are capable of capturing diverse types of deformations, complementing to each other. In this article, we propose a novel framework of fusing different deformation fields to acquire an overall field to describe all-round deformations, in which multiple complementary cues regarding deformable 3D voxels can be strategically leveraged to improve the alignment of the given image pair. The key to the effect of deformation field fusion for registration lies in two aspects: the fusion network architecture and the loss function. Thus, we develop a well-designed fusion block using ingenious operations based on different types of pooling, convolution, and concatenation. Moreover, since calculating the deformation field using a conventional similarity loss cannot describe the contextual variations which are inter-dependent in each pair of fixed and moving images, we propose a novel Contrast-Structural loss to enhance the motion displacement between the image pair by calculating the similarity of pixels in density values, while being ranged in their spatial proximity. Extensive experimental results demonstrate that our proposed method achieves state-of-the-art performance on currently mainstream benchmark datasets.
Haifeng Zhao 0001, Chi Zhang 0082, Deyin Liu, Lin Wu 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 ICAF-4: An Integrated Framework of Category-level Articulated Object Perception and Manipulation for Embodied Intelligence
Li Zhang 0104, Qiankun Li 0004, Qi Wu 0007, Lin Wu 0001, Liu Liu 0012
BMVC5
2024 De novo Protein Design Using Geometric Vector Field Networks
abstract
Advances like protein diffusion have marked revolutionary progress in $\textit{de novo}$ protein design, a central topic in life science. These methods typically depend on protein structure encoders to model residue backbone frames, where atoms do not exist. Most prior encoders rely on atom-wise features, such as angles and distances between atoms, which are not available in this context. Only a few basic encoders, like IPA, have been proposed for this scenario, exposing the frame modeling as a bottleneck. In this work, we introduce the Vector Field Network (VFN), that enables network layers to perform learnable vector computations between coordinates of frame-anchored virtual atoms, thus achieving a higher capability for modeling frames. The vector computation operates in a manner similar to a linear layer, with each input channel receiving 3D virtual atom coordinates instead of scalar values. The multiple feature vectors output by the vector computation are then used to update the residue representations and virtual atom coordinates via attention aggregation. Remarkably, VFN also excels in modeling both frames and atoms, as the real atoms can be treated as the virtual atoms for modeling, positioning VFN as a potential $\textit{universal encoder}$. In protein diffusion (frame modeling), VFN exhibits a impressive performance advantage over IPA, excelling in terms of both designability ($\textbf{67.04}$\% vs. 53.58\%) and diversity ($\textbf{66.54}$\% vs. 51.98\%). In inverse folding(frame and atom modeling), VFN outperforms the previous SoTA model, PiFold ($\textbf{54.7}$\% vs. 51.66\%), on sequence recovery rate; we also propose a method of equipping VFN with the ESM model, which significantly surpasses the previous ESM-based SoTA ($\textbf{62.67}$\% vs. 55.65\%), LM-Design, by a substantial margin. Code is available at https://github.com/aim-uofa/VFN
Weian Mao, Muzhi Zhu, Shuaike Shen, Lin Wu 0001, Hao Chen 0041, Chunhua Shen
ICLR5
2024 EfficientCAPER: An End-to-End Framework for Fast and Robust Category-Level Articulated Object Pose Estimation
abstract
Human life is populated with articulated objects. Pose estimation for category-level articulated objects is a significant challenge due to their inherent complexity and diverse kinematic structures. Current methods for this task usually meet the problems of insufficient consideration of kinematic constraints, self-occlusion, and optimization requirements. In this paper, we propose EfficientCAPER, an end-to-end Category-level Articulated object Pose EstimatoR, eliminating the need for optimization functions as post-processing and utilizing the kinematic structure for joint-centric pose modeling, thus enhancing the efficiency and applicability. Given a partial point cloud as input, the EfficientCAPER firstly estimates the pose for the free part of an articulated object using decoupled rotation representation. Next, we canonicalize the input point cloud to estimate constrained parts' poses by predicting the joint parameters and states as replacements. Evaluations on three diverse datasets, ArtImage, ReArtMix, and RobotArm, show EfficientCAPER's effectiveness and generalization ability to real-world scenarios. The framework exhibits excellent static pose estimation performance for articulated objects, contributing to the advancement of category-level pose estimation. Codes will be made publicly available.
Li Zhang 0104, Lin Wu 0001, Linlin Ou, Liu Liu 0012
NeurIPS4
2024 Generalizing sentence-level lipreading to unseen speakers: a two-stream end-to-end approach
Yu Li 0053, Feng Xue 0002, Lin Wu 0001, Yincen Xie, Shujie Li 0002
Multim. Syst.3
2024 Jacobian norm with Selective Input Gradient Regularization for interpretable adversarial defense
abstract
Deep neural networks (DNNs) can be easily deceived by imperceptible alterations known as adversarial examples. These examples can lead to misclassification , posing a significant threat to the reliability of deep learning systems in real-world applications. Adversarial training (AT) is a popular technique used to enhance robustness by training models on a combination of corrupted and clean data. However, existing AT-based methods often struggle to handle transferred adversarial examples that can fool multiple defense models, thereby falling short of meeting the generalization requirements for real-world scenarios. Furthermore, AT typically fails to provide interpretable predictions, which are crucial for domain experts seeking to understand the behavior of DNNs. To overcome these challenges, we present a novel approach called Jacobian norm and Selective Input Gradient Regularization (J-SIGR). Our method leverages Jacobian normalization to improve robustness and introduces regularization of perturbation-based saliency maps, enabling interpretable predictions. By adopting J-SIGR, we achieve enhanced defense capabilities and promote high interpretability of DNNs. We evaluate the effectiveness of J-SIGR across various architectures by subjecting it to powerful adversarial attacks. Our experimental evaluations provide compelling evidence of the efficacy of J-SIGR against transferred adversarial attacks, while preserving interpretability. The project code can be found at https://github.com/Lywu-github/jJ-SIGR.git .
Deyin Liu, Lin Wu 0001, Bo Li 0090, Farid Boussaïd, Mohammed Bennamoun, Xianghua Xie, Chengwu Liang
Pattern Recognit.2
2023 CTVIS: Consistent Training for Online Video Instance Segmentation
abstract
The discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is directly supervised by the contrastive loss computed upon the contrastive items (CIs), which are sets of anchor/positive/negative embeddings. Recent online VIS methods leverage CIs sourced from one reference frame only, which we argue is insufficient for learning highly discriminative embeddings. Intuitively, a possible strategy to enhance CIs is replicating the inference phase during training. To this end, we propose a simple yet effective training strategy, called Consistent Training for Online VIS (CTVIS), which devotes to aligning the training and inference pipelines in terms of building CIs. Specifically, CTVIS constructs CIs by referring inference the momentum-averaged embedding and the memory bank storage mechanisms, and adding noise to the relevant embeddings. Such an extension allows a reliable comparison between embeddings of current instances and the stable representations of historical instances, thereby conferring an advantage in modeling VIS challenges such as occlusion, re-identification, and deformation. Empirically, CTVIS outstrips the SOTA VIS models by up to +5.0 points on three VIS benchmarks, including YTVIS19 (55.1% AP), YTVIS21 (50.1% AP) and OVIS (35.5% AP). Furthermore, we find that pseudo-videos transformed from images can train robust models surpassing fully-supervised ones.
Kaining Ying, Weian Mao, Zhenhua Wang 0003, Hao Chen 0041, Lin Wu 0001, Yifan Liu 0001, Chengxiang Fan, Yunzhi Zhuge, Chunhua Shen
ICCV6
2023 Image Template Matching via Dense and Consistent Contrastive Learning
abstract
Image template matching refers to localizing a small query image as opposed to a large reference image map. The query image a.k.a template has to be screened across every equal-sized region in the reference map to perform inner-product at pixel-level and the resulting similarity indicates the template location. Due to the domain heterogeneity between template and reference images, the matching performance degrades under dramatic appearance changes. More severely, the asymmetric matching easily leads to over-fitting by suggesting excessively false positive regions. To these ends, we propose an effective template matching method based on contrastive learning to perform a dense and consistent InfoNCEloss during matching. This can increase the matching at finer details, and thus effectively regularizes network training to prevent over-fitting. Extensive experiments on the synthetic aperture radar (SAR) and optical datasets, i.e., SEN1-2 and OS datasets demonstrate that our proposed method outperforms state-of-the-art methods by a large margin.
Bo Li 0090, Lin Wu 0001, Deyin Liu, Hongyang Chen 0001, Yuanxin Ye, Xianghua Xie
ICME2
2023 Contrastive pre-training and linear interaction attention-based transformer for universal medical reports generation
Donghao Zhang 0004, Danli Shi, Renjing Xu, Qingyi Tao, Lin Wu 0001, Mingguang He, ZongYuan Ge
J. Biomed. Informatics6
2023 LipFormer: Learning to Lipread Unseen Speakers Based on Visual-Landmark Transformers
abstract
Lipreading refers to understanding and further translating the speech of a video speaker into textual outputs. State-of-the-art lipreading methods excel in interpreting overlap speakers, i.e., speakers appear in both training and inference. However, generalizing those methods to unseen speakers incurs catastrophic performance degradation due to the limited number of speakers in training bank as well as the dominant visual variations caused by the shape/color of lips presented by different speakers. Therefore, merely depending on the visible changes of lips tends to overfit the model. To improve to generalise, in this paper we propose to use multi-modal features, i.e., visual and landmark, to describe the lip motion while being irrespective to speaker characteristics. The proposed sentence-level framework, dubbed LipFormer, is based on visual-landmark transformer architecture wherein a lip motion stream, a facial landmark stream, and a cross-modal fusion are interconnected. More specifically, the two-stream embeddings produced by self-attention are prompted into a cross-attention module to achieve the alignment across visual and landmark variations. The resulting fused features are decoded into linguistic texts by a cascaded sequence-to-sequence translation. Extensive experiments demonstrate that our method can generalise well to unseen speakers in multiple datasets.
Feng Xue 0002, Yu Li 0053, Deyin Liu, Yincen Xie, Lin Wu 0001, Richang Hong
IEEE Trans. Circuits Syst. Video Technol.5
2023 Semantic-Aware Adversarial Training for Reliable Deep Hashing Retrieval
abstract
Deep hashing has been intensively studied and successfully applied in large-scale image retrieval systems due to its efficiency and effectiveness. Recent studies have recognized that the existence of adversarial examples poses a security threat to deep hashing models, that is, adversarial vulnerability. Notably, it is challenging to efficiently distill reliable semantic representatives for deep hashing to guide adversarial learning, and thereby it hinders the enhancement of adversarial robustness of deep hashing-based retrieval models. Moreover, current researches on adversarial training for deep hashing are hard to be formalized into a unifiedminimaxstructure. In this paper, we explore Semantic-Aware Adversarial Training (SAAT) for improving the adversarial robustness of deep hashing models. Specifically, we conceive a discriminative mainstay features learning (DMFL) scheme to construct semantic representatives for guiding adversarial learning in deep hashing. Particularly, our DMFL with the strict theoretical guarantee is adaptively optimized in a discriminative learning manner, where both discriminative and semantic properties are jointly considered. Moreover, adversarial examples are fabricated by maximizing the Hamming distance between the hash codes of adversarial samples and mainstay features, the efficacy of which is validated in the adversarial attack trials. Further, we,for the first time, formulate the formalized adversarial training of deep hashing into a unified minimax optimization under the guidance of the generated mainstay codes. Extensive experiments on benchmark datasets show superb attack performance against the state-of-the-art algorithms, meanwhile, the proposed adversarial training can effectively eliminate adversarial perturbations for trustworthy deep hashing-based retrieval.
Xu Yuan 0007, Zheng Zhang 0006, Xunguang Wang, Lin Wu 0001
IEEE Trans. Inf. Forensics Secur.4
2023 Learning Resolution-Adaptive Representations for Cross-Resolution Person Re-Identification
abstract
Cross-resolution person re-identification (CRReID) is a challenging and practical problem that involves matching low-resolution (LR) query identity images against high-resolution (HR) gallery images. Query images often suffer from resolution degradation due to the different capturing conditions from real-world cameras. State-of-the-art solutions for CRReID either learn a resolution-invariant representation or adopt a super-resolution (SR) module to recover the missing information from the LR query. In this paper, we propose an alternative SR-free paradigm to directly compare HR and LR images via a dynamic metric that is adaptive to the resolution of a query image. We realize this idea by learning resolution-adaptive representations for cross-resolution comparison. We propose two resolution-adaptive mechanisms to achieve this. The first mechanism encodes the resolution specifics into different subvectors in the penultimate layer of the deep neural network, creating a varying-length representation. To better extract resolution-dependent information, we further propose to learn resolution-adaptive masks for intermediate residual feature blocks. A novel progressive learning strategy is proposed to train those masks properly. These two mechanisms are combined to boost the performance of CRReID. Experimental results show that the proposed method outperforms existing approaches and achieves state-of-the-art performance on multiple CRReID benchmarks.
Lin Wu 0001, Lingqiao Liu, Yang Wang 0023, Zheng Zhang 0006, Farid Boussaïd, Mohammed Bennamoun, Xianghua Xie
IEEE Trans. Image Process.1
2023 Asymmetric Cross-Scale Alignment for Text-Based Person Search
abstract
Text-based person search (TBPS) is of significant importance in intelligent surveillance, which aims to retrieve pedestrian images with high semantic relevance to a given text description. This retrieval task is characterized with both modal heterogeneity and fine-grained matching. To implement this task, one needs to extract multi-scale features from both image and text domains, and then perform the cross-modal alignment. However, most existing approaches only consider the alignment confined at their individual scales, e.g., an image-sentence or a region-phrase scale. Such a strategy adopts the presumable alignment in feature extraction, while overlooking the cross-scale alignment, e.g., image-phrase. In this paper, we present a transformer-based model to extract multi-scale representations, and perform Asymmetric Cross-Scale Alignment (ACSA) to precisely align the two modalities. Specifically, ACSA consists of a global-level alignment module and an asymmetric cross-attention module, where the former aligns an image and texts on a global scale, and the latter applies the cross-attention mechanism to dynamically align the cross-modal entities in region/image-phrase scales. Extensive experiments on two benchmark datasets CUHK-PEDES and RSTPReid demonstrate the effectiveness of our approach.
Zhong Ji, Junhua Hu, Deyin Liu, Lin Wu 0001, Ye Zhao 0001
IEEE Trans. Multim.4
2023 Verbal-Person Nets: Pose-Guided Multi-Granularity Language-to-Person Generation
abstract
Person image generation conditioned on natural language allows us to personalize image editing in a user-friendly manner. This fashion, however, involves different granularities of semantic relevance between texts and visual content. Given a sentence describing an unknown person, we propose a novel pose-guided multi-granularity attention architecture to synthesize the person image in an end-to-end manner. To determine what content to draw at a global outline, the sentence-level description and pose feature maps are incorporated into a U-Net architecture to generate a coarse person image. To further enhance the fine-grained details, we propose to draw the human body parts with highly correlated textual nouns and determine the spatial positions with respect to target pose points. Our model is premised on a conditional generative adversarial network (GAN) that translates language description into a realistic person image. The proposed model is coupled with two-stream discriminators: 1) text-relevant local discriminators to improve the fine-grained appearance by identifying the region-text correspondences at the finer manipulation and 2) a global full-body discriminator to regulate the generation via a pose-weighting feature selection. Extensive experiments conducted on benchmarks validate the superiority of our method for person image generation.
Deyin Liu, Lin Wu 0001, Feng Zheng 0001, Lingqiao Liu, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Generative Metric Learning for Adversarially Robust Open-world Person Re-Identification
abstract
The vulnerability of re-identification (re-ID) models under adversarial attacks is of significant concern as criminals may use adversarial perturbations to evade surveillance systems. Unlike a closed-world re-ID setting (i.e., a fixed number of training categories), a reliable re-ID system in the open world raises the concern of training a robust yet discriminative classifier, which still shows robustness in the context of unknown examples of an identity. In this work, we improve the robustness of open-world re-ID models by proposing a generative metric learning approach to generate adversarial examples that are regularized to produce robust distance metric. The proposed approach leverages the expressive capability of generative adversarial networks to defend the re-ID models against feature disturbance attacks. By generating the target people variants and sampling the triplet units for metric learning, our learned distance metrics are regulated to produce accurate predictions in the feature metric space. Experimental results on the three re-ID datasets, i.e., Market-1501, DukeMTMC-reID, and MSMT17 demonstrate the robustness of our method.
Deyin Liu, Lin Wu 0001, Richang Hong, ZongYuan Ge, Jialie Shen 0001, Farid Boussaïd, Mohammed Bennamoun
ACM Trans. Multim. Comput. Commun. Appl.2
2022 Cross-Modal Retrieval with Heterogeneous Graph Embedding
abstract
Conventional methods address the cross-modal retrieval problem by projecting the multi-modal data into a shared representation space. Such a strategy will inevitably lose the modality-specific information, leading to decreased retrieval accuracy. In this paper, we propose heterogeneous graph embeddings to preserve more abundant cross-modal information. The embedding from one modality will be compensated with the aggregated embeddings from the other modality. In particular, a self-denoising tree search is designed to reduce the "label noise" problem, making the heterogeneous neighborhood more semantically relevant. The dual-path aggregation tackles the "modality imbalance" problem, giving each sample comprehensive dual-modality information. The final heterogeneous graph embedding is obtained by feeding the aggregated dual-modality features to the cross-modal self-attention module. Experiments conducted on cross-modality person re-identification and image-text retrieval task validate the superiority and generality of the proposed method.
Dapeng Chen, Lin Wu 0001, Harry Qin, Wei Peng 0011
ACM Multimedia4
2022 Pseudo-Pair Based Self-Similarity Learning for Unsupervised Person Re-Identification
abstract
Person re-identification (re-ID) is of great importance to video surveillance systems by estimating the similarity between a pair of cross-camera person shorts. Current methods for estimating such similarity require a large number of labeled samples for supervised training. In this paper, we present a pseudo-pair based self-similarity learning approach for unsupervised person re-ID without human annotations. Unlike conventional unsupervised re-ID methods that use pseudo labels based on global clustering, we construct patch surrogate classes as initial supervision, and propose to assign pseudo labels to images through the pairwise gradient-guided similarity separation. This can cluster images in pseudo pairs, and the pseudos can be updated during training. Based on pseudo pairs, we propose to improve the generalization of similarity function via a novel self-similarity learning:it learns local discriminative features from individual images via intra-similarity, and discovers the patch correspondence across images via inter-similarity. The intra-similarity learning is based on channel attention to detect diverse local features from an image. The inter-similarity learning employs a deformable convolution with a non-local block to align patches for cross-image similarity. Experimental results on several re-ID benchmark datasets demonstrate the superiority of the proposed method over the state-of-the-arts.
Lin Wu 0001, Deyin Liu, Dapeng Chen, ZongYuan Ge, Farid Boussaïd, Mohammed Bennamoun, Jialie Shen 0001
IEEE Trans. Image Process.1
2021 Expression of Concern "Crossing generative adversarial networks for cross-view person re-identification" [Neurocomputing 340 (2019) 259-269]
Chengyuan Zhang 0001, Lin Wu 0001, Yang Wang 0023
Neurocomputing2
2021 Deep Coattention-Based Comparator for Relative Representation Learning in Person Re-Identification
abstract
Person re-identification (re-ID) favors discriminative representations over unseen shots to recognize identities in disjoint camera views. Effective methods are developed via pair-wise similarity learning to detect a fixed set of region features, which can be mapped to compute the similarity value. However, relevant parts of each image are detected independently without referring to the correlation on the other image. Also, region-based methods spatially position local features for their aligned similarities. In this article, we introduce the deep coattention-based comparator (DCC) to fuse codependent representations of paired images so as to correlate the best relevant parts and produce their relative representations accordingly. The proposed approach mimics the human foveation to detect the distinct regions concurrently across images and alternatively attends to fuse them into the similarity learning. Our comparator is capable of learning representations relative to a test shot and well-suited to reidentifying pedestrians in surveillance. We perform extensive experiments to provide the insights and demonstrate the state of the arts achieved by our method in benchmark data sets: 1.2 and 2.5 points gain in mean average precision (mAP) on DukeMTMC-reID and Market-1501, respectively.
Lin Wu 0001, Yang Wang 0023, Junbin Gao, Meng Wang 0001, Zhengjun Zha, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2020 Unsupervised Domain Adaptive Object Detection Using Forward-Backward Cyclic Adaptation
Siqi Yang 0001, Lin Wu 0001, Arnold Wiliem, Brian C. Lovell
ACCV (3)2
2020 MDPL-net: Multi-layer Dictionary Learning Network with Added Skip Dense Connections
abstract
Dictionary learning (DL) is powerful for representation learning, while it fails to capture the deep hierarchical information hidden in data. In this paper, we propose a new generalized end-to-end mulita-layer representation learning architecture referred to as Multi-layer Dictionary Pair Learning Network (MDPL-net) for the deep sparse and hierarchical representation of images. To enable MDPL-net to conduct accurate classification, MDPL-net clearly integrates the skip connection end-to-end network and multi-layer deep sparse dictionary learning into a unified architecture. The representation learning module has several hidden DL blocks, where each hidden DL block has a dictionary pair learning (DPL) layer, a batch-norm layer and an activation function layer, and the DL blocks are connected in a feed-forward manner. To further improve the information flow and maintain the privileged features between different DL blocks, a novel skip dense connectivity pattern is deployed between hidden DL blocks, which can obtain more stable and discriminative features. The DPL layer jointly formulates the discriminative synthesis dictionary and analysis dictionary by minimizing reconstruction error within each batch over the feature maps from front layers. Extensive results on benchmark databases demonstrate the effectiveness of MDPL-net for discriminative representation and robust image classification.
Zhao Zhang 0001, Zheng Zhang 0006, Yang Wang 0023, Lin Wu 0001, Meng Wang 0001
ICDM5
2020 Zero-Shot Object Detection via Learning an Embedding from Semantic Space to Visual Space
abstract
Zero-shot object detection (ZSD) has received considerable attention from the community of computer vision in recent years. It aims to simultaneously locate and categorize previously unseen objects during inference. One crucial problem of ZSD is how to accurately predict the label of each object proposal, i.e. categorizing object proposals, when conducting ZSD for unseen categories. Previous ZSD models generally relied on learning an embedding from visual space to semantic space or learning a joint embedding between semantic description and visual representation. As the features in the learned semantic space or the joint projected space tend to suffer from the hubness problem, namely the feature vectors are likely embedded to an area of incorrect labels, and thus it will lead to lower detection precision. In this paper, instead, we propose to learn a deep embedding from the semantic space to the visual space, which enables to well alleviate the hubness problem, because, compared with semantic space or joint embedding space, the distribution in visual space has smaller variance. After learning a deep embedding model, we perform $k$ nearest neighbor search in the visual space of unseen categories to determine the category of each semantic description. Extensive experiments on two public datasets show that our approach significantly outperforms the existing methods.
Xianzhi Wang 0001, Lina Yao 0001, Lin Wu 0001, Feng Zheng 0001
IJCAI4
2020 Mutual kNN based spectral clustering
Malong Tan, Shichao Zhang 0001, Lin Wu 0001
Neural Comput. Appl.3
2020 Medi-Care AI: Predicting medications from billing codes via robust recurrent neural networks
Deyin Liu, Lin Wu 0001, Xue Li 0001, Lin Qi 0001
Neural Networks2
2020 Cross-Entropy Adversarial View Adaptation for Person Re-Identification
abstract
Person re-identification (re-ID) is a task of matching pedestrians under disjoint camera views. To recognize paired snapshots, it has to cope with large cross-view variations caused by the camera view shift. The supervised deep neural networks are effective in producing a set of non-linear projections that can transform cross-view images into a common feature space. However, they typically impose a symmetric architecture, leaving the network ill-conditioned on its optimization. In this paper, we learn view-invariant subspace for person re-ID, and its corresponding similarity metric using an adversarial view adaptation approach. The main contribution is to learn coupled asymmetric mappings regarding view characteristics which are adversarially trained to address the view discrepancy by optimizing the cross-entropy view confusion objective. To determine the similarity value, the network is empowered with a similarity discriminator to promote features that are highly discriminant in distinguishing positive and negative pairs. The other contribution includes an adaptive weighing on the most difficult samples to address the imbalance of within-/between-identity pairs. Our approach achieves notably improved performance in comparison with the state-of-the-arts on benchmark datasets.
Lin Wu 0001, Richang Hong, Yang Wang 0023, Meng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Few-Shot Deep Adversarial Learning for Video-Based Person Re-Identification
abstract
Video-based person re-identification (re-ID) refers to matching people across camera views from arbitrary unaligned video footages. Existing methods rely on supervision signals to optimise a projected space under which the distances between inter/intra-videos are maximised/minimised. However, this demands exhaustively labelling people across camera views, rendering them unable to be scaled in large networked cameras. Also, it is noticed that learning effective video representations with view invariance is not explicitly addressed for which features exhibit different distributions otherwise. Thus, matching videos for person re-ID demands flexible models to capture the dynamics in time-series observations and learn view-invariant representations with access to limited labeled training samples. In this paper, we propose a novel few-shot deep learning approach to videobased person re-ID, to learn comparable representations that are discriminative and view-invariant. The proposed method is developed on the variational recurrent neural networks (VRNNs) and trained adversarially to produce latent variables with temporal dependencies that are highly discriminative yet view-invariant in matching persons. Through extensive experiments conducted on three benchmark datasets, we empirically show the capability of our method in creating view-invariant temporal features and state-of-the-art performance achieved by our method.
Lin Wu 0001, Yang Wang 0023, Hongzhi Yin, Meng Wang 0001, Ling Shao 0001
IEEE Trans. Image Process.1
2019 Deep Instance-Level Hard Negative Mining Model for Histopathology Images
Meng Li 0087, Lin Wu 0001, Arnold Wiliem, Kun Zhao 0001, Teng Zhang 0004, Brian C. Lovell
MICCAI (1)2
2019 CORAL8: Concurrent Object Regression for Area Localization in Medical Image Panels
Sam Maksoud, Arnold Wiliem, Kun Zhao 0001, Teng Zhang 0004, Lin Wu 0001, Brian C. Lovell
MICCAI (1)5
2019 Crossing generative adversarial networks for cross-view person re-identification
Chengyuan Zhang 0001, Lin Wu 0001, Yang Wang 0023
Neurocomputing2
2019 Deep Attention-Based Spatially Recursive Networks for Fine-Grained Visual Recognition
abstract
Fine-grained visual recognition is an important problem in pattern recognition applications. However, it is a challenging task due to the subtle interclass difference and large intraclass variation. Recent visual attention models are able to automatically locate critical object parts and represent them against appearance variations. However, without consideration of spatial dependencies in discriminative feature learning, these methods are underperformed in classifying fine-grained objects. In this paper, we present a deep attention-based spatially recursive model that can learn to attend to critical object parts and encode them into spatially expressive representations. Our network is technically premised on bilinear pooling, enabling local pairwise feature interactions between outputs from two different convolutional neural networks (CNNs) that correspond to distinct region detection and relevant feature extraction. Then, spatial long-short term memory (LSTMs) units are introduced to generate spatially meaningful hidden representations via the long-range dependency on all features in two dimensions. The attention model is leveraged between bilinear outcomes and spatial LSTMs for dynamic selection on varied inputs. Our model, which is composed of two-stream CNN layers, bilinear pooling, and spatial recursive encoding with attention, is end-to-end trainable to serve as the part detector and feature extractor whereby relevant features are localized, extracted, and encoded spatially for recognition purpose. We demonstrate the superiority of our method over two typical fine-grained recognition tasks: fine-grained image classification and person re-identification.
Lin Wu 0001, Yang Wang 0023, Xue Li 0001, Junbin Gao
IEEE Trans. Cybern.1
2019 Cycle-Consistent Deep Generative Hashing for Cross-Modal Retrieval
abstract
In this paper, we propose a novel deep generative approach to cross-modal retrieval to learn hash functions in the absence of paired training samples through the cycle consistency loss. Our proposed approach employs adversarial training scheme to learn a couple of hash functions enabling translation between modalities while assuming the underlying semantic relationship. To induce the hash codes with semantics to the input-output pair, cycle consistency loss is further delved into the adversarial training to strengthen the correlation between the inputs and corresponding outputs. Our approach is generative to learn hash functions, such that the learned hash codes can maximally correlate each input-output correspondence and also regenerate the inputs so as to minimize the information loss. The learning to hash embedding is thus performed to jointly optimize the parameters of the hash functions across modalities as well as the associated generative models. Extensive experiments on a variety of large-scale cross-modal data sets demonstrate that our proposed method outperforms the state of the arts.
Lin Wu 0001, Yang Wang 0023, Ling Shao 0001
IEEE Trans. Image Process.1
2019 Where-and-When to Look: Deep Siamese Attention Networks for Video-Based Person Re-Identification
abstract
Video-based person re-identification (re-id) is a central application in surveillance systems with a significant concern in security. Matching persons across disjoint camera views in their video fragments are inherently challenging due to the large visual variations and uncontrolled frame rates. There are two steps crucial to person re-id, namely, discriminative feature learning and metric learning. However, existing approaches consider the two steps independently, and they do not make full use of the temporal and spatial information in the videos. In this paper, we propose a Siamese attention architecture that jointly learns spatiotemporal video representations and their similarity metrics. The network extracts local convolutional features from regions of each frame and enhances their discriminative capability by focusing on distinct regions when measuring the similarity with another pedestrian video. The attention mechanism is embedded into spatial gated recurrent units to selectively propagate relevant features and memorize their spatial dependencies through the network. The model essentially learns which parts (where) from which frames (when) are relevant and distinctive for matching persons and attaches higher importance therein. The proposed Siamese model is end-to-end trainable to jointly learn comparable hidden representations for paired pedestrian videos and their similarity value. Extensive experiments on three benchmark datasets show the effectiveness of each component of the proposed deep network while outperforming state-of-the-art methods.
Lin Wu 0001, Yang Wang 0023, Junbin Gao, Xue Li 0001
IEEE Trans. Multim.1
2019 3-D PersonVLAD: Learning Deep Global Representations for Video-Based Person Reidentification
abstract
We present the global deep video representation learning to video-based person reidentification (re-ID) that aggregates local 3-D features across the entire video extent. Existing methods typically extract frame-wise deep features from 2-D convolutional networks (ConvNets) which are pooled temporally to produce the video-level representations. However, 2-D ConvNets lose temporal priors immediately after the convolutions, and a separate temporal pooling is limited in capturing human motion in short sequences. In this paper, we present global video representation learning, to be complementary to 3-D ConvNets as a novel layer to capture the appearance and motion dynamics in full-length videos. Nevertheless, encoding each video frame in its entirety and computing aggregate global representations across all frames is tremendously challenging due to the occlusions and misalignments. To resolve this, our proposed network is further augmented with the 3-D part alignment to learn local features through the soft-attention module. These attended features are statistically aggregated to yield identity-discriminative representations. Our global 3-D features are demonstrated to achieve the state-of-the-art results on three benchmark data sets: MARS, Imagery Library for Intelligent Detection Systems-Video Re-identification, and PRID2011.
Lin Wu 0001, Yang Wang 0023, Ling Shao 0001, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 TADA: Trend Alignment with Dual-Attention Multi-task Recurrent Neural Networks for Sales Prediction
abstract
As a common strategy in sales-supply chains, the prediction of sales volume offers precious information for companies to achieve a healthy balance between supply and demand. In practice, the sales prediction task is formulated as a time series prediction problem which aims to predict the future sales volume for different products with the observation of various influential factors (e.g., brand, season, discount, etc.) and corresponding historical sales records. However, with the development of contemporary commercial markets, the dynamic interaction between influential factors with different semantic meanings becomes more subtle, causing challenges in fully capturing dependencies among these variables. Besides, though seeking similar trends from the history benefits the accuracy for the prediction of upcoming sales, existing methods hardly suit sales prediction tasks because the trends in sales time series are more irregular and complex. Hence, we gain insights from the encoder-decoder recurrent neural network (RNN) structure, and propose a novel framework named TADA to carry out trend alignment with dualattention, multi-task RNNs for sales prediction. In TADA, we innovatively divide the influential factors into internal feature and external feature, which are jointly modelled by a multi-task RNN encoder. In the decoding stage, TADA utilizes two attention mechanisms to compensate for the unknown states of influential factors in the future and adaptively align the upcoming trend with relevant historical trends to ensure precise sales prediction. Experimental results on two real-world datasets comprehensively show the superiority of TADA in sales prediction tasks against other state-of-the-art competitors.
Tong Chen 0005, Hongzhi Yin, Hongxu Chen 0002, Lin Wu 0001, Hao Wang 0005, Xiaofang Zhou 0001, Xue Li 0001
ICDM4
2018 Structured deep hashing with convolutional neural networks for fast person re-identification
Lin Wu 0001, Yang Wang 0023, ZongYuan Ge, Qichang Hu, Xue Li 0001
Comput. Vis. Image Underst.1
2018 Beyond Low-Rank Representations: Orthogonal clustering basis reconstruction with optimized graph structure for multi-view spectral clustering
Yang Wang 0023, Lin Wu 0001
Neural Networks2
2018 Deep adaptive feature embedding with local sample distributions for person re-identification
Lin Wu 0001, Yang Wang 0023, Junbin Gao, Xue Li 0001
Pattern Recognit.1
2018 What-and-where to match: Deep spatially multiplicative integration networks for person re-identification
Lin Wu 0001, Yang Wang 0023, Xue Li 0001, Junbin Gao
Pattern Recognit.1
2018 Multiview Spectral Clustering via Structured Low-Rank Matrix Factorization
abstract
Multiview data clustering attracts more attention than their single-view counterparts due to the fact that leveraging multiple independent and complementary information from multiview feature spaces outperforms the single one. Multiview spectral clustering aims at yielding the data partition agreement over their local manifold structures by seeking eigenvalue-eigenvector decompositions. Among all the methods, low-rank representation (LRR) is effective, by exploring the multiview consensus structures beyond the low rankness to boost the clustering performance. However, as we observed, such classical paradigm still suffers from the following stand-out limitations for multiview spectral clustering of overlooking the flexible local manifold structure, caused by aggressively enforcing the low-rank data correlation agreement among all views, and such a strategy, therefore, cannot achieve the satisfied between-views agreement; worse still, LRR is not intuitively flexible to capture the latent data clustering structures. In this paper, first, we present the structured LRR by factorizing into the latent low-dimensional data-cluster representations, which characterize the data clustering structure for each view. Upon such representation, second, the Laplacian regularizer is imposed to be capable of preserving the flexible local manifold structure for each view. Third, we present an iterative multiview agreement strategy by minimizing the divergence objective among all factorized latent data-cluster representations during each iteration of optimization process, where such latent representation from each view serves to regulate those from other views, and such an intuitive process iteratively coordinates all views to be agreeable. Fourth, we remark that such data-cluster representation can flexibly encode the data clustering structure from any view with an adaptive input cluster number. To this end, finally, a novel nonconvex objective function is proposed via the efficient alternating minimization strategy. The complexity analysis is also presented. The extensive experiments conducted against the real-world multiview data sets demonstrate the superiority over the state of the arts.
Yang Wang 0023, Lin Wu 0001, Xuemin Lin 0001, Junbin Gao
IEEE Trans. Neural Networks Learn. Syst.2
2017 Generating Life Course Trajectory Sequences with Recurrent Neural Networks and Application to Early Detection of Social Disadvantage
Lin Wu 0001, Michele Haynes, Tong Chen 0005, Xue Li 0001
ADMA1
2017 Structured learning of metric ensembles with application to person re-identification
Sakrapee Paisitkriangkrai, Lin Wu 0001, Chunhua Shen, Anton van den Hengel
Comput. Vis. Image Underst.2
2017 Robust hashing for multi-view data: Jointly learning low-rank kernelized similarity consensus and hash functions
Lin Wu 0001, Yang Wang 0023
Image Vis. Comput.1
2017 Deep linear discriminant analysis on fisher networks: A hybrid architecture for person re-identification
Lin Wu 0001, Chunhua Shen, Anton van den Hengel
Pattern Recognit.1
2017 Exploiting Attribute Correlations: A Novel Trace Lasso-Based Weakly Supervised Dictionary Learning Method
abstract
It is now well established that sparse representation models are working effectively for many visual recognition tasks, and have pushed forward the success of dictionary learning therein. Recent studies over dictionary learning focus on learning discriminative atoms instead of purely reconstructive ones. However, the existence of intraclass diversities (i.e., data objects within the same category but exhibit large visual dissimilarities), and interclass similarities (i.e., data objects from distinct classes but share much visual similarities), makes it challenging to learn effective recognition models. To this end, a large number of labeled data objects are required to learn models which can effectively characterize these subtle differences. However, labeled data objects are always limited to access, committing it difficult to learn a monolithic dictionary that can be discriminative enough. To address the above limitations, in this paper, we propose a weakly-supervised dictionary learning method to automatically learn a discriminative dictionary by fully exploiting visual attribute correlations rather than label priors. In particular, the intrinsic attribute correlations are deployed as a critical cue to guide the process of object categorization, and then a set of subdictionaries are jointly learned with respect to each category. The resulting dictionary is highly discriminative and leads to intraclass diversity aware sparse representations. Extensive experiments on image classification and object recognition are conducted to show the effectiveness of our approach.
Lin Wu 0001, Yang Wang 0023, Shirui Pan
IEEE Trans. Cybern.1
2017 Effective Multi-Query Expansions: Collaborative Deep Networks for Robust Landmark Retrieval
abstract
Given a query photo issued by a user (q-user), the landmark retrieval is to return a set of photos with their landmarks similar to those of the query, while the existing studies on the landmark retrieval focus on exploiting geometries of landmarks for similarity matches between candidate photos and a query photo. We observe that the same landmarks provided by different users over social media community may convey different geometry information depending on the viewpoints and/or angles, and may, subsequently, yield very different results. In fact, dealing with the landmarks with low quality shapes caused by the photography of q-users is often nontrivial and has seldom been studied. In this paper, we propose a novel framework, namely, multi-query expansions, to retrieve semantically robust landmarks by two steps. First, we identify the top- k photos regarding the latent topics of a query landmark to construct multi-query set so as to remedy its possible low quality shape. For this purpose, we significantly extend the techniques of Latent Dirichlet Allocation. Then, motivated by the typical collaborative filtering methods, we propose to learn a collaborative deep networks-based semantically, nonlinear, and high-level features over the latent factor for landmark photo as the training set, which is formed by matrix factorization over collaborative user-photo matrix regarding the multi-query set. The learned deep network is further applied to generate the features for all the other photos, meanwhile resulting into a compact multi-query set within such space. Then, the final ranking scores are calculated over the high-level feature space between the multi-query set and all other photos, which are ranked to serve as the final ranking list of landmark retrieval. Extensive experiments are conducted on real-world social media data with both landmark photos together with their user information to show the superior performance over the existing methods, especially our recently proposed multi-query based mid-level pattern representation method [1].
Yang Wang 0023, Xuemin Lin 0001, Lin Wu 0001, Wenjie Zhang 0001
IEEE Trans. Image Process.3
2017 Unsupervised Metric Fusion Over Multiview Data by Graph Random Walk-Based Cross-View Diffusion
abstract
Learning an ideal metric is crucial to many tasks in computer vision. Diverse feature representations may combat this problem from different aspects; as visual data objects described by multiple features can be decomposed into multiple views, thus often provide complementary information. In this paper, we propose a cross-view fusion algorithm that leads to a similarity metric for multiview data by systematically fusing multiple similarity measures. Unlike existing paradigms, we focus on learning distance measure by exploiting a graph structure of data samples, where an input similarity matrix can be improved through a propagation of graph random walk. In particular, we construct multiple graphs with each one corresponding to an individual view, and a cross-view fusion approach based on graph random walk is presented to derive an optimal distance measure by fusing multiple metrics. Our method is scalable to a large amount of data by enforcing sparsity through an anchor graph representation. To adaptively control the effects of different views, we dynamically learn view-specific coefficients, which are leveraged into graph random walk to balance multiviews. However, such a strategy may lead to an over-smooth similarity metric where affinities between dissimilar samples may be enlarged by excessively conducting cross-view fusion. Thus, we figure out a heuristic approach to controlling the iteration number in the fusion process in order to avoid over smoothness. Extensive experiments conducted on real-world data sets validate the effectiveness and efficiency of our approach.
Yang Wang 0023, Wenjie Zhang 0001, Lin Wu 0001, Xuemin Lin 0001, Xiang Zhao 0002
IEEE Trans. Neural Networks Learn. Syst.3
2016 Iterative Views Agreement: An Iterative Low-Rank Based Structured Optimization Method to Multi-View Spectral Clustering
Yang Wang 0023, Wenjie Zhang 0001, Lin Wu 0001, Xuemin Lin 0001, Shirui Pan
IJCAI3
2016 Shifting multi-hypergraphs via collaborative probabilistic voting
Yang Wang 0023, Xuemin Lin 0001, Lin Wu 0001, Qing Zhang 0001, Wenjie Zhang 0001
Knowl. Inf. Syst.3
2015 Effective Multi-Query Expansions: Robust Landmark Retrieval
abstract
Given a query photo issued by a user (q-user), the landmark retrieval is to return a set of photos with their landmarks similar to those of the query, while the existing studies on the landmark retrieval focus on exploiting geometries of landmarks for similarity matches between candidate photos and a query photo. We observe that the same landmarks provided by different users may convey different geometry information depending on the viewpoints and/or angles, and may subsequently yield very different results. In fact, dealing with the landmarks with shapes caused by the photography of q-users is often nontrivial and has never been studied.
Yang Wang 0023, Xuemin Lin 0001, Lin Wu 0001, Wenjie Zhang 0001
ACM Multimedia3
2015 Robust User Community-Aware Landmark Photo Retrieval
Lin Wu 0001, John Shepherd 0001, Xiaodi Huang 0001, Chunzhi Hu
MMM (2)1
2015 LBMCH: Learning Bridging Mapping for Cross-modal Hashing
abstract
Hashing has gained considerable attention on large-scale similarity search, due to its enjoyable efficiency and low storage cost. In this paper, we study the problem of learning hash functions in the context of multi-modal data for cross-modal similarity search. Notwithstanding the progress achieved by existing methods, they essentially learn only one common hamming space, where data objects from all modalities are mapped to conduct similarity search. However, such method is unable to well characterize the flexible and discriminative local (neighborhood) structure in all modalities simultaneously, hindering them to achieve better performance. Bearing such stand-out limitation, we propose to learn heterogeneous hamming spaces with each preserving the local structure of data objects from an individual modality. Then, a novel method to learning bridging mapping for cross-modal hashing, named LBMCH, is proposed to characterize the cross-modal semantic correspondence by seamlessly connecting these distinct hamming spaces. Meanwhile, the local structure of each data object in a modality is preserved by constructing an anchor based representation, enabling LBMCH to characterize a linear complexity w.r.t the size of training set. The efficacy of LBMCH is experimentally validated against real-world cross-modal datasets.
Yang Wang 0023, Xuemin Lin 0001, Lin Wu 0001, Wenjie Zhang 0001, Qing Zhang 0001
SIGIR3
2015 Multi-Query Augmentation-Based Web Landmark Photo Retrieval
abstract
Given a query photo characterizing a location-aware landmark shot by a user, landmark retrieval is about returning a set of photos ranked in their similarities to the query. Existing studies on landmark retrieval focus on conducting a matching process between candidate photos and a query photo by exploiting location-aware visual features. Notwithstanding the good results achieved, these approaches are based on an assumption that a landmark of interest is well-captured and distinctive enough to be distinguished from others. In fact, distinctive landmarks may be badly selected, e.g. changes on viewpoints or angles. This will discourage the recognition results if a biased query photo is issued. In this paper, we present a novel technique that exploits user communities in social media networks. Given a biased query photo containing some landmarks taken by a user, we select multiple users to complement this user for retrieval. Multiple photos are then used to enrich the query photo, constituting a more representative yet robust multi-query set. A pattern mining method is developed to obtain a compact feature representation of photos from the multi-query set. Such a representation is utilized to efficiently query the database so as to improve retrieval results. Extensive experiments on real-world datasets demonstrate the effectiveness and efficiency of our approach.
Lin Wu 0001, Xiaodi Huang 0001, John Shepherd 0001, Yang Wang 0023
Comput. J.1
2015 An efficient framework of Bregman divergence optimization for co-ranking images and tags in a heterogeneous network
Lin Wu 0001, Xiaodi Huang 0001, Chengyuan Zhang 0001, John Shepherd 0001, Yang Wang 0023
Multim. Tools Appl.1
2015 Robust Subspace Clustering for Multi-View Data by Exploiting Correlation Consensus
abstract
More often than not, a multimedia data described by multiple features, such as color and shape features, can be naturally decomposed of multi-views. Since multi-views provide complementary information to each other, great endeavors have been dedicated by leveraging multiple views instead of a single view to achieve the better clustering performance. To effectively exploit data correlation consensus among multi-views, in this paper, we study subspace clustering for multi-view data while keeping individual views well encapsulated. For characterizing data correlations, we generate a similarity matrix in a way that high affinity values are assigned to data objects within the same subspace across views, while the correlations among data objects from distinct subspaces are minimized. Before generating this matrix, however, we should consider that multi-view data in practice might be corrupted by noise. The corrupted data will significantly downgrade clustering results. We first present a novel objective function coupled with an angular based regularizer. By minimizing this function, multiple sparse vectors are obtained for each data object as its multiple representations. In fact, these sparse vectors result from reaching data correlation consensus on all views. For tackling noise corruption, we present a sparsity-based approach that refines the angular-based data correlation. Using this approach, a more ideal data similarity matrix is generated for multi-view data. Spectral clustering is then applied to the similarity matrix to obtain the final subspace clustering. Extensive experiments have been conducted to validate the effectiveness of our proposed approach.
Yang Wang 0023, Xuemin Lin 0001, Lin Wu 0001, Wenjie Zhang 0001, Qing Zhang 0001, Xiaodi Huang 0001
IEEE Trans. Image Process.3
2014 Exploiting Correlation Consensus: Towards Subspace Clustering for Multi-modal Data
abstract
Often, a data object described by many features can be decomposed as multi-modalities, which always provide complementary information to each other. In this paper, we study subspace clustering for multi-modal data by effectively exploiting data correlation consensus across modalities, while keeping individual modalities well encapsulated. Our technique can yield a more ideal data similarity matrix, which encodes strong data correlations for the cross-modal data objects in the same subspace.
Yang Wang 0023, Xuemin Lin 0001, Lin Wu 0001, Wenjie Zhang 0001, Qing Zhang 0001
ACM Multimedia3
2014 Shifting Hypergraphs by Probabilistic Voting
Yang Wang 0002, Xuemin Lin 0001, Qing Zhang 0001, Lin Wu 0001
PAKDD (2)4
2013 Efficient image and tag co-ranking: a bregman divergence optimization method
abstract
Ranking on image search has attracted considerable attentions. Many graph-based algorithms have been proposed to solve this problem. Despite their remarkable success, these approaches are restricted to their separated image networks. To improve the ranking performance, one effective strategy is to work beyond the separated image graph by leveraging fruitful information from manual semantic labeling (i.e., tags) associated with images, which leads to the technique of co-ranking images and tags, a representative method that aims to explore the reinforcing relationship between image and tag graphs. The idea of co-ranking is implemented by adopting the paradigm of random walks. However, there are two problems hidden in co-ranking remained to be open: the high computational complexity and the problem of out-of-sample. To address the challenges above, in this paper, we cast the co-ranking process into a Bregman divergence optimization framework under which we transform the original random walk into an equivalent optimal kernel matrix learning problem. Enhanced by this new formulation, we derive a novel extension to achieve a better performance for both in-sample and out-of-sample cases. Extensive experiments are conducted to demonstrate the effectiveness and efficiency of our approach.
Lin Wu 0001, Yang Wang 0023, John Shepherd 0001
ACM Multimedia1
2013 Co-ranking Images and Tags via Random Walks on a Heterogeneous Graph
Lin Wu 0001, Yang Wang 0023, John Shepherd 0001
MMM (1)1
2013 An Optimization Method for Proportionally Diversifying Search Results
Lin Wu 0001, Yang Wang 0023, John Shepherd 0001, Xiang Zhao 0002
PAKDD (1)1
2013 Max-sum diversification on image ranking with non-uniform matroid constraints
Lin Wu 0001, Yang Wang 0023, John Shepherd 0001, Xiang Zhao 0002
Neurocomputing1
2013 Clustering via geometric median shift over Riemannian manifolds
Yang Wang 0023, Xiaodi Huang 0001, Lin Wu 0001
Inf. Sci.3
2012 Human Action Recognition from Video Sequences by Enforcing Tri-view Constraints
abstract
Two-view methods have been well developed to identify human actions. However, in a case where the corresponding imaged points cannot induce distinguished measures, the performance of the methods deteriorates. For this reason, we propose a new view-invariant measure for human action recognition by enforcing tri-view constraints in this paper. This new measurement method can be tolerant to different rates of human actions and the anthropometric proportions. We apply our approach to video synchronization by imposing both the similarity ratio and the consistency in the trifocal tensor over entire video sequences. By testing on both synthetic and real data, our method has achieved higher tolerance to noise levels, as well as higher identification accuracy than the traditional two-view method. Experimental results demonstrate that our approach can identify human pose transitions, in spite of dynamic time-lines, different viewpoints and unknown camera parameters.
Yang Wang 0023, Lin Wu 0001, Xiaodi Huang 0001, Xuemin Lin 0001
Comput. J.2
2012 Detecting image forgeries using metrology
Lin Wu 0001, Xiaochun Cao, Wei Zhang 0031, Yang Wang 0023
Mach. Vis. Appl.1
2011 Action recognition using tri-view constraints
abstract
Two-view methods have been well developed to identify human actions. However, in a case where the corresponding imaged points cannot induce distinguished measures, the performance of the methods deteriorates. For this reason, we propose a new view-invariant measure for human action recognition by enforcing tri-view constraints in this paper. We apply our approach to video synchronization by imposing both the similarity ratio and the consistency in the trifocal tensor over entire video sequences. By testing on both synthetic and real data, our method has achieved higher tolerance to noise levels, as well as higher identification accuracy than the traditional two-view method. Experimental results demonstrate that our approach can identify human pose transitions, despite of dynamic time-lines, different viewpoints, and unknown camera parameters.
Yang Wang 0023, Lin Wu 0001, Xiaodi Huang 0001
AVSS2
2010 Geo-location estimation from two shadow trajectories
abstract
The position of a world point's solar shadow depends on its geographical location, the geometrical relationship between the orientation of the sunshine and the ground plane where the shadow casts. This paper investigates the property of solar shadow trajectories on a planar surface and shows that camera parameters, latitude, longitude can be estimated from two observed shadow trajectories. Our contribution is that we use the design of the analemmatic sundial to get the shadow conic and furthermore recover the camera's geographical location. The proposed method does not require the shadow casting objects or a vertical object to be visible in the recovery of camera calibration. This approach is thoroughly validated on both synthetic and real data, and tested against various sources of errors including noise and number of observations.
Lin Wu 0001, Xiaochun Cao
CVPR1
2010 Camera calibration and geo-location estimation from two shadow trajectories
Lin Wu 0001, Xiaochun Cao, Hassan Foroosh
Comput. Vis. Image Underst.1
2010 Automatic Geo-Registration for Port Surveillance
abstract
This paper proposes a new solution to geo-register the nearly feature-less maritime video feeds. We detect the horizon using sizable or uniformly moving vessels, and estimate the vertical apex using water reflections of the street lamps. The computed horizon and apex provide a metric rectification that removes the affine distortions and reduces the searching space for geo-registration. Geo-registration is obtained by searching the best orientation where the estimated water masks on satellite images and camera views are matched. The proposed solution has the following contributions: first, water and coastlines are used as features for registration between horizontally looking maritime views and satellite images. Second, water reflections are proposed to estimate the vertical vanishing point. Third, we give algorithms for the detection of water areas in both satellite images and camera views. Experimental results and applications on cross camera tracking are demonstrated. We also discuss several observations, as well as limitations of the proposed approach.
Xiaochun Cao, Lin Wu 0001, Zeeshan Rasheed 0002, Tae Eun Choe, Feng Guo 0006, Niels Haering
Int. J. Pattern Recognit. Artif. Intell.2
2010 Video synchronization and its application to object transfer
Xiaochun Cao, Lin Wu 0001, Jiangjian Xiao, Hassan Foroosh, Jigui Zhu, Xiaohong Li 0001
Image Vis. Comput.2