VLDB 2026 Research / reviewers in the wild / expert
Zechao Li
dblp:51/8693
· DBLP profile ↗
189ranked-venue papers
25as first author
93since 2021 · last 2026
0000-0002-5341-5985ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 113 · 11 first-author · 50 since 2021Artificial intelligence and machine learning · 91 · 13 first-author · 44 since 2021Databases, data management, data science and information retrieval · 13 · 3 first-author · 5 since 2021Computer networks · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-Grained Image Retrieval via Dual-Vision AdaptationabstractFine-Grained Image Retrieval~(FGIR) faces challenges in learning discriminative visual representations to retrieve images with similar fine-grained features. Current leading FGIR solutions typically follow two regimes: enforce pairwise similarity constraints in the semantic embedding space, or incorporate a localization sub-network to fine-tune the entire model. However, such two regimes tend to overfit the training data while forgetting the knowledge gained from large-scale pre-training, thus reducing their generalization ability. In this paper, we propose a Dual-Vision Adaptation (DVA) approach for FGIR, which guides the frozen pre-trained model to perform FGIR through collaborative sample and feature adaptation. Specifically, we design Object-Perceptual Adaptation, which modifies input samples to help the pre-trained model perceive critical objects and elements within objects that are helpful for category prediction. Meanwhile, we propose In-Context Adaptation, which introduces a small set of parameters for feature adaptation without modifying the pre-trained parameters. This makes the FGIR task using these adapted features closer to the task solved during the pre-training. Additionally, to balance retrieval efficiency and performance, we propose Discrimination Perception Transfer to transfer the discriminative knowledge in the object-perceptual adaptation to the image encoder using the knowledge distillation mechanism. Extensive experiments show that DVA performs well on three fine-grained datasets. Xin Jiang 0010, Meiqi Cao, Hao Tang 0007, Fei Shen 0004, Zechao Li |
AAAI | 5 |
| 2026 | The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool HallucinationabstractEnhancing the reasoning capabilities of Large Language Models (LLMs) is a key strategy for building Agents that "think then act."However, recent observations, like OpenAI's o3, suggest a paradox: stronger reasoning often coincides with increased hallucination, yet no prior work has systematically examined whether reasoning enhancement itself causes tool hallucination.To address this gap, we pose the central question: Does strengthening reasoning increase tool hallucination of LLM Agents?We address this gap by introducing SIMPLETOOL-HALLUBENCH, a diagnostic benchmark measuring tool hallucination.Through controlled experiments, we show that across RL, distillation, and toggleable reasoning modes, gains in task performance are consistently accompanied by higher tool hallucination rates.This effect is training method-agnostic and transcends simple overfitting, as training even on non-tool-related tasks (e.g., mathematics) still amplifies tool hallucination.Controlled ablations further reveal that the reasoning itself, rather than RL training in general, is most closely associated with the hallucination increase.We evaluate mitigation strategies including Prompt Engineering and Direct Preference Optimization (DPO), revealing a fundamental reliability-capability trade-off: reducing hallucination unavoidably degrades utility.Our findings demonstrate that under current reasoning enhancement methods, improved reasoning is systematically associated with increased tool hallucination, highlighting the need for training objectives that jointly optimize capability and reliability.Our Chenlong Yin, Zeyang Sha, Shiwen Cui, Changhua Meng, Zechao Li |
ACL (1) | 5 |
| 2026 | GradAlign: Detecting Out-of-Distribution Samples via Gradient Concentration
Jiawei Gu, Yanpeng Sun, Hao Tang 0007, Zechao Li |
Int. J. Comput. Vis. | 4 |
| 2026 | CylindFormer: Image-to-Point Cloud Registration with Cylindrical Transformer
Hao Tang 0007, Yanpeng Sun, Shengfeng He, Zechao Li |
Int. J. Comput. Vis. | 5 |
| 2026 | Jo-SNC: Combating Noisy Labels Through Fostering Self- and Neighbor-ConsistencyabstractLabel noise is pervasive in various real-world scenarios, posing challenges in supervised deep learning. Deep networks are vulnerable to such label-corrupted samples due to the memorization effect. One major stream of previous methods concentrates on identifying clean data for training. However, these methods often neglect imbalances in label noise across different mini-batches and devote insufficient attention to out-of-distribution noisy data. To this end, we propose a noise-robust method named Jo-SNC (Joint sample selection and model regularization based on Self- and Neighbor-Consistency). Specifically, we propose to employ the Jensen-Shannon divergence to measure the "likelihood" of a sample being clean or out-of-distribution. This process factors in the nearest neighbors of each sample to reinforce the reliability of clean sample identification. We design a self-adaptive, data-driven thresholding scheme to adjust per-class selection thresholds. While clean samples undergo conventional training, detected in-distribution and out-of-distribution noisy samples are trained following partial label learning and negative learning, respectively. Finally, we advance the model performance further by proposing a triplet consistency regularization that promotes self-prediction consistency, neighbor-prediction consistency, and feature consistency. Extensive experiments on various benchmark datasets and comprehensive ablation studies demonstrate the effectiveness and superiority of our approach over existing state-of-the-art methods. Zeren Sun, Yazhou Yao, Tongliang Liu, Zechao Li, Fumin Shen, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Refine, Control and Distill: A Text-to-Image Framework for Faithful Image GenerationabstractWhile text-to-image diffusion models exhibit outstanding results, they struggle to faithfully generate key subjects with corresponding attributes in prompts, challenges known as catastrophic neglect and attribute binding. Previous works typically utilize attention adjustments to solve the above problems, whereas we observe that they may still generate unfaithful images. In this paper, we carefully analyze the text-to-image process and pinpoint three pivotal bottlenecks that hinder image faithful generation: (1) unequal responses of neglected subjects in text embedding, (2) competition and entanglement between subjects' attention, and (3) suboptimal quality of intermediate features from U-Net. Based on the aforementioned observations, we propose a Refine, Control, and Distill (RCD) framework built upon the stable diffusion model to alleviate the negative effects raised by the bottlenecks mentioned above, respectively. Specifically, we achieve the above goals through a text embedding refinement module, three region-level attention control losses, and self-distillation of intermediate semantic features in the denoising process. Our approach exhibits promising capability in generating faithful and high-quality images and outperforms state-of-the-art methods through extensive quantitative and qualitative evaluations on recent advanced base diffusion models. Peng Xing, Ning Wang 0020, Yanpeng Sun, Jinhui Tang 0001, Zechao Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Trajectory-enhanced transferable attacks for vision-language pre-trained models
Haiqi Zhang 0001, Ziqiang Li 0001, Hao Tang 0007, Zechao Li |
Pattern Recognit. | 4 |
| 2026 | Robust object detection in adverse weather with feature decorrelation via independence learning
Shiyu Xuan, Zechao Li |
Pattern Recognit. | 3 |
| 2026 | Beyond aggregate metrics: Scale-balanced temporal grounding via semantic diffusion encoding and structure-aware decoding
Henghao Zhao, Zechao Li, Xuezhi He |
Pattern Recognit. | 3 |
| 2026 | FineG-RAG: Fine-Grained Retrieval-Augmented Generation for Multimodal Large Language ModelsabstractFine-grained visual recognition refers to the ability to distinguish subtle differences between visually similar objects— a fundamental yet challenging capability for Multimodal Large Language Models (MLLMs). In this paper, we observe that even strong open-source MLLMs, such as Qwen2-VL and InternVL2, still struggle with accurately identifying fine-grained categories. These models often fail to attend to subtle but critical details for precise discrimination. To unlock this potential, we propose FineG-RAG, a retrieval-augmented generation pipeline designed to enhance the fine-grained recognition capabilities of MLLMs. FineG-RAG integrates external fine-grained knowledge into the recognition process via a generalized retriever. To support this, we construct fine-grained visual-language knowledge database containing representative images with wide visual diversity and expert-crafted attribute descriptions from multiple perspectives. Relevant fine-grained knowledge is retrieved from this database and fed into a visual-language augmented prompt, which provides rich multimodal context to guide MLLMs in generating accurate labels. To better evaluate the fine-grained recognition capabilities of MLLMs, we design a multiple-choice evaluation strategy based on publicly four fine-grained datasets. Extensive experiments demonstrate that FineG-RAG consistently outperforms baseline methods, achieving superior recognition accuracy across a range of off-the-shelf, open-source MLLMs. Lu Jin 0001, Xinguang Xiang, Yanpeng Sun, Zechao Li, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | URA-Net: Uncertainty-Integrated Anomaly Perception and Restoration Attention Network for Unsupervised Anomaly DetectionabstractUnsupervised anomaly detection plays a pivotal role in industrial defect inspection and medical image analysis, with most methods relying on the reconstruction framework. However, these methods may suffer from over-generalization, enabling them to reconstruct anomalies well, which leads to poor detection performance. To address this issue, instead of focusing solely on normality reconstruction, we propose an innovative Uncertainty-Integrated Anomaly Perception and Restoration Attention Network (URA-Net), which explicitly restores abnormal patterns to their corresponding normality. First, unlike traditional image reconstruction methods, we utilize a pre-trained convolutional neural network to extract multi-level semantic features as the reconstruction target. To assist the URA-Net learning to restore anomalies, we introduce a novel feature-level artificial anomaly synthesis module to generate anomalous samples for training. Subsequently, a novel uncertainty-integrated anomaly perception module based on Bayesian neural networks is introduced to learn the distributions of anomalous and normal features. This facilitates the estimation of anomalous regions and ambiguous boundaries, laying the foundation for subsequent anomaly restoration. Then, we propose a novel restoration attention mechanism that leverages global normal semantic information to restore detected anomalous regions, thereby obtaining defect-free restored features. Finally, we employ residual maps between input features and restored features for anomaly detection and localization. The comprehensive experimental results on two industrial datasets, MVTec AD and BTAD, along with a medical image dataset, OCT-2017, unequivocally demonstrate the effectiveness and superiority of the proposed method. Peng Xing, Yunkang Cao, Haiming Yao, Weiming Shen 0001, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | SSP-SAM: SAM With Semantic-Spatial Prompt for Referring Expression SegmentationabstractThe Segment Anything Model (SAM) excels at general image segmentation but has limited ability to understand natural language, which restricts its direct application in Referring Expression Segmentation (RES). Toward this end, we propose SSP-SAM, a framework that fully utilizes SAM’s segmentation capabilities by integrating a Semantic-Spatial Prompt (SSP) encoder. Specifically, we incorporate both visual and linguistic attention adapters into the SSP encoder, which highlight salient objects within the visual features and discriminative phrases within the linguistic features. This design enhances the referent representation for the prompt generator, resulting in high-quality SSPs that enable SAM to generate precise masks guided by language. Although not specifically designed for Generalized RES (GRES), where the referent may correspond to zero, one, or multiple objects, SSP-SAM naturally supports this more flexible setting without additional modifications. Extensive experiments on widely used RES and GRES benchmarks confirm the superiority of our method. Notably, our approach generates segmentation masks of high quality, achieving strong precision even at strict thresholds such as [email protected]. Further evaluation on the PhraseCut dataset demonstrates improved performance in open-vocabulary scenarios compared to existing state-of-the-art RES methods. The code and checkpoints are available at: https://github.com/WayneTomas/SSP-SAM. Wei Tang 0011, Xuejing Liu, Yanpeng Sun, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | VideoExpert: Augmented LLM for Temporal-Sensitive Video UnderstandingabstractThe core challenge in video understanding lies in perceiving dynamic content changes over time. However, multi-modal large language models (MLLMs) struggle with temporal-sensitive video tasks, such as video temporal grounding, which requires generating timestamps to mark the occurrence of specific events. Existing strategies require MLLMs to generate absolute or relative timestamps directly. We have observed that those MLLMs tend to rely more on language patterns than visual cues when generating timestamps, affecting their performance. To address this problem, we propose VideoExpert, a general-purpose MLLM suitable for several temporal-sensitive video tasks. Inspired by the expert concept, VideoExpert integrates two parallel modules: the Temporal Expert and the Spatial Expert. The Temporal Expert is responsible for modeling time sequences and performing temporal grounding. It processes high-frame-rate yet compressed tokens to capture dynamic variations in videos and includes a lightweight prediction head for precise event localization. The Spatial Expert focuses on content detail analysis and instruction following. It handles specially designed spatial tokens and language input, aiming to generate content-related responses. These two experts collaborate seamlessly via a special token, ensuring coordinated temporal grounding and content generation. Notably, the Temporal and Spatial Experts maintain independent parameter sets. This parameter decoupling design enables specialized learning within each part without mutual interference. By offloading temporal grounding from content generation, VideoExpert prevents text pattern biases in timestamp predictions. Moreover, we introduce a Spatial Compress module to obtain spatial tokens. This module filters and compresses patch tokens while preserving key information, delivering compact yet detail-rich input for the Spatial Expert. Extensive experiments conducted on four widely-used benchmarks (i.e. Charades-STA, QVHighlight, YouCookII and NextGQA) across four tasks (temporal grounding, highlight detection, dense video captioning and grounding question answering) demonstrate the effectiveness and versatility of the VideoExpert. Henghao Zhao, Ge-Peng Ji, Rui Yan 0010, Huan Xiong, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Gradient Pruning Interactive Attack for Vision-Language Pre-Training ModelsabstractVision-Language Pre-training (VLP) models exhibit pronounced vulnerability to multimodal adversarial examples, necessitating rigorous robustness research, particularly for transferable attacks in black-box scenarios. Current research predominantly enhances attack transferability across VLP models by diversifying image and text inputs. However, during adversarial example generation, these methods often prioritize amplifying inter-modal semantic discrepancies (i.e., modality-discrepancy features) while overlooking model-specific semantic features critical to transferable attacks. To address this limitation, we pro pose a transferable Gradient Pruning Interactive Attack (GPI Attack), which integrates gradient-pruned image perturbations with semantic-oriented text perturbations through modality interaction. For image attacks, extreme backpropagated gradients may cause adversarial examples to highlight certain model specific features, leading to poor transferability. To suppress this feature, the textual modality guides the pruning of extreme gradients within intermediate VLP blocks, and these pruned gradients are subsequently employed to direct the generation of adversarial images. For text attacks, we consolidate the perturbation process solely at the embedding level, which reduces semantic discrepancies across hierarchical structures and significantly enhances the generalizability of adversarial texts. Experimental results demonstrate the effectiveness of GPI-Attack in image text retrieval tasks on multimodal datasets such as Flickr30K and MSCOCO. Additionally, the proposed gradient pruning technique is plug-and-play, showing performance improvements even when applied to baseline methods, indicating its potential as a valuable enhancement for attack performance. Haiqi Zhang 0001, Hao Tang 0007, Yanpeng Sun, Zechao Li |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | Progressive Feature Encoding With Background Perturbation Learning for Ultra-Fine-Grained Visual CategorizationabstractUltra-Fine-Grained Visual Categorization (Ultra-FGVC) aims to classify objects into sub-granular categories, presenting the challenge of distinguishing visually similar objects with limited data. Existing methods primarily address sample scarcity but often overlook the importance of leveraging intrinsic object features to construct highly discriminative representations. This limitation significantly constrains their effectiveness in Ultra-FGVC tasks. To address these challenges, we propose SV-Transformer that progressively encodes object features while incorporating background perturbation modeling to generate robust and discriminative representations. At the core of our approach is a progressive feature encoder, which hierarchically extracts global semantic structures and local discriminative details from backbone-generated representations. This design enhances inter-class separability while ensuring resilience to intra-class variations. Furthermore, our background perturbation learning mechanism introduces controlled variations in the feature space, effectively mitigating the impact of sample limitations and improving the model's capacity to capture fine-grained distinctions. Comprehensive experiments demonstrate that SV-Transformer achieves state-of-the-art performance on benchmark Ultra-FGVC datasets, showcasing its efficacy in addressing the challenges of Ultra-FGVC task. Xin Jiang 0010, Ziye Fang, Fei Shen 0004, Junyao Gao 0002, Zechao Li |
IEEE Trans. Image Process. | 5 |
| 2026 | Through the Looking Glass: A Dual Perspective on Weakly Supervised Few-Shot SegmentationabstractMeta-learning aims to uniformly sample homologous support-query pairs, characterized by the same categories and similar attributes, and extract useful inductive biases through identical network architectures. However, this identical network design results in over-semantic homogenization. To address this, we propose a novel homologous but heterogeneous network. By treating support-query pairs as dual perspectives, we introduce heterogeneous visual aggregation (HA) modules to enhance complementarity while preserving semantic commonality. To further reduce semantic noise and amplify the uniqueness of heterogeneous semantics, we design a heterogeneous transport (HT) module. Finally, we propose heterogeneous CLIP (HC) textual information to enhance the generalization capability of multimodal models. In the weakly-supervised few-shot semantic segmentation (WFSS) task, with only 1/24 of the parameters of existing state-of-the-art models, TLG achieves a 13.2% improvement on Pascal- $5{^{\text {i}}}$ and a 7.9% improvement on COCO- $20{^{\text {i}}}$ . To the best of our knowledge, TLG is also the first weakly-supervised (image-level) model that outperforms fully supervised (pixel-level) models under the same backbone architectures. The code is available at https://github.com/jarch-ma/TLG. Jiaqi Ma 0006, Guosen Xie, Fang Zhao 0006, Zechao Li |
IEEE Trans. Image Process. | 4 |
| 2026 | Pseudo-Text Guided Robust Learning for Noisy Correspondence in Cross-Modal RetrievalabstractNoisy Correspondence (NC), caused by mismatched pairs in multimedia datasets, poses major challenges for cross-modal retrieval, especially under high noise levels. Existing solutions often suffer from substantial performance degradation as noise levels increase. To address this issue, we propose Pseudo-Text guided Robust Learning (PTRL), a novel framework designed to identify noisy pairs and enhance model robustness. Specifically, PTRL leverages pseudo-text as explicit supervision signals and introduces a new data division criterion to accurately distinguish between clean and noisy pairs. Instead of discarding or directly using noisy data, PTRL proposes a pseudo-text replacement strategy to maintain semantic consistency of the training set, thereby facilitating more reliable learning. In addition, pseudo-text-image pairs serve as a form of data augmentation, enriching data diversity and improving model generalization. To further stabilize training and mitigate overfitting, PTRL incorporates a robust InfoNCE loss that is particularly effective in the presence of noise. Extensive experiments demonstrate that PTRL achieves state-of-the-art performance and robustness, with an RSum improvements of +60.1% on Flickr30K and +22.6% on MS-COCO at an 80% noise level, significantly outperforming existing methods. The datasets and source code are available at https://github.com/shidan0122/PTRL.git. Dan Shi 0003, Zechao Li, Lei Zhu 0002, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | Contrastive Graph Modeling for Cross-Domain Few-Shot Medical Image SegmentationabstractCross-domain few-shot medical image segmentation (CD-FSMIS) offers a promising and data-efficient solution for medical applications where annotations are severely scarce and multimodal analysis is required. However, existing methods typically filter out domain-specific information to improve generalization, which inadvertently limits cross-domain performance and degrades source-domain accuracy. To address this, we present Contrastive Graph Modeling (C-Graph), a framework that leverages the structural consistency of medical images as a reliable domain-transferable prior. We represent image features as graphs, with pixels as nodes and semantic affinities as edges. A Structural Prior Graph (SPG) layer is proposed to capture and transfer target-category node dependencies and enable global structure modeling through explicit node interactions. Building upon SPG layers, we introduce a Subgraph Matching Decoding (SMD) mechanism that exploits semantic relations among nodes to guide prediction. Furthermore, we design a Confusion-minimizing Node Contrast (CNC) loss to mitigate node ambiguity and subgraph heterogeneity by contrastively enhancing node discriminability in the graph space. Our method significantly outperforms prior CD-FSMIS approaches across multiple cross-domain benchmarks, achieving state-of-the-art performance while simultaneously preserving strong segmentation accuracy on the source domain. Our code is available at https://github.com/primebo1/C-Graph. Yuntian Bo, Tao Zhou 0002, Zechao Li, Haofeng Zhang 0001, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2026 | Rethinking Vision Transformer for Large-Scale Fine-Grained Image RetrievalabstractLarge-scale fine-grained image retrieval (FGIR) aims to retrieve images belonging to the same subcategory as a given query by capturing subtle differences in a large-scale setting. Recently, Vision Transformers (ViT) have been employed in FGIR due to their powerful self-attention mechanism for modeling long-range dependencies. However, most Transformer-based methods focus primarily on leveraging self-attention to distinguish fine-grained details, while overlooking the high computational complexity and redundant dependencies inherent to these models, limiting their scalability and effectiveness in large-scale FGIR. In this paper, we propose an Efficient and Effective ViT-based framework, termedEET, which integrates token pruning module with a discriminative transfer strategy to address these limitations. Specifically, we introduce a content-based token pruning scheme to enhance the efficiency of the vanilla ViT, progressively removing background or low-discriminative tokens at different stages by exploiting feature responses and self-attention mechanism. To ensure the resulting efficient ViT retains strong discriminative power, we further present a discriminative transfer strategy comprising bothdiscriminative knowledge transferanddiscriminative region guidance. Using a distillation paradigm, these components transfer knowledge from a larger “teacher” ViT to a more efficient “student” model, guiding the latter to focus on subtle yet crucial regions in a cost-free manner. Extensive experiments on two widely-used fine-grained datasets and four large-scale fine-grained datasets demonstrate the effectiveness of our method. Specifically, EET reduces the inference latency of ViT-Small by 42.7% and boosts the retrieval performance of 16-bit hash codes by 5.15% on the challenging NABirds dataset. The code is publicly available at:https://github.com/WhiteJiang/EET. Xin Jiang 0010, Hao Tang 0007, Yonghua Pan, Zechao Li |
IEEE Trans. Multim. | 4 |
| 2026 | Multi-Modal Knowledge Distillation Hashing Based on CLIP for Weakly Supervised Image RetrievalabstractExisting weakly supervised hashing often suffers from the imprecision of user-provided tags and over-reliance on textual knowledge from pre-trained word embeddings, neglecting crucial visual knowledge associated with image labels. As a result, this leads to unsatisfactory performance in closed-vocabulary tasks and limited generalization in open-vocabulary scenarios. To address this issue, we propose Multi-modal Knowledge Distillation Hashing (MKDH), a novel method leveraging visual and language pre-training (VLP) model such as CLIP to learn robust hash codes. Our method designs a dual-layer attention adapter to generate joint representations by capturing fine-grained visual and textual knowledge from the CLIP teacher network. Additionally, we introduce a knowledge extraction contrastive loss to enhance the robustness of joint representations and a knowledge distillation contrastive loss to transfer the extracted multi-modal knowledge to the hash codes. To further mitigate the negative impact of false negative pairs in these contrastive losses, we introduce false negative weighting strategy that reduces the weights assigned to such pairs. Extensive experiments on three widely used datasets demonstrate that our method achieves robust retrieval performance with significant improvements in both closed- and open-vocabulary settings. The source code is available athttps://github.com/IMAG-LZY/MKDH. Zhengyun Lu, Lu Jin 0001, Zechao Li, Jinhui Tang 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | Visual Position Prompt for MLLM Based Visual GroundingabstractAlthough Multimodal Large Language Models (MLLMs) excel at various image-related tasks, they encounter challenges in precisely aligning coordinates with spatial information within images, particularly in position-aware tasks such as visual grounding. This limitation arises from two key factors. First, MLLMs lack explicit spatial references, making it difficult to associate textual descriptions with precise image locations. Second, their feature extraction processes prioritize global context over fine-grained spatial details, leading to weak localization capability. To address these issues, we introduce VPP-LLaVA, an MLLM enhanced with Visual Position Prompt (VPP) to improve its grounding capability. VPP-LLaVA integrates two complementary mechanisms: the global VPP overlays a learnable, axis-like tensor onto the input image to provide structured spatial cues, while the local VPP incorporates position-aware queries to support fine-grained localization. To effectively train our model with spatial guidance, we further introduce VPP-SFT, a curated dataset of 0.6 M high-quality visual grounding samples. Designed in a compact format, it enables efficient training and is significantly smaller than datasets used by other MLLMs (e.g., 21 M samples in MiniGPT-v2), yet still provides a strong performance boost. The resulting model, VPP-LLaVA, not only achieves state-of-the-art results on standard visual grounding benchmarks but also demonstrates strong zero-shot generalization to challenging unseen datasets. Wei Tang 0011, Yanpeng Sun, Qinying Gu, Zechao Li |
IEEE Trans. Multim. | 4 |
| 2026 | THMM-CLIP: Task-Guided Hierarchical Multi-Modal Alignment for Rehearsal-Free Class Incremental LearningabstractClass incremental learning (CIL) requires models to acquire knowledge from sequential tasks containing non-overlapping classes while avoiding catastrophic forgetting. While vision-language foundation models like CLIP demonstrate remarkable potential for CIL through their pre-trained cross-modal alignment capabilities, existing CLIP-based approaches critically overlook the progressive degradation of visual representations in incremental scenarios . Through feature space analysis, we identify a crucial dichotomy : textual embeddings maintain stable discriminative power across sequential tasks, whereas visual features exhibit progressive deterioration manifested by intra-task confusion (ambiguous decision boundaries between co-occurring classes) and inter-task interference (semantic collision between historical and novel categories). To address these dual challenges, we propose task-guided hierarchical multi-modal alignment (THMM-CLIP), a framework that establishes persistent visual-textual coherence through hierarchical multi-modal alignment (HMA) and robust prompt selection (RPS). HMA adapts lightweight task-specific prompt vectors to dynamically recalibrate the CLIP image encoder, thereby achieving: (i) intra-task alignment, (ii) inter-task discriminability alignment, and (iii) global structural alignment with textual features. RPS incorporates a dual-level task identifier that integrates class-level and task-level representative features to ensure precise prompt retrieval during inference. Ablation studies validate all components’ contributions, while t-SNE visualizations, confusion matrices, and Grad-CAM analyses confirm strengthened cross-modal alignment. Yuankang Pan, Zhaoquan Yuan, Xiao Wu 0001, Zechao Li, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2026 | Object Detection under Low-light Conditions via Degradation Learning Driven by Foundation ModelsabstractThe previous approach tackled object detection challenges in low-light scenes by training on images captured in such conditions. However, the limited availability of annotated data in such environments has impeded the development of specialized detectors. As a result, current detectors continue to struggle with images degraded by low illumination. To overcome these limitations, a novel framework for low-light image generation is proposed. These generated images alter only the illumination conditions while preserving the original content, thereby enabling fine-tuning of object detectors. The framework incorporates a trainable degradation module and integrates two frozen Foundation Models: CLIP and a low-light image enhancement network. Under the guidance of CLIP, the degradation module transfers low-light characteristics from textual descriptions to well-lit images by aligning embeddings in a shared vision-language space. This process effectively simulates realistic low-light conditions. With the assistance of the low-light enhancement network, the generated low-light images are recovered. By enforcing similarity between these enhanced versions and the original well-lit counterparts in both the spatial and frequency domains, the content of the generated low-light images is effectively preserved. Extensive experiments on real-world low-light datasets show that fine-tuning detectors using the generated low-light images significantly improves detection performance, even when only well-lit images are available. This validates the effectiveness of the proposed method. Shiyu Xuan, Zechao Li |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | IMAGDressing-v1: Customizable Virtual DressingabstractExisting virtual try-on (VTON) methods provide only limited user control over garment attributes and generally overlook essential factors such as face, pose, and scene context. To address these limitations, we introduce the virtual dressing (VD) task, which aims to synthesize freely editable human images conditioned on fixed garments and optional user-defined inputs. We further propose a comprehensive affinity metric index (CAMI) to quantify the consistency between generated outputs and reference garments. We present IMAGDressing-v1, which leverages a garment-specific U-Net to integrate semantic features from CLIP and texture features from a VAE. To incorporate these garment features into a frozen denoising U-Net for flexible text-driven scene control, we employ a hybrid attention mechanism composed of frozen self-attention and trainable cross-attention layers. IMAGDressing-v1 seamlessly integrates with extension modules, such as ControlNet and IP-Adapter, enabling enhanced diversity and controllability. To alleviate data constraints, we introduce the Interactive Garment Pairing (IGPair) dataset, comprising over 300,000 garment–image pairs and a standardized data assembly pipeline. Extensive experiments demonstrate that IMAGDressing-v1 achieves state-of-the-art performance in controlled human image synthesis. The code and model will be available at https://github.com/muzishen/IMAGDressing. Fei Shen 0004, Xin Jiang 0010, Hu Ye, Cong Wang 0034, Xiaoyu Du 0002, Zechao Li, Jinhui Tang 0001 |
AAAI | 7 |
| 2025 | 3CAD: A Large-Scale Real-World 3C Product Dataset for Unsupervised Anomaly DetectionabstractIndustrial anomaly detection achieves progress thanks to datasets such as MVTec-AD and VisA. However, they suffer from limitations in terms of the number of defect samples, types of defects, and availability of real-world scenes. These constraints inhibit researchers from further exploring the performance of industrial detection with higher accuracy. To this end, we propose a new large-scale anomaly detection dataset called 3CAD, which is derived from real 3C production lines. Specifically, the proposed 3CAD includes eight different types of manufactured parts, totaling 27,039 high-resolution images labeled with pixel-level anomalies. The key features of 3CAD are that it covers anomalous regions of different sizes, multiple anomaly types, and the possibility of multiple anomalous regions and multiple anomaly types per anomaly image. This is the largest and first anomaly detection dataset dedicated to 3C product quality control for community exploration and development. Meanwhile, we introduce a simple yet effective framework for unsupervised anomaly detection: a Coarse-to-Fine detection paradigm with Recovery Guidance (CFRG). To detect small defect anomalies, the proposed CFRG utilizes a coarse-to-fine detection paradigm. Specifically, we utilize a heterogeneous distillation model for coarse localization and then fine localization through a segmentation model. In addition, to better capture normal patterns, we introduce recovery features as guidance. Finally, we report the results of our CFRG framework and popular anomaly detection methods on the 3CAD dataset, demonstrating strong competitiveness and providing a highly challenging benchmark to promote the development of the anomaly detection field. Enquan Yang, Peng Xing, Hanyang Sun, Wenbo Guo 0017, Yuanwei Ma, Zechao Li, Dan Zeng 0001 |
AAAI | 6 |
| 2025 | A Unified Interpretation of Training-Time Out-Of-Distribution Detection
Xin Jiang 0010, Zechao Li |
ICCV | 3 |
| 2025 | Gradient Short-Circuit: Efficient Out-of-Distribution Detection via Feature InterventionabstractOut-of-Distribution (OOD) detection is critical for safely deploying deep models in open-world environments, where inputs may lie outside the training distribution. During inference on a model trained exclusively with In-Distribution (ID) data, we observe a salient gradient phenomenon: around an ID sample, the local gradient directions for "enhancing" that sample's predicted class remain relatively consistent, whereas OOD samples--unseen in training--exhibit disorganized or conflicting gradient directions in the same neighborhood. Motivated by this observation, we propose an inference-stage technique to short-circuit those feature coordinates that spurious gradients exploit to inflate OOD confidence, while leaving ID classification largely intact. To circumvent the expense of recomputing the logits after this gradient short-circuit, we further introduce a local first-order approximation that accurately captures the post-modification outputs without a second forward pass. Experiments on standard OOD benchmarks show our approach yields substantial improvements. Moreover, the method is lightweight and requires minimal changes to the standard inference pipeline, offering a practical path toward robust OOD detection in real-world applications. Jiawei Gu, Ziyue Qiao, Zechao Li |
ICCV | 3 |
| 2025 | ReDet: Effective Real-time Object Detection via Efficient Multi-scale Extraction AggregationabstractReal-time object detection demands detectors that excel in both speed and accuracy. However, existing methods rely on complex Feature Pyramid Networks and computationally intensive post-processing to boost performance, often struggling to balance efficiency and accuracy. In this paper, we propose ReDet, an efficient real-time end-to-end object detection framework that improves detection capability while maintaining low computational cost. Specifically, we propose a Multi-scale Extraction Aggregation module to enhance feature fusion ability with minimal overhead, boosting feature representation across scales. Additionally, a Regression Enhancement Module is incorporated to mitigate the assignment inconsistency between localization and classification, further enhancing detection accuracy with negligible computational cost. Moreover, we introduce a Multi-label Auxiliary Strategy to eliminate the reliance on post-processing by enabling one-to-one label assignment. The experimental results demonstrate that ReDet-T achieves 39.2% AP and 500 FPS on the COCO val2017 dataset, while ReDet-S achieves 45.3% AP and 294 FPS, outperforming many other detectors in both speed and accuracy. Xin Jiang 0010, Lu Jin 0001, Zechao Li |
ICME | 4 |
| 2025 | SpectralGap: Graph-Level Out-of-Distribution Detection via Laplacian Eigenvalue GapsabstractThe task of graph-level out-of-distribution (OOD) detection is crucial for deploying graph neural networks in real-world settings. In this paper, we observe a significant difference in the relationship between the largest and second-largest eigenvalues of the Laplacian matrix for in-distribution (ID) and OOD graph samples: OOD samples often exhibit anomalous spectral gaps (the difference between the largest and second-largest eigenvalues). This observation motivates us to propose SpecGap, an effective post-hoc approach for OOD detection on graphs. SpecGap adjusts features by subtracting the component associated with the second-largest eigenvalue, scaled by the spectral gap, from the high-level features (i.e., X - (λn - λn-1) u_n-1 v_n-1^T). SpecGap achieves state-of-the-art performance across multiple benchmark datasets. We present extensive ablation studies and comprehensive theoretical analyses to support our empirical results. As a parameter-free post-hoc method, SpecGap can be easily integrated into existing graph neural network models without requiring any additional training or model modification. Jiawei Gu, Ziyue Qiao, Zechao Li |
IJCAI | 3 |
| 2025 | OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection is crucial for ensuring the reliability and safety of machine learning models in real-world applications. While zero-shot OOD detection, which requires no training on in-distribution (ID) data, has become feasible with the emergence of vision-language models like CLIP, existing methods primarily focus on semantic matching and fail to fully capture distributional discrepancies. To address these limitations, we propose OT-DETECTOR, a novel framework that employs Optimal Transport (OT) to quantify both semantic and distributional discrepancies between test samples and ID labels. Specifically, we introduce cross-modal transport mass and transport cost as semantic-wise and distribution-wise OOD scores, respectively, enabling more robust detection of OOD samples. Additionally, we present a semantic-aware content refinement (SaCR) module, which utilizes semantic cues from ID labels to amplify the distributional discrepancy between ID and hard OOD samples. Extensive experiments on several benchmarks demonstrate that OT-DETECTOR achieves state-of-the-art performance across various OOD detection tasks, particularly in challenging hard-OOD scenarios. Yu Liu 0158, Hao Tang 0007, Haiqi Zhang 0001, Harry Qin, Zechao Li |
IJCAI | 5 |
| 2025 | Activation Shape Matters: OOD Detection with Norm-Entropy FusionabstractOut-of-distribution (OOD) detection is crucial for safe ML deployment, yet neural networks often exhibit overconfidence on unseen data. While activation norms provide useful OOD signals, they remain vulnerable---OOD inputs can artificially inflate norms through sparse, high-magnitude activations, while valid in-distribution samples with moderate norms may be misclassified. We propose that activation distributional shape, not just magnitude, is essential for robust detection. Our method, Activation Norm and Entropy Weighting (ANEW), combines L2-norm (strength) with Shannon entropy (spread) to distinguish between genuine in-distribution patterns and OOD samples, including adversarial examples mimicking high norms via low-entropy spikes. ANEW requires only a single forward pass without retraining, making it highly practical. Extensive experiments across diverse architectures and benchmarks show ANEW significantly outperforms norm-only baselines, reducing both false positives and false negatives in challenging scenarios. Code available upon acceptance. Jiawei Gu, Ziyue Qiao, Zechao Li |
ACM Multimedia | 3 |
| 2025 | Time-IC: Empowering MLLM with Interleaved Context for Temporal-Sensitive Video UnderstandingabstractTemporal-sensitive video tasks, such as dense video captioning, require models to describe event content in detail and generate timestamps marking their occurrence. Existing approaches tend to over-rely on visuals while neglecting the role of rich contexts, limiting temporal reasoning and semantic comprehension in videos. To this end, we propose Time-IC, a multimodal large language model that incorporates video context for improved understanding. Specifically, Time-IC integrates both internal context (timestamps, speech transcripts) and external context (titles, author descriptions), providing complementary semantic cues and high-level semantic priors. More importantly, it unifies context, video frames, and task instructions into a temporally aligned, interleaved sequence. With instruction tuning, this design enables the model to associate contextual elements and generate comprehensive responses across diverse tasks. Extensive experiments on widely-used benchmarks (YouCookII, QVHighlights) across three tasks (dense video captioning, temporal grounding, highlight detection) demonstrate the effectiveness and flexibility of the proposed Time-IC. Henghao Zhao, Rui Yan 0010, Zechao Li |
MMAsia | 4 |
| 2025 | FedMGP: Personalized Federated Learning with Multi-Group Text-Visual PromptsabstractIn this paper, we introduce FedMGP, a new paradigm for personalized federated prompt learning in vision-language models (VLMs). Existing federated prompt learning (FPL) methods often rely on a single, text-only prompt representation, which leads to client-specific overfitting and unstable aggregation under heterogeneous data distributions. Toward this end, FedMGP equips each client with multiple groups of paired textual and visual prompts, enabling the model to capture diverse, fine-grained semantic and instance-level cues. A diversity loss is introduced to drive each prompt group to specialize in distinct and complementary semantic aspects, ensuring that the groups collectively cover a broader range of local characteristics.During communication, FedMGP employs a dynamic prompt aggregation strategy based on similarity-guided probabilistic sampling: each client computes the cosine similarity between its prompt groups and the global prompts from the previous round, then samples s groups via a softmax-weighted distribution. This soft selection mechanism preferentially aggregates semantically aligned knowledge while still enabling exploration of underrepresented patterns—effectively balancing the preservation of common knowledge with client-specific features. Notably, FedMGP maintains parameter efficiency by redistributing a fixed prompt capacity across multiple groups, achieving state-of-the-art performance with the lowest communication parameters (5.1k) among all federated prompt learning methods. Theoretical analysis shows that our dynamic aggregation strategy promotes robust global representation learning by reinforcing shared semantics while suppressing client-specific noise. Extensive experiments demonstrate that FedMGP consistently outperforms prior approaches in both personalization and domain generalization across diverse federated vision-language benchmarks.The code will be released on https://github.com/weihao-bo/FedMGP.git. Weihao Bo, Yanpeng Sun, Xinyu Zhang 0017, Zechao Li |
NeurIPS | 5 |
| 2025 | Refining Norms: A Post-hoc Framework for OOD Detection in Graph Neural NetworksabstractGraph Neural Networks (GNNs) are increasingly deployed in mission-critical tasks, yet they often encounter inputs that lie outside their training distribution, leading to unreliable or overconfident predictions. To address this limitation, we present RAGNOR (Robust Aggregation Graph Norm for Outlier Recognition), a post-hoc approach that leverages embedding norms for robust out-of-distribution (OOD) detection on both node-level and graph-level tasks. Unlike previous methods designed primarily for image domains, RAGNOR directly tackles the relational challenges intrinsic to graphs: local contamination by anomalous neighbors, disparate norm scales across classes or roles, and insufficient references for boundary or low-degree nodes. By combining global Z-score normalization, median-based local aggregation, and multi-hop blending, RAGNOR effectively refines raw norm signals into robust OOD scores while incurring minimal overhead and requiring no retraining of the original GNN. Experimental evaluations on multiple benchmarks demonstrate that RAGNOR not only achieves competitive or superior detection performance compared to alternative techniques, but also provides an intuitive, modular design that can be readily integrated into existing graph pipelines. Jiawei Gu, Ziyue Qiao, Zechao Li |
NeurIPS | 3 |
| 2025 | Revitalizing SVD for Global Covariance Pooling: Halley's Method to Overcome Over-FlatteningabstractGlobal Covariance Pooling (GCP) has garnered increasing attention in visual recognition tasks, where second-order statistics frequently yield stronger representations than first-order approaches. However, two main streams of GCP---Newton--Schulz-based iSQRT-COV and exact or near-exact SVD methods---struggle at opposite ends of the training spectrum. While iSQRT-COV stabilizes early learning by avoiding large gradient explosions, it over-compresses significant eigenvalues in later stages, causing an \emph{over-flattening} phenomenon that stalls final accuracy. In contrast, SVD-based methods excel at preserving the high-eigenvalue structure essential for deep networks but suffer from sensitivity to small eigenvalue gaps early on. We propose \textbf{Halley-SVD}, a high-order iterative method that unites the smooth gradient advantages of iSQRT-COV with the late-stage fidelity of SVD. Grounded in Halley's iteration, our approach obviates explicit divisions by $(\lambda_i - \lambda_j)$ and forgoes threshold- or polynomial-based heuristics. As a result, it prevents both early gradient explosions and the excessive compression of large eigenvalues. Extensive experiments on CNNs and transformer architectures show that Halley-SVD consistently and robustly outperforms iSQRT-COV at large model scales and batch sizes, achieving higher overall accuracy without mid-training switches or custom truncations. This work provides a new solution to the long-standing dichotomy in GCP, illustrating how high-order methods can balance robustness and spectral precision to fully harness the representational power of modern deep networks. Jiawei Gu, Ziyue Qiao, Zechao Li |
NeurIPS | 4 |
| 2025 | CSGO: Content-Style Composition in Text-to-Image GenerationabstractThe advancement of image style transfer has been fundamentally constrained by the absence of large-scale, high-quality datasets with explicit content-style-stylized supervision. Existing methods predominantly adopt training-free paradigms (e.g., image inversion), which limit controllability and generalization due to the lack of structured triplet data. To bridge this gap, we design a scalable and automated pipeline that constructs and purifies high-fidelity content-style-stylized image triplets. Leveraging this pipeline, we introduce IMAGStyle—the first large-scale dataset of its kind, containing 210K diverse and precisely aligned triplets for style transfer research. Empowered by IMAGStyle, we propose CSGO, a unified, end-to-end trainable framework that decouples content and style representations via independent feature injection. CSGO jointly supports image-driven style transfer, text-driven stylized generation, and text-editing-driven stylized synthesis within a single architecture. Extensive experiments show that CSGO achieves state-of-the-art controllability and fidelity, demonstrating the critical role of structured synthetic data in unlocking robust and generalizable style transfer. Source code: \url{https://github.com/instantX-research/CSGO} Peng Xing, Yanpeng Sun, Qixun Wang 0001, Hao Ai, Jen-Yuan Huang, Zechao Li |
NeurIPS | 8 |
| 2025 | A recover-then-discriminate framework for robust anomaly detection
Peng Xing, Jinhui Tang 0001, Zechao Li |
Sci. China Inf. Sci. | 4 |
| 2025 | SSA: semantic structure aware inference on CNN networks for weakly pixel-wise dense predictions without cost
Yanpeng Sun, Zechao Li |
Frontiers Comput. Sci. | 2 |
| 2025 | Imbuing, Enrichment and Calibration: Leveraging Language for Unseen Domain Extension
Chenyi Jiang, Jianqin Zhao, Jingjing Deng 0001, Zechao Li, Haofeng Zhang 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Dynamic semantic prototype perception for text-video retrieval
Henghao Zhao, Rui Yan 0001, Zechao Li |
Image Vis. Comput. | 3 |
| 2025 | Imaginary-Connected Embedding in Complex Space for Unseen Attribute-Object DiscriminationabstractCompositional Zero-Shot Learning (CZSL) aims to recognize novel compositions of seen primitives. Prior studies have attempted to either learn primitives individually (non-connected) or establish dependencies among them in the composition (fully-connected). In contrast, human comprehension of composition diverges from the aforementioned methods as humans possess the ability to make composition-aware adaptation for these primitives, instead of inferring them rigidly through the aforementioned methods. However, developing a comprehension of compositions akin to human cognition proves challenging within the confines of real space. This arises from the limitation of real-space-based methods, which often categorize attributes, objects, and compositions using three independent measures, without establishing a direct dynamic connection. To tackle this challenge, we expand the CZSL distance metric scheme to encompass complex spaces to unify the independent measures, and we establish an imaginary-connected embedding in complex space to model human understanding of attributes. To achieve this representation, we introduce an innovative visual bias-based attribute extraction module that selectively extracts attributes based on object prototypes. As a result, we are able to incorporate phase information in training and inference, serving as a metric for attribute-object dependencies while preserving the independent acquisition of primitives. We evaluate the effectiveness of our proposed approach on three benchmark datasets, illustrating its superiority compared to baseline methods. Chenyi Jiang, Yang Long 0001, Zechao Li, Haofeng Zhang 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object DetectionabstractRGB-Thermal Salient Object Detection (RGB-T SOD) aims to pinpoint prominent objects within aligned pairs of visible and thermal infrared images. A key challenge lies in bridging the inherent disparities between RGB and Thermal modalities for effective saliency map prediction. Traditional encoder-decoder architectures, while designed for cross-modality feature interactions, may not have adequately considered the robustness against noise originating from defective modalities, thereby leading to suboptimal performance in complex scenarios. Inspired by hierarchical human visual systems, we propose the ConTriNet, a robust Confluent Triple-Flow Network employing a "Divide-and-Conquer" strategy. This framework utilizes a unified encoder with specialized decoders, each addressing different subtasks of exploring modality-specific and modality-complementary information for RGB-T SOD, thereby enhancing the final saliency map prediction. Specifically, ConTriNet comprises three flows: two modality-specific flows explore cues from RGB and Thermal modalities, and a third modality-complementary flow integrates cues from both modalities. ConTriNet presents several notable advantages. It incorporates a Modality-induced Feature Modulator (MFM) in the modality-shared union encoder to minimize inter-modality discrepancies and mitigate the impact of defective samples. Additionally, a foundational Residual Atrous Spatial Pyramid Module (RASPM) in the separated flows enlarges the receptive field, allowing for the capture of multi-scale contextual information. Furthermore, a Modality-aware Dynamic Aggregation Module (MDAM) in the modality-complementary flow dynamically aggregates saliency-related cues from both modality-specific flows. Leveraging the proposed parallel triple-flow framework, we further refine saliency maps derived from different flows through a flow-cooperative fusion strategy, yielding a high-quality, full-resolution saliency map for the final prediction. To evaluate the robustness and stability of our approach, we collect a comprehensive RGB-T SOD benchmark, VT-IMAG, covering various real-world challenging scenarios. Extensive experiments on public benchmarks and our VT-IMAG dataset demonstrate that ConTriNet consistently outperforms state-of-the-art competitors in both common and challenging scenarios, even when dealing with incomplete modality data. The code and VT-IMAG will be available at: https://cser-tang-hao.github.io/contrinet.html. Hao Tang 0007, Zechao Li, Shengfeng He, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Parameter AdapterabstractThe remarkable advancement in text-to-image generation models significantly boosts the research in ID customization generation. However, existing personalization methods cannot simultaneously satisfy high-fidelity and low-costs requirements. Their main bottleneck lies in the additional prompt image encoder (i.e., CLIP vision encoder), which produces weak alignment signals with the text-to-image model that may lose face information and is not well 'absorbed' by the text-to-image model. Towards this end, we propose Inv-Adapter, which first introduces a more reasonable and efficient token representation of ID image features and introduces a lightweight parameter adaptor to inject ID features. Specifically, our Inv-Adapter extracts diffusion-domain representations of ID images utilizing a pre-trained text-to-image model via DDIM image inversion, without an additional image encoder. Benefiting from the high alignment of the extracted ID prompt features and the intermediate features of the text-to-image model, we then introduce a lightweight attention adapter to embed them efficiently into the base text-to-image model. We conduct extensive experiments on different text-to-image models to assess ID fidelity, generation loyalty, speed, training costs, model scale and generalization ability in scenarios of general object, all of which show that the proposed Inv-Adapter is highly competitive in ID customization generation and model scale. Peng Xing, Ning Wang 0020, Jianbo Ouyang, Zechao Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Modality-Specific Interactive Attack for Vision-Language Pre-Training ModelsabstractRecent advances have heightened the interest in the adversarial transferability of Vision-Language Pre-training (VLP) models. However, most existing strategies constrained by two persistent limitations: suboptimal utilization of cross-modal interactive information, and inherent discrepancies across hierarchical textual representation. To address these challenges, we propose the Modality-Specific Interactive Attack (MSI-Attack), a novel approach that integrates semantic-level image perturbations with embedding-level text perturbations, all while maintaining minimal inter-modal constraints. In our image attack methodology, we introduce Multi-modal Integrated Gradients (MIG) to guide perturbations toward the core semantics of images, enriched by their associated deeply text information. This technique enhances transferability by capturing consistent features across various models, thereby effectively misleading similar-model perception areas. Additionally, we employ a momentum iteration strategy in conjunction with MIG, which amalgamates current and historical gradients to expedite the perturbation updates. For text attacks, we streamline the perturbation process by operating exclusively at the embedding level. This reduces semantic gaps across hierarchical structures and significantly enhances the generalizability of adversarial text. Moreover, we delve deeper into how semantic perturbations with varying degrees of similarity affect the overall attack effectiveness. Our experimental results on image-text retrieval tasks using the multi-modal datasets Flickr30K and MSCOCO underscore the efficacy of MSI-Attack. Our method achieves superior performance, setting a new state-of-the-art benchmark, all without the need for additional mechanisms. Haiqi Zhang 0001, Hao Tang 0007, Yanpeng Sun, Shengfeng He, Zechao Li |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Causal Inference Hashing for Long-Tailed Image RetrievalabstractIn hashing-based long-tailed image retrieval, the dominance of data-rich head classes often hinders the learning of effective hash codes for data-poor tail classes due to inherent long-tailed bias. Interestingly, this bias also contains valuable prior knowledge by revealing inter-class dependencies, which can be beneficial for hash learning. However, previous methods have not thoroughly analyzed this tangled negative and positive effects of long-tailed bias from a causal inference perspective. In this paper, we propose a novel hash framework that employs causal inference to disentangle detrimental bias effects from beneficial ones. To capture good bias in long-tailed datasets, we construct hash mediators that conserve valuable prior knowledge from class centers. Furthermore, we propose a de-biased hash loss To enhance the beneficial bias effects while mitigating adverse ones, leading to more discriminative hash codes. Specifically, this loss function leverages the beneficial bias captured by hash mediators to support accurate class label prediction, while mitigating harmful bias by blocking its causal path to the hash codes and refining predictions through backdoor adjustment. Extensive experimental results on four widely used datasets demonstrate that the proposed method improves retrieval performance against the state-of-the-art methods by large margins. The source code is available at https://github.com/IMAG-LuJin/CIH. Lu Jin 0001, Zhengyun Lu, Zechao Li, Yonghua Pan, Longquan Dai, Jinhui Tang 0001, Ramesh Jain 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Exploring Effective Factors for Improving Visual In-Context LearningabstractThe In-Context Learning (ICL) is to understand a new task via a few demonstrations (aka. prompt) and predict new inputs without tuning the models. While it has been widely studied in NLP, it is still a relatively new area of research in computer vision. To reveal the factors influencing the performance of visual in-context learning, this paper shows that Prompt Selection and Prompt Fusion are two major factors that have a direct impact on the inference performance of visual in-context learning. Prompt selection is the process of selecting the most suitable prompt for query image. This is crucial because high-quality prompts assist large-scale visual models in rapidly and accurately comprehending new tasks. Prompt fusion involves combining prompts and query images to activate knowledge within large-scale visual models. However, altering the prompt fusion method significantly impacts its performance on new tasks. Based on these findings, we propose a simple framework prompt-SelF to improve visual in-context learning. Specifically, we first use the pixel-level retrieval method to select a suitable prompt, and then use different prompt fusion methods to activate diverse knowledge stored in the large-scale vision model, and finally, ensemble the prediction results obtained from different prompt fusion methods to obtain the final prediction results. We conducted extensive experiments on single-object segmentation and detection tasks to demonstrate the effectiveness of prompt-SelF. Remarkably, prompt-SelF has outperformed OSLSM method-based meta-learning in 1-shot segmentation for the first time. This indicated the great potential of visual in-context learning. The source code and models will be available at https://github.com/syp2ysy/prompt-SelF. Yanpeng Sun, Qiang Chen 0007, Jian Wang 0066, Jingdong Wang 0001, Zechao Li |
IEEE Trans. Image Process. | 5 |
| 2025 | Multi-View Clustering With Incremental Instances and ViewsabstractMulti-view clustering (MVC) has attracted increasing attention with the emergence of various data collected from multiple sources. In real-world dynamic environment, instances are continually gathered, and the number of views expands as new data sources become available. Learning for such simultaneous increment of instances and views, particularly in unsupervised scenarios, is crucial yet underexplored. In this paper, we address this problem by proposing a novel MVC method with Incremental Instances and Views, MVC-IIV for short. MVC-IIV contains two stages, an initial stage and an incremental stage. In the initial stage, a basic latent multi-view subspace clustering model is constructed to handle existing data, which can be viewed as traditional static MVC. In the incremental stage, the previously trained model is reused to guide learning for newly arriving instances with new views, transferring historical knowledge while avoiding redundant computations. In specific, we design and reuse two modules, i.e., multi-view embedding module for low-dimensional representation learning, and consensus centroids module for cluster probability learning. By adding consistency regularization on the two modules, the knowledge acquired from previous data is used, which not only enhances the exploration within current data batch, but also extracts the between-batch data correlations. The proposed model can be efficiently solved with linear space and time complexity. Extensive experiments demonstrate the effectiveness and efficiency of our method compared with the state-of-the-art approaches. Chao Zhang 0078, Zhi Wang 0001, Xiuyi Jia, Zechao Li, Chunlin Chen 0001, Huaxiong Li |
IEEE Trans. Image Process. | 4 |
| 2025 | Cross-Domain Few-Shot Medical Image Segmentation via Dynamic Semantic MatchingabstractCross-domain few-shot medical image segmentation (CDFSMIS) presents the fundamental challenge of segmenting novel anatomical or tissue structures on unfamiliar medical imaging domains with limited annotated data. In this paper, we conduct an in-depth investigation of CDFSMIS and identify two critical observations: 1) the conventional matching mechanisms from existing few-shot models are particularly vulnerable to discrepancies in local characteristics between different domains and 2) the semantic representations learned from source domains often lack robustness when generalizing to unfamiliar target domains. Motivated by these insights, we propose a novel Dynamic Semantic Matching (DSM) framework that addresses these challenges through a three-component approach. First, we design a support-query feature re-weighting (SFR) mechanism that leverages multilevel hidden features to suppress domain-specific contents. Second, we introduce a dynamic semantic information selection (DSIS) strategy that adaptively identifies and combines domain-robust channels to construct generalizable representations. Third, we develop a dual-perspective semantic center calculation method to address the inherent texture imbalance in medical images. Extensive experiments on four unfamiliar target domains (MS-CMR, PI-PMR, Chest-X-Ray and ISIC2018) demonstrate that our approach significantly outperforms state-of-the-art few-shot segmentation and cross-domain few-shot segmentation models, validating the effectiveness of DSM in simultaneously addressing domain generalization and semantic matching challenges in medical image segmentation. The source code is available at https://github.com/YazhouZhu19/DSM. Yazhou Zhu 0001, Tao Zhou 0002, Zechao Li, Haofeng Zhang 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | AFANet: Adaptive Frequency-Aware Network for Weakly-Supervised Few-Shot Semantic SegmentationabstractFew-shot learning aims to recognize novel concepts by leveraging prior knowledge learned from a few samples. However, for visually intensive tasks such as few-shot semantic segmentation, pixel-level annotations are time-consuming and costly. Therefore, in this paper, we utilize the more challenging image-level annotations and propose an adaptive frequency-aware network (AFANet) for weakly-supervised few-shot semantic segmentation (WFSS). Specifically, we first propose a cross-granularity frequency-aware module (CFM) that decouples RGB images into high-frequency and low-frequency distributions and further optimizes semantic structural information by realigning them. Unlike most existing WFSS methods using the textual information from the multi-modal language-vision model, e.g., CLIP, in an offline learning manner, we further propose a CLIP-guided spatial-adapter module (CSM), which performs spatial domain adaptive transformation on textual information through online learning, thus providing enriched cross-modal semantic information for CFM. Extensive experiments on the Pascal-5iand COCO-20idatasets demonstrate that AFANet has achieved state-of-the-art performance. Jiaqi Ma 0006, Guosen Xie, Fang Zhao 0006, Zechao Li |
IEEE Trans. Multim. | 4 |
| 2025 | Fast Disentangled Slim Tensor Learning for Multi-View ClusteringabstractTensor-based multi-view clustering has recently received significant attention due to its exceptional ability to explore cross-view high-order correlations. However, most existing methods still encounter some limitations. (1) Most of them explore the correlations among different affinity matrices, making them unscalable to large-scale data. (2) Although some methods address it by introducing bipartite graphs, they may result in sub-optimal solutions caused by an unstable anchor selection process. (3) They generally ignore the negative impact of latent semantic-unrelated information in each view. To tackle these issues, we propose a new approach termed fast Disentangled Slim Tensor Learning (DSTL) for multi-view clustering. Instead of focusing on the multi-view graph structures, DSTL directly explores the high-order correlations among multi-view latent semantic representations based on matrix factorization. To alleviate the negative influence of feature redundancy, inspired by robust PCA, DSTL disentangles the latent low-dimensional representation into a semantic-unrelated part and a semantic-related part for each view. Subsequently, two slim tensors are constructed with tensor-based regularization. To further enhance the quality of feature disentanglement, the semantic-related representations are aligned across views through a consensus alignment indicator. Our proposed model is computationally efficient and can be solved effectively. Extensive experiments demonstrate the superiority and efficiency of DSTL over state-of-the-art approaches. Deng Xu, Chao Zhang 0078, Zechao Li, Chunlin Chen 0001, Huaxiong Li |
IEEE Trans. Multim. | 3 |
| 2025 | Relational Consistency Induced Self-Supervised Hashing for Image RetrievalabstractThis article proposes a new hashing framework named relational consistency induced self-supervised hashing (RCSH) for large-scale image retrieval. To capture the potential semantic structure of data, RCSH explores the relational consistency between data samples in different spaces, which learns reliable data relationships in the latent feature space and then preserves the learned relationships in the Hamming space. The data relationships are uncovered by learning a set of prototypes that group similar data samples in the latent feature space. By uncovering the semantic structure of the data, meaningful data-to-prototype and data-to-data relationships are jointly constructed. The data-to-prototype relationships are captured by constraining the prototype assignments generated from different augmented views of an image to be the same. Meanwhile, these data-to-prototype relationships are preserved to learn informative compact hash codes by matching them with these reliable prototypes. To accomplish this, a novel dual prototype contrastive loss is proposed to maximize the agreement of prototype assignments in the latent feature space and Hamming space. The data-to-data relationships are captured by enforcing the distribution of pairwise similarities in the latent feature space and Hamming space to be consistent, which makes the learned hash codes preserve meaningful similarity relationships. Extensive experimental results on four widely used image retrieval datasets demonstrate that the proposed method significantly outperforms the state-of-the-art methods. Besides, the proposed method achieves promising performance in out-of-domain retrieval tasks, which shows its good generalization ability. The source code and models are available at https://github.com/IMAG-LuJin/RCSH. Lu Jin 0001, Zechao Li, Yonghua Pan, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Knowledge-Guided Semantic Transfer Network for Few-Shot Image RecognitionabstractDeep learning-based models have been shown to outperform human beings in many computer vision tasks with massive available labeled training data in learning. However, humans have an amazing ability to easily recognize images of novel categories by browsing only a few examples of these categories. In this case, few-shot learning comes into being to make machines learn from extremely limited labeled examples. One possible reason why human beings can well learn novel concepts quickly and efficiently is that they have sufficient visual and semantic prior knowledge. Toward this end, this work proposes a novel knowledge-guided semantic transfer network (KSTNet) for few-shot image recognition from a supplementary perspective by introducing auxiliary prior knowledge. The proposed network jointly incorporates vision inferring, knowledge transferring, and classifier learning into one unified framework for optimal compatibility. A category-guided visual learning module is developed in which a visual classifier is learned based on the feature extractor along with the cosine similarity and contrastive loss optimization. To fully explore prior knowledge of category correlations, a knowledge transfer network is then developed to propagate knowledge information among all categories to learn the semantic-visual mapping, thus inferring a knowledge-based classifier for novel categories from base categories. Finally, we design an adaptive fusion scheme to infer the desired classifiers by effectively integrating the above knowledge and visual information. Extensive experiments are conducted on two widely used Mini-ImageNet and Tiered-ImageNet benchmarks to validate the effectiveness of KSTNet. Compared with the state of the art, the results show that the proposed method achieves favorable performance with minimal bells and whistles, especially in the case of one-shot learning. Zechao Li, Hao Tang 0007, Zhimao Peng, Guo-Jun Qi, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Robust Low-Rank Latent Feature Analysis for Spatiotemporal Signal RecoveryabstractWireless sensor network (WSN) is an emerging and promising developing area in the intelligent sensing field. Due to various factors like sudden sensors breakdown or saving energy by deliberately shutting down partial nodes, there are always massive missing entries in the collected sensing data from WSNs. Low-rank matrix approximation (LRMA) is a typical and effective approach for pattern analysis and missing data recovery in WSNs. However, existing LRMA-based approaches ignore the adverse effects of outliers inevitably mixed with collected data, which may dramatically degrade their recovery accuracy. To address this issue, this article innovatively proposes a latent feature analysis (LFA) based spatiotemporal signal recovery (STSR) model, named LFA-STSR. Its main idea is twofold: 1) incorporating the spatiotemporal correlation into an LFA model as the regularization constraint to improve its recovery accuracy and 2) aggregating the -norm into the loss part of an LFA model to improve its robustness to outliers. As such, LFA-STSR can accurately recover missing data based on partially observed data mixed with outliers in WSNs. To evaluate the proposed LFA-STSR model, extensive experiments have been conducted on four real-world WSNs datasets. The results demonstrate that LFA-STSR significantly outperforms the related six state-of-the-art models in terms of both recovery accuracy and robustness to outliers. Di Wu 0056, Zechao Li, Zhikai Yu, Yi He 0007, Xin Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | ADPS: Asymmetric Distillation Postsegmentation for Image Anomaly DetectionabstractKnowledge distillation-based anomaly detection (KDAD) methods rely on the teacher-student paradigm to detect and segment anomalous regions by contrasting the unique features extracted by both networks. However, existing KDAD methods suffer from two main limitations: 1) the student network can effortlessly replicate the teacher network's representations and 2) the features of the teacher network serve solely as a "reference standard" and are not fully leveraged. Toward this end, we depart from the established paradigm and instead propose an innovative approach called asymmetric distillation postsegmentation (ADPS). Our ADPS employs an asymmetric distillation paradigm that takes distinct forms of the same image as the input of the teacher-student networks, driving the student network to learn discriminating representations for anomalous regions. Meanwhile, a customized Weight Mask Block (WMB) is proposed to generate a coarse anomaly localization mask that transfers the distilled knowledge acquired from the asymmetric paradigm to the teacher network. Equipped with WMB, the proposed postsegmentation module (PSM) can effectively detect and segment abnormal regions with fine structures and clear boundaries. Experimental results demonstrate that the proposed ADPS outperforms the state-of-the-art methods in detecting and segmenting anomalies. Surprisingly, ADPS significantly improves average precision (AP) metric by $\mathbf {9}\%$ and $\mathbf {20}\%$ on the MVTec anomaly detection (AD) and KolektorSDD2 datasets, respectively. Peng Xing, Hao Tang 0007, Jinhui Tang 0001, Zechao Li |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight DetectionabstractVideo moment retrieval and highlight detection have received attention in the current era of video content proliferation, aiming to localize moments and estimate clip relevances based on user-specific queries. Most existing methods approach these challenges from a discriminative learning perspective, focusing on learning the correspondence between query and activity boundary locations through complex cross-modal interactions. However, the continuous nature of video content often results in unclear boundaries between temporal events. This boundary ambiguity may confuse models, resulting in the subpar performance in predicting target boundaries. To alleviate this problem, we propose to solve the two tasks jointly from the perspective of denoising generation. Moreover, the target boundary can be localized clearly by iterative refinement from coarse to fine. Specifically, a novel framework, DiffusionVMR, is proposed to redefine the two tasks as a unified conditional denoising generation process by combining the diffusion model. During training, the Gaussian noise is added to corrupt the ground truth (GT), with noisy candidates produced as input. The model is trained to reverse this noise addition process. In the inference phase, DiffusionVMR initiates directly from Gaussian noise and progressively refines the proposals from the noise to the meaningful output. Notably, the proposed DiffusionVMR inherits the advantages of diffusion models that allow for iteratively refined results during inference, enhancing the boundary transition from coarse to fine. Furthermore, the training and inference of DiffusionVMR are decoupled. An arbitrary setting can be used in DiffusionVMR during inference without consistency with the training phase. Extensive experiments conducted on five widely used benchmarks (i.e., QVHighlight, Charades-STA, TACoS, YouTubeHighlights, and TVSum) across two tasks (moment retrieval and/or highlight detection) demonstrate the effectiveness and flexibility of the proposed DiffusionVMR. Henghao Zhao, Qinghong Lin, Rui Yan 0001, Zechao Li |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Delving into Multimodal Prompting for Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) involves categorizing fine subdivisions within a broader category, which poses challenges due to subtle inter-class discrepancies and large intra-class variations. However, prevailing approaches primarily focus on uni-modal visual concepts. Recent advancements in pre-trained vision-language models have demonstrated remarkable performance in various high-level vision tasks, yet the applicability of such models to FGVC tasks remains uncertain. In this paper, we aim to fully exploit the capabilities of cross-modal description to tackle FGVC tasks and propose a novel multimodal prompting solution, denoted as MP-FGVC, based on the contrastive language-image pertaining (CLIP) model. Our MP-FGVC comprises a multimodal prompts scheme and a multimodal adaptation scheme. The former includes Subcategory-specific Vision Prompt (SsVP) and Discrepancy-aware Text Prompt (DaTP), which explicitly highlights the subcategory-specific discrepancies from the perspectives of both vision and language. The latter aligns the vision and text prompting elements in a common semantic space, facilitating cross-modal collaborative reasoning through a Vision-Language Fusion Module (VLFM) for further improvement on FGVC. Moreover, we tailor a two-stage optimization strategy for MP-FGVC to fully leverage the pre-trained CLIP model and expedite efficient adaptation for FGVC. Extensive experiments conducted on four FGVC datasets demonstrate the effectiveness of our MP-FGVC. Xin Jiang 0010, Hao Tang 0007, Junyao Gao 0002, Xiaoyu Du 0002, Shengfeng He, Zechao Li |
AAAI | 6 |
| 2024 | Learning Cluster-Wise Anchors for Multi-View ClusteringabstractDue to its effectiveness and efficiency, anchor based multi-view clustering (MVC) has recently attracted much attention. Most existing approaches try to adaptively learn anchors to construct an anchor graph for clustering. However, they generally focus on improving the diversity among anchors by using orthogonal constraint and ignore the underlying semantic relations, which may make the anchors not representative and discriminative enough. To address this problem, we propose an adaptive Cluster-wise Anchor learning based MVC method, CAMVC for short. We first make an anchor cluster assumption that supposes the prior cluster structure of target anchors by pre-defining a consensus cluster indicator matrix. Based on the prior knowledge, an explicit cluster structure of latent anchors is enforced by learning diverse cluster centroids, which can explore both inter-cluster diversity and intra-cluster consistency of anchors, and improve the subspace representation discrimination. Extensive results demonstrate the effectiveness and superiority of our proposed method compared with some state-of-the-art MVC approaches. Chao Zhang 0078, Xiuyi Jia, Zechao Li, Chunlin Chen 0001, Huaxiong Li |
AAAI | 3 |
| 2024 | Soft Knowledge Prompt: Help External Knowledge Become a Better Teacher to Instruct LLM in Knowledge-based VQAabstractLLM has achieved impressive performance on multi-modal tasks, which have received everincreasing research attention.Recent research focuses on improving prediction performance and reliability (e.g., addressing the hallucination problem).They often prepend relevant external knowledge to the input text as an extra prompt.However, these methods would be affected by the noise in the knowledge and the context length limitation of LLM.In our work, we focus on making better use of external knowledge and propose a method to actively extract valuable information in the knowledge to produce the latent vector as a soft prompt, which is then fused with the image embedding to form a knowledge-enhanced context to instruct LLM.The experimental results on knowledge-based VQA benchmarks show that the proposed method enjoys better utilization of external knowledge and helps the model achieve better performance. Qunbo Wang, Ruyi Ji, Tianhao Peng 0002, Wenjun Wu 0001, Zechao Li, Jing Liu 0001 |
ACL (1) | 5 |
| 2024 | VRP-SAM: SAM with Visual Reference PromptabstractIn this paper, we propose a novel Visual Reference Prompt (VRP) encoder that empowers the Segment Any-thing Model (SAM) to utilize annotated reference images as prompts for segmentation, creating the VRP-SAM model. In essence, VRP-SAM can utilize annotated reference images to comprehend specific objects and perform segmen-tation of specific objects in target image. It is note that the VRP encoder can support a variety of annotation for-mats for reference images, including point, box, scribble, and mask. VRP-SAM achieves a breakthrough within the SAM framework by extending its versatility and applicabil-ity while preserving SAM's inherent strengths, thus enhancing user-friendliness. To enhance the generalization abil-ity of VRP-SAM, the VRP encoder adopts a meta-learning strategy. To validate the effectiveness of VRP-SAM, we con-ducted extensive empirical studies on the Pascal and COCO datasets. Remarkably, VRP-SAM achieved state-of-the-art performance in visual reference segmentation with mini-mal learnable parameters. Furthermore, VRP-SAM demon-strates strong generalization capabilities, allowing it to per-form segmentation of unseen objects and enabling cross-domain segmentation. The source code and models will be available at https://github.com/syp2ysy/VRP-SAM Yanpeng Sun, Shan Zhang 0002, Xinyu Zhang 0017, Qiang Chen 0007, Errui Ding, Jingdong Wang 0001, Zechao Li |
CVPR | 9 |
| 2024 | DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval GuidelinesabstractFine-grained image retrieval (FGIR) is to learn visual representations that distinguish visually similar objects while maintaining generalization. Existing methods propose to generate discriminative features, but rarely consider the particularity of the FGIR task itself. This paper presents a meticulous analysis leading to the proposal of practical guidelines to identify subcategory-specific discrepancies and generate discriminative features to design effective FGIR models. These guidelines include emphasizing the object (G1), highlighting subcategory-specific discrepancies (G2), and employing effective training strategy (G3). Following G1 and G2, we design a novel Dual Visual Filtering mechanism for the plain visual transformer, denoted as DVF, to capture subcategory-specific discrepancies. Specifically, the dual visual filtering mechanism comprises an object-oriented module and a semantic-oriented module. These components serve to magnify objects and identify discriminative regions, respectively. Following G3, we implement a discriminative model training strategy to improve the discriminability and generalization ability of DVF. Extensive analysis and ablation studies confirm the efficacy of our proposed guidelines. Without bells and whistles, the proposed DVF achieves state-of-the-art performance on three widely-used fine-grained datasets in closed-set and open-set settings. Xin Jiang 0010, Hao Tang 0007, Rui Yan 0010, Jinhui Tang 0001, Zechao Li |
ACM Multimedia | 5 |
| 2024 | 3DPCP-Net: A Lightweight Progressive 3D Correspondence Pruning Network for Accurate and Efficient Point Cloud Registration
Zechao Li |
ACM Multimedia | 2 |
| 2024 | TMM-CLIP: Task-guided Multi-Modal Alignment for Rehearsal-Free Class Incremental Learning
Yuankang Pan, Zhaoquan Yuan, Xiao Wu 0001, Zechao Li, Changsheng Xu |
MMAsia | 4 |
| 2024 | Context Disentangling and Prototype Inheriting for Robust Visual GroundingabstractVisual grounding (VG) aims to locate a specific target in an image based on a given language query. The discriminative information from context is important for distinguishing the target from other objects, particularly for the targets that have the same category as others. However, most previous methods underestimate such information. Moreover, they are usually designed for the standard scene (without any novel object), which limits their generalization to the open-vocabulary scene. In this paper, we propose a novel framework with context disentangling and prototype inheriting for robust visual grounding to handle both scenes. Specifically, the context disentangling disentangles the referent and context features, which achieves better discrimination between them. The prototype inheriting inherits the prototypes discovered from the disentangled visual features by a prototype bank to fully utilize the seen data, especially for the open-vocabulary scene. The fused features, obtained by leveraging Hadamard product on disentangled linguistic and visual features of prototypes to avoid sharp adjusting the importance between the two types of features, are then attached with a special token and feed to a vision Transformer encoder for bounding box regression. Extensive experiments are conducted on both standard and open-vocabulary scenes. The performance comparisons indicate that our method outperforms the state-of-the-art methods in both scenarios. Wei Tang 0011, Liang Li 0003, Xuejing Liu, Lu Jin 0001, Jinhui Tang 0001, Zechao Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Learning Contrastive Self-Distillation for Ultra-Fine-Grained Visual Categorization Targeting Limited SamplesabstractIn the field of intelligent multimedia analysis, ultra-fine-grained visual categorization (Ultra-FGVC) plays a vital role in distinguishing intricate subcategories within broader categories. However, this task is inherently challenging due to the complex granularity of category subdivisions and the limited availability of data for each category. To address these challenges, this work proposes CSDNet, a pioneering framework that effectively explores contrastive learning and self-distillation to learn discriminative representations specifically designed for Ultra-FGVC tasks. CSDNet comprises three main modules: Subcategory-Specific Discrepancy Parsing (SSDP), Dynamic Discrepancy Learning (DDL), and Subcategory-Specific Discrepancy Transfer (SSDT), which collectively enhance the generalization of deep models across instance, feature, and logit prediction levels. To increase the diversity of training samples, the SSDP module introduces adaptive augmented samples to spotlight subcategory-specific discrepancies. Simultaneously, the proposed DDL module stores historical intermediate features by a dynamic memory queue, which optimizes the feature learning space through iterative contrastive learning. Furthermore, the SSDT module effectively distills subcategory-specific discrepancies knowledge from the inherent structure of limited training data using a self-distillation paradigm at the logit prediction level. Experimental results demonstrate that CSDNet outperforms current state-of-the-art Ultra-FGVC methods, emphasizing its powerful efficacy and adaptability in addressing Ultra-FGVC tasks. Ziye Fang, Xin Jiang 0010, Hao Tang 0007, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Normal Image Guided Segmentation Framework for Unsupervised Anomaly DetectionabstractUnsupervised anomaly detection is required to detect/segment anomalous samples/regions that deviate from the normal pattern while learning only through the normal sample category. Towards this end, this paper proposes a novel framework for anomaly detection by introducing normal images as guidance called Normal Image Guided Segmentation Framework (NIGSF). It consists of a Normal Guided Network (NGN) and a Saliency Augmentation Module (SAM). NGN constructs the contrast set, which is a candidate set for extracting normal sample features. Then, a normal feature extractor is developed to extract detailed and complete features containing normal semantic information as guidance features. Meanwhile, the guidance feature fusion module is introduced to realize normal semantic guidance in the feature space, and then the segmentation module discriminates the features that are different from the normal guidance features as anomalies. SAM aims to generate forged anomaly samples utilizing available normal samples. It introduces saliency maps and random Perlin noise to generate saliency Perlin noise maps and then to generate diverse forged anomaly samples. Extensive experiments are conducted to evaluate the performance of NIGSF on three anomaly detection benchmark datasets. The results demonstrate the effectiveness of each proposed module and the superiority of the proposed method. Specifically, NIGSF outperforms the runner-up by 5.4% in terms of anomaly segmentation AP metric. Peng Xing, Yanpeng Sun, Dan Zeng 0001, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Spatial Structure Constraints for Weakly Supervised Semantic SegmentationabstractThe image-level label has prevailed in weakly supervised semantic segmentation tasks due to its easy availability. Since image-level labels can only indicate the existence or absence of specific categories of objects, visualization-based techniques have been widely adopted to provide object location clues. Considering class activation maps (CAMs) can only locate the most discriminative part of objects, recent approaches usually adopt an expansion strategy to enlarge the activation area for more integral object localization. However, without proper constraints, the expanded activation will easily intrude into the background region. In this paper, we propose spatial structure constraints (SSC) for weakly supervised semantic segmentation to alleviate the unwanted object over-activation of attention expansion. Specifically, we propose a CAM-driven reconstruction module to directly reconstruct the input image from deep CAM features, which constrains the diffusion of last-layer object attention by preserving the coarse spatial structure of the image content. Moreover, we propose an activation self-modulation module to refine CAMs with finer spatial structure details by enhancing regional consistency. Without external saliency models to provide background clues, our approach achieves 72.7% and 47.0% mIoU on the PASCAL VOC 2012 and COCO datasets, respectively, demonstrating the superiority of our proposed approach. The source codes and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/SSC. Tao Chen 0012, Yazhou Yao, Xingguo Huang, Zechao Li, Liqiang Nie, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Global Meets Local: Dual Activation Hashing Network for Large-Scale Fine-Grained Image RetrievalabstractIn the Internet era, the exponential growth of fine-grained image databases poses a considerable challenge for efficient information retrieval. Hashing-based approaches gained traction for their computational and storage efficiency, yet fine-grained hashing retrieval presents unique challenges due to small inter-class and large intra-class variations inherent to fine-grained entities. Thus, traditional hashing algorithms falter in discerning these subtle, yet critical, visual differences and fail to generate compact yet semantically rich hash codes. To address this, we introduce a Dual Activation Hashing Network (DAHNet) designed to convert high-dimensional image data into optimized binary codes via an innovative feature activation paradigm. The architecture consists of dual branches specifically tailored for global and local semantic activation, thereby establishing direct correspondences between hash codes and distinguishable object parts through a hierarchical activation pipeline. Specifically, our spatial-oriented semantic activation module modulates dominant visual regions while amplifying the activations of subtle yet semantically rich areas in a controlled manner. Building on these activated visual representations, the proposed inter-region semantic enrichment module further enriches them by unearthing semantically complementary cues. Concurrently,DAHNetintegrates a channel-oriented semantic activation module that exploits channel-specific correlations to distill contextual cues from spatially-activated visual features, thereby reinforcing robust learning to hash. To maintain the similarity of the original entities, we amalgamate final hash codes from both activation branches, capturing both local textural details and global structural information. Comprehensive evaluations on five fine-grained image retrieval benchmarks demonstrateDAHNet's superior performance over existing state-of-the-art hashing solutions, especially on 12-bit, improving performance by 4%-15% compared to the current best results on the five benchmarks. Moreover, generalization studies validate the efficacy of our dual-activation framework in the domain of content-based fine-grained image retrieval. The code is publicly available at:https://github.com/WhiteJiang/DAHNet. Xin Jiang 0010, Hao Tang 0007, Zechao Li |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Self-Paced Relational Contrastive Hashing for Large-Scale Image RetrievalabstractSupervised deep hashing aims to learn hash functions using label information. Existing methods learn hash functions by employing either pairwise/triplet loss to explore the point-to-point relation or center loss to explore the point-to-class relation. However, these methods overlook the collaboration between the above two kinds of relations and the hardness of pairs. In this work, we propose a novel Self-Paced Relational Contrastive Hashing (SPRCH) method with a single learning objective to capture valuable discriminative information from hard pairs using both the point-to-point and point-to-class relations. To exploit the above two kinds of relations, the Relational Contrastive Hash (RCH) loss is proposed, which ensures that each data anchor is closer to all similar data points and corresponding class centers in the Hamming space compared to dissimilar ones. Moreover, the proposed RCH loss reduces the drastic imbalance between point-to-point pairs and point-to-class pairs by rebalancing their weights. To prioritize hard pairs, a self-paced learning schedule is proposed, assigning higher weights to these pairs in the RCH loss. The self-paced learning schedule assigns dynamic weights to pairs according to their similarities and the training process. In this way, deep hash model can initially learn universal patterns from the entire set of pairs and then gradually acquire more valuable discriminative information from hard pairs. Experimental results on four widely-used image retrieval datasets demonstrate that our proposed SPRCH method significantly outperforms the state-of-the-art supervised deep hash methods. Zhengyun Lu, Lu Jin 0001, Zechao Li, Jinhui Tang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Unbiased Visual Question Answering by Leveraging Instrumental VariableabstractExisting unbiased VQA models reduce the spurious correlation between questions and answers to force the models to focus on visual information. However, the visual information captured by these unbiased models is irrelevant to the correct answer, resulting in leveraging spurious correlation to predict incorrect answers. This makes these unbiased methods fail to obtain critical visual information, thus performing poorly on questions dominated by the visual information. To capture the valuable visual information, this paper proposes a novel unbiased VQA model based on causal inference, leveraging Instrumental Variable (IVar) to increase the causal effect between visual features and answers. First, to obtain suitable instrumental variables, the noise generator is proposed according to the constraints of IVar. The generated noise can be regarded as IVar, which is used to pollute the original visual features. Then, this paper proposes IVar loss which utilizes the generated IVar to increase the causal effect between visual features and answers. When the visual feature is polluted by IVar, IVar loss guides the model to predict incorrect answers to enhance the correlation between IVar and the answer. Since the correlation between IVar and the answer is proportional to the causal effect between the visual feature and the answer, IVar loss enhances the importance of the visual information, thereby rectifying the model to capture critical visual information. The extensive experimental results on widely-used benchmarks demonstrate the advantages of the proposed method. The proposed method gains the best accuracy on answer typeOtherof VQA-CP v2. These results demonstrate the superiority of the proposed method in capturing critical visual information since most questions on the answer typeOtherare dominated by visual information. Yonghua Pan, Jing Liu 0001, Lu Jin 0001, Zechao Li |
IEEE Trans. Multim. | 4 |
| 2024 | Alleviating Over-Fitting in Hashing-Based Fine-Grained Image Retrieval: From Causal Feature Learning to Binary-Injected Hash LearningabstractHashing-based fine-grained image retrieval pursues learning diverse local features to generate inter-class discriminative hash codes. However, existing fine-grained hash methods with attention mechanisms usually tend to just focus on a few obvious areas, which misguides the network to over-fit some salient features. Such a problem raises two main limitations. 1) It overlooks some subtle local features, degrading the generalization capability of learned embedding. 2) It causes the over-activation of some hash bits correlated to salient features, which breaks the binary code balance and further weakens the discrimination abilities of hash codes. To address these limitations of the over-fitting problem, we propose a novel hash framework fromCausalFeature learning toBinary-injectedHash learning (CFBH), which captures various local information and suppresses over-activated hash bits simultaneously. For causal feature learning, we adopt causal inference theory to alleviate the bias towards the salient regions in fine-grained images. In detail, we obtain local features from the feature map and combine this local information with original image information followed by this theory. Theoretically, these fused embeddings help the network to re-weight the retrieval effort of each local feature and exploit more subtle variations without observational bias. For binary-injected hash learning, we propose a Binary Noise Injection (BNI) module inspired by Dropout. The BNI module not only mitigates over-activation to particular bits, but also makes hash codes uncorrelated and balanced in the Hamming space. Extensive experimental results on six popular fine-grained image datasets demonstrate the superiority of CFBH over several State-of-the-Art methods. Xinguang Xiang, Xinhao Ding, Lu Jin 0001, Zechao Li, Jinhui Tang 0001, Ramesh Jain 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Extraordinarily Time- and Memory-Efficient Large-Scale Canonical Correlation Analysis in Fourier Domain: From Shallow to DeepabstractCanonical correlation analysis (CCA) is a correlation analysis technique that is widely used in statistics and the machine-learning community. However, the high complexity involved in the training process lays a heavy burden on the processing units and memory system, making CCA nearly impractical in large-scale data. To overcome this issue, a novel CCA method that tries to carry out analysis on the dataset in the Fourier domain is developed in this article. Appling Fourier transform on the data, we can convert the traditional eigenvector computation of CCA into finding some predefined discriminative Fourier bases that can be learned with only element-wise dot product and sum operations, without complex time-consuming calculations. As the eigenvalues come from the sum of individual sample products, they can be estimated in parallel. Besides, thanks to the data characteristic of pattern repeatability, the eigenvalues can be well estimated with partial samples. Accordingly, a progressive estimate scheme is proposed, in which the eigenvalues are estimated through feeding data batch by batch until the eigenvalues sequence is stable in order. As a result, the proposed method shows its characteristics of extraordinarily fast and memory efficiencies. Furthermore, we extend this idea to the nonlinear kernel and deep models and obtained satisfactory accuracy and extremely fast training time consumption as expected. An extensive discussion on the fast Fourier transform (FFT)-CCA is made in terms of time and memory efficiencies. Experimental results on several large-scale correlation datasets, such as MNIST8M, X-RAY MICROBEAM SPEECH, and Twitter Users Data, demonstrate the superiority of the proposed algorithm over state-of-the-art (SOTA) large-scale CCA methods, as our proposed method achieves almost same accuracy with the training time of our proposed method being 1000 times faster. This makes our proposed models best practice models for dealing with large-scale correlation datasets. The source code is available at https://github.com/Mrxuzhao/FFTCCA. Xiangjun Shen, Zhaorui Xu, Liangjun Wang, Zechao Li, Guangcan Liu, Jianping Fan 0007, Zhengjun Zha |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | M3Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action RecognitionabstractDue to the scarcity of manually annotated data required for fine-grained video understanding, few-shot fine-grained (FS-FG) action recognition has gained significant attention, with the aim of classifying novel fine-grained action categories with only a few labeled instances. Despite the progress made in FS coarse-grained action recognition, current approaches encounter two challenges when dealing with the fine-grained action categories: the inability to capture subtle action details and the insufficiency of learning from limited data that exhibit high intra-class variance and inter-class similarity. To address these limitations, we propose M3Net, a matching-based framework for FS-FG action recognition, which incorporates multi-view encoding, multi-view matching, and multi-view fusion to facilitate embedding encoding, similarity matching, and decision making across multiple viewpoints.Multi-view encoding captures rich contextual details from the intra-frame, intra-video, and intra-episode perspectives, generating customized higher-order embeddings for fine-grained data.Multi-view matching integrates various matching functions enabling flexible relation modeling within limited samples to handle multi-scale spatio-temporal variations by leveraging the instance-specific, category-specific, and task-specific perspectives. Multi-view fusion consists of matching-predictions fusion and matching-losses fusion over the above views, where the former promotes mutual complementarity and the latter enhances embedding generalizability by employing multi-task collaborative learning. Explainable visualizations and experimental results on three challenging benchmarks demonstrate the superiority of M3Net in capturing fine-grained action details and achieving state-of-the-art performance for FS-FG action recognition. Hao Tang 0007, Jun Liu 0036, Shuanglin Yan, Rui Yan 0010, Zechao Li, Jinhui Tang 0001 |
ACM Multimedia | 5 |
| 2023 | Robust Spectral Embedding Completion Based Incomplete Multi-view ClusteringabstractGraph based methods have been widely used in incomplete multi-view clustering (IMVC). Most recent methods try to fill the original missing samples or incomplete affinity matrices to obtain a complete similarity graph for the subsequent spectral clustering. However, recovering the original high-dimensional data or complete n X n similarity matrix is usually time-consuming and noise-sensitive. Besides, they generally separate the cluster indicator learning into an individual step, which may result in sub-optimal graphs or spectral embeddings for clustering. To address these problems, this paper proposes a robust Spectral Embedding Completion based IMVC (SEC-IMVC) method, which incorporates spectral embedding completion and discrete cluster indicator learning into a unified framework. SEC-IMVC performs completion on spectral embeddings, and the embedding noise is eliminated to reduce the negative influence of original data noise. The discrete cluster indicator matrix is seamlessly learned by using spectral rotation, and it can explore the first-order feature consistency among different views. To further improve the completion robustness, the second-order correlation consistency is also captured by pairwise relations alignment. We compare our method with some state-of-the-art approaches on several datasets, and the experimental results show the effectiveness and advantages of our method. Chao Zhang 0078, Jingwen Wei, Bo Wang 0027, Zechao Li, Chunlin Chen 0001, Huaxiong Li |
ACM Multimedia | 4 |
| 2023 | Who is partner: A new perspective on data association of multi-object tracking
Yuqing Ding, Yanpeng Sun, Zechao Li |
Image Vis. Comput. | 3 |
| 2023 | Entity-Enhanced Adaptive Reconstruction Network for Weakly Supervised Referring Expression GroundingabstractWeakly supervised Referring Expression Grounding (REG) aims to ground a particular target in an image described by a language expression while lacking the correspondence between target and expression. Two main problems exist in weakly supervised REG. First, the lack of region-level annotations introduces ambiguities between proposals and queries. Second, most previous weakly supervised REG methods ignore the discriminative location and context of the referent, causing difficulties in distinguishing the target from other same-category objects. To address the above challenges, we design an entity-enhanced adaptive reconstruction network (EARN). Specifically, EARN includes three modules: entity enhancement, adaptive grounding, and collaborative reconstruction. In entity enhancement, we calculate semantic similarity as supervision to select the candidate proposals. Adaptive grounding calculates the ranking score of candidate proposals upon subject, location and context with hierarchical attention. Collaborative reconstruction measures the ranking result from three perspectives: adaptive reconstruction, language reconstruction and attribute classification. The adaptive mechanism helps to alleviate the variance of different referring expressions. Experiments on five datasets show EARN outperforms existing state-of-the-art methods. Qualitative results demonstrate that the proposed EARN can better handle the situation where multiple objects of a particular category are situated together. Xuejing Liu, Liang Li 0003, Shuhui Wang, Zhengjun Zha, Zechao Li, Qi Tian 0001, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Visual Anomaly Detection via Partition Memory Bank Module and Error EstimationabstractReconstruction method based on the memory module for visual anomaly detection attempts to narrow the reconstruction error for normal samples while enlarging it for anomalous samples. Unfortunately, the existing memory module is not fully applicable to the anomaly detection task, and the reconstruction error of the anomaly samples remains small. Towards this end, this work proposes a new unsupervised visual anomaly detection method to jointly learn effective normal features and eliminate unfavorable reconstruction errors. Specifically, a novel Partition Memory Bank (PMB) module is proposed to effectively learn and store detailed features with semantic integrity of normal samples. It develops a new partition mechanism and a unique query generation method to preserve the context information and then improves the learning ability of the memory module. The proposed PMB and the skip connection are alternatively explored to make the reconstruction of abnormal samples worse. To obtain more precise anomaly localization results and solve the problem of cumulative reconstruction error, a novel Histogram Error Estimation module is proposed to adaptively eliminate the unfavorable errors by the histogram of the difference image. It improves the anomaly localization performance without increasing the cost. To evaluate the effectiveness of the proposed method for anomaly detection and localization, extensive experiments are conducted on three widely-used anomaly detection datasets. The encouraging performance of the proposed method compared to the recent approaches based on the memory module demonstrates its superiority. Peng Xing, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Neulft: A Novel Approach to Nonlinear Canonical Polyadic Decomposition on High-Dimensional Incomplete TensorsabstractA High-Dimensional and Incomplete (HDI) tensor is frequently encountered in a big data-related application concerning the complex dynamic interactions among numerous entities. Traditional tensor factorization-based models cannot handle an HDI tensor efficiently, while existing latent factorization of tensors models are all linear models unable to model an HDI tensor's nonlinearity. Motivated by this critical discovery, this paper proposes a Neural Latent Factorization of Tensors model, which provides a novel approach to nonlinear Canonical Polyadic decomposition on an HDI tensor. It is implemented with three-fold interesting ideas: a) adopting the density-oriented modeling principle to build rank-one tensor series with high computational efficiency and affordable storage cost; b) treating each rank-one tensor as a hidden neuron to achieve an efficient neural network structure; and c) developing an adaptive backward propagation (ABP) learning scheme for efficient model training. Experimental results on six HDI tensors from a real system demonstrate that compared with state-of-the-art models, the proposed model achieves significant performance gain in both convergence rate and accuracy. Hence, it is of great significance in performing challenging HDI tensor analysis. Xin Luo 0001, Hao Wu 0061, Zechao Li |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | 3D3M: 3D Modulated Morphable Model for Monocular Face Reconstructionabstract3D face reconstruction from a single image is a vital task in various multimedia applications. A key challenge for 3D face shape reconstruction is to build the correct dense face correspondence between the monocular input face and the deformable mesh. Most existing methods rely on shape labels fitted by traditional methods or strong priors such as multi-view geometry consistency. In contrast, we propose an innovative 3D Modulated Morphable Model (3D3M) to learn the dense shape correspondence from monocular images in a self-supervised manner. Specifically, given a batch of input faces, 3D3M encodes their 3DMM attributes (shape, texture, lighting, etc.) and then randomly shuffles the 3DMM attributes to generate the attribute-changed faces. The attribute-changed faces can be encoded and rendered back in a cycle-consistent manner, which enables us to utilize the self-supervised consistencies in dense mesh vertices and reconstructed pixels. The dense shape and pixel correspondence enable us to adopt a series of self-supervised constraints to fit the 3D face model accurately and learn the per-vertex correctives end-to-end. 3D3M builds excellent high-quality 3D face reconstruction results from monocular images. Both quantitative and qualitative experimental results have verified the superiority of 3D3M over prior arts on 3D face reconstruction and face alignment. Yong Li 0032, Jianguo Hu, Xinmiao Pan, Zechao Li, Zhen Cui 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Tube-Embedded Transformer for Pixel PredictionabstractMulti-task pixel-level learning, which aims to exploit the inter-task interactions to improve the learning of each task, is an important but challenging issue in visual perception and multimedia applications. Measuring the inter-task correlation and intra-task specificity, we propose a tube-embedded transformer (TET) framework for robust multi-task pixel prediction. To facilitate inter-task interactions, we aggregate and project all tasks into a shared tube pool to generate the latent multi-task representation during the coarse-to-fine decoding stages. The resulting task-tube interactions replace the two-by-two task-task interactions to reduce the model complexity significantly. In addition, we introduce the transformer mechanism to adaptively transfer tube features to the target task. Concretely, on the one hand, multi-task features aggregate in the tube to generate the shared feature representation bases; on the other hand, based on the task-tube association and complementarity, the tube outputs the query entry and the weighting coefficients of the target task. Experimentally, on the joint learning of semantic segmentation, depth estimation, and surface normal estimation, the comparison experiments show the superiority of the TET multi-task learning method over other state-of-the-art approaches, and the ablation experiments verify the effectiveness of the TET mechanism. Zhen Cui 0001, Zechao Li, Jin Xie 0001, Jian Yang 0003 |
IEEE Trans. Multim. | 4 |
| 2023 | Deep Semantic Multimodal Hashing Network for Scalable Image-Text and Video-Text RetrievalsabstractHashing has been widely applied to multimodal retrieval on large-scale multimedia data due to its efficiency in computation and storage. In this article, we propose a novel deep semantic multimodal hashing network (DSMHN) for scalable image-text and video-text retrieval. The proposed deep hashing framework leverages 2-D convolutional neural networks (CNN) as the backbone network to capture the spatial information for image-text retrieval, while the 3-D CNN as the backbone network to capture the spatial and temporal information for video-text retrieval. In the DSMHN, two sets of modality-specific hash functions are jointly learned by explicitly preserving both intermodality similarities and intramodality semantic labels. Specifically, with the assumption that the learned hash codes should be optimal for the classification task, two stream networks are jointly trained to learn the hash functions by embedding the semantic labels on the resultant hash codes. Moreover, a unified deep multimodal hashing framework is proposed to learn compact and high-quality hash codes by exploiting the feature representation learning, intermodality similarity-preserving learning, semantic label-preserving learning, and hash function learning with different types of loss functions simultaneously. The proposed DSMHN method is a generic and scalable deep hashing framework for both image-text and video-text retrievals, which can be flexibly integrated with different types of loss functions. We conduct extensive experiments for both single-modal- and cross-modal-retrieval tasks on four widely used multimodal-retrieval data sets. Experimental results on both image-text- and video-text-retrieval tasks demonstrate that the DSMHN significantly outperforms the state-of-the-art methods. Lu Jin 0001, Zechao Li, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Probabilistic Regularized Extreme Learning for Robust Modeling of Traffic Flow ForecastingabstractThe adaptive neurofuzzy inference system (ANFIS) is a structured multioutput learning machine that has been successfully adopted in learning problems without noise or outliers. However, it does not work well for learning problems with noise or outliers. High-accuracy real-time forecasting of traffic flow is extremely difficult due to the effect of noise or outliers from complex traffic conditions. In this study, a novel probabilistic learning system, probabilistic regularized extreme learning machine combined with ANFIS (probabilistic R-ELANFIS), is proposed to capture the correlations among traffic flow data and, thereby, improve the accuracy of traffic flow forecasting. The new learning system adopts a fantastic objective function that minimizes both the mean and the variance of the model bias. The results from an experiment based on real-world traffic flow data showed that, compared with some kernel-based approaches, neural network approaches, and conventional ANFIS learning systems, the proposed probabilistic R-ELANFIS achieves competitive performance in terms of forecasting ability and generalizability. Jungang Lou, Yunliang Jiang, Qing Shen 0005, Ruiqin Wang, Zechao Li |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Time and Memory Efficient Large-Scale Canonical Correlation Analysis in Fourier DomainabstractCanonical correlation analysis (CCA) is a linear correlation analysis technique used widely in the statistics and machine learning community. However, the high complexity involved in pursuing eigenvector lays a heavy burden on the memory and computational time, making CCA nearly impractical in large-scale cases. In this paper, we attempt to overcome this issue by representing the data in the Fourier domain. Thanks to the data characteristic of pattern repeatability, one can translate projection-seeking of CCA into choosing some discriminative Fourier bases with only element-wise dot product and sum operations, without time-consuming eigenvector computation. Another merit of this scheme is that the eigenvalues can be approximated asymptotically in contrast to existing methods. Specifically, the eigenvalues can be estimated progressively, and the accuracy goes up as the number of data samples increases monotonously. This makes it possible to use partial data samples to obtain satisfactory accuracy. All the facts above make the proposed method extremely fast and memory efficient. Experimental results on several large-scale datasets, such as MNIST 8M, X-RAY MICROBEAM SPEECH, and TWITTER USERS Data, demonstrate the superiority of the proposed algorithm over SOTA large-scale CCA methods, as our proposed method achieves almost same accuracy with the training time being 1,000 times faster than SOTA methods. Xiangjun Shen, Zhaorui Xu, Liangjun Wang, Zechao Li |
ACM Multimedia | 4 |
| 2022 | Singular Value Fine-tuning: Few-shot Segmentation requires Few-parameters Fine-tuningabstractFreezing the pre-trained backbone has become a standard paradigm to avoid overfitting in few-shot segmentation. In this paper, we rethink the paradigm and explore a new regime: {\em fine-tuning a small part of parameters in the backbone}. We present a solution to overcome the overfitting problem, leading to better model generalization on learning novel classes. Our method decomposes backbone parameters into three successive matrices via the Singular Value Decomposition (SVD), then {\em only fine-tunes the singular values} and keeps others frozen. The above design allows the model to adjust feature representations on novel classes while maintaining semantic clues within the pre-trained backbone. We evaluate our {\em Singular Value Fine-tuning (SVF)} approach on various few-shot segmentation methods with different backbones. We achieve state-of-the-art results on both Pascal-5$^i$ and COCO-20$^i$ across 1-shot and 5-shot settings. Hopefully, this simple baseline will encourage researchers to rethink the role of backbone fine-tuning in few-shot settings. Yanpeng Sun, Qiang Chen 0007, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Jian Cheng 0001, Zechao Li, Jingdong Wang 0001 |
NeurIPS | 9 |
| 2022 | CTNet: Context-Based Tandem Network for Semantic SegmentationabstractContextual information has been shown to be powerful for semantic segmentation. This work proposes a novel Context-based Tandem Network (CTNet) by interactively exploring the spatial contextual information and the channel contextual information, which can discover the semantic context for semantic segmentation. Specifically, the Spatial Contextual Module (SCM) is leveraged to uncover the spatial contextual dependency between pixels by exploring the correlation between pixels and categories. Meanwhile, the Channel Contextual Module (CCM) is introduced to learn the semantic features including the semantic feature maps and class-specific features by modeling the long-term semantic dependence between channels. The learned semantic features are utilized as the prior knowledge to guide the learning of SCM, which can make SCM obtain more accurate long-range spatial dependency. Finally, to further improve the performance of the learned representations for semantic segmentation, the results of the two context modules are adaptively integrated to achieve better results. Extensive experiments are conducted on four widely-used datasets, i.e., PASCAL-Context, Cityscapes, ADE20K and PASCAL VOC2012. The results demonstrate the superior performance of the proposed CTNet by comparison with several state-of-the-art methods. The source code and models are available at https://github.com/syp2ysy/CTNet. Zechao Li, Yanpeng Sun, Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Learning attention-guided pyramidal features for few-shot fine-grained recognition
Hao Tang 0007, Chengcheng Yuan, Zechao Li, Jinhui Tang 0001 |
Pattern Recognit. | 3 |
| 2022 | MMatch: Semi-Supervised Discriminative Representation Learning for Multi-View ClassificationabstractSemi-supervised multi-view learning has been an important research topic due to its capability to exploit complementary information from unlabeled multi-view data. This work proposes MMatch, a new semi-supervised discriminative representation learning method for multi-view classification. Unlike existing multi-view representation learning methods that seldom consider the negative impact caused by particular views with unclear classification structures (weak discriminative views). MMatch jointly learns view-specific representations and class probabilities of training data. The representations concatenated to integrate multiple views’ information to form a global representation. Moreover, MMatch performs the smoothness constraint on the class probabilities of the global representation to improve pseudo labels, whereas the pseudo labels regularize the structure of view-specific representations. A discriminative global representation is mined with the training process, and the negative impact of weak discriminative views is overcome. Besides, MMatch learns consistent classification while preserving diverse information from multiple views. Experiments on several multi-view datasets demonstrate the effectiveness of MMatch. Xiaoli Wang 0003, Liyong Fu, Yudong Zhang 0001, Yongli Wang 0002, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Sub-Region Localized Hashing for Fine-Grained Image RetrievalabstractFine-grained image hashing is challenging due to the difficulties of capturing discriminative local information to generate hash codes. On the one hand, existing methods usually extract local features with the dense attention mechanism by focusing on dense local regions, which cannot contain diverse local information for fine-grained hashing. On the other hand, hash codes of the same class suffer from large intra-class variation of fine-grained images. To address the above problems, this work proposes a novel sub-Region Localized Hashing (sRLH) to learn intra-class compact and inter-class separable hash codes that also contain diverse subtle local information for efficient fine-grained image retrieval. Specifically, to localize diverse local regions, a sub-region localization module is developed to learn discriminative local features by locating the peaks of non-overlap sub-regions in the feature map. Different from localizing dense local regions, these peaks can guide the sub-region localization module to capture multifarious local discriminative information by paying close attention to dispersive local regions. To mitigate intra-class variations, hash codes of the same class are enforced to approach one common binary center. Meanwhile, the gram-schmidt orthogonalization is performed on the binary centers to make the hash codes inter-class separable. Extensive experimental results on four widely used fine-grained image retrieval datasets demonstrate the superiority of sRLH to several state-of-the-art methods. The source code of sRLH will be released at https://github.com/ZhangYajie-NJUST/sRLH.git. Xinguang Xiang, Lu Jin 0001, Zechao Li, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Learning Robust Discriminant Subspace Based on Joint L₂, ₚ- and L₂, ₛ-Norm Distance Metricsabstract-norm as the distance metric. However, both of their robustness and discriminant power are limited. In this article, we present a new robust discriminant subspace (RDS) learning method for feature extraction, with an objective function formulated in a different form. To guarantee the subspace to be robust and discriminative, we measure the within-class distances based on [Formula: see text]-norm and use [Formula: see text]-norm to measure the between-class distances. This also makes our method include rotational invariance. Since the proposed model involves both [Formula: see text]-norm maximization and [Formula: see text]-norm minimization, it is very challenging to solve. To address this problem, we present an efficient nongreedy iterative algorithm. Besides, motivated by trace ratio criterion, a mechanism of automatically balancing the contributions of different terms in our objective is found. RDS is very flexible, as it can be extended to other existing feature extraction techniques. An in-depth theoretical analysis of the algorithm's convergence is presented in this article. Experiments are conducted on several typical databases for image classification, and the promising results indicate the effectiveness of RDS. Liyong Fu, Zechao Li, Qiaolin Ye, Qingwang Liu, Xiaobo Chen 0001, Xijian Fan, Wankou Yang, Guowei Yang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Global-Guided Selective Context Network for Scene ParsingabstractRecent studies on semantic segmentation are exploiting contextual information to address the problem of inconsistent parsing prediction in big objects and ignorance in small objects. However, they utilize multilevel contextual information equally across pixels, overlooking those different pixels may demand different levels of context. Motivated by the above-mentioned intuition, we propose a novel global-guided selective context network (GSCNet) to adaptively select contextual information for improving scene parsing. Specifically, we introduce two global-guided modules, called global-guided global module (GGM) and global-guided local module (GLM), to, respectively, select global context (GC) and local context (LC) for pixels. When given an input feature map, GGM jointly employs the input feature map and its globally pooled feature to learn its global contextual demand based on which per-pixel GC is selected. While GLM adopts low-level feature from the adjacent stage as LC and synthetically models the input feature map, its globally pooled feature and LC to generate local contextual demand, based on which per-pixel LC is selected. Furthermore, we combine these two modules as a selective context block and import such SCBs in different levels of the network to propagate contextual information in a coarse-to-fine manner. Finally, we conduct extensive experiments to verify the effectiveness of our proposed model and achieve state-of-the-art performance on four challenging scene parsing data sets, i.e., Cityscapes, ADE20K, PASCAL Context, and COCO Stuff. Especially, GSCNet-101 obtains 82.6% on Cityscapes test set without using coarse data and 56.22% on ADE20K test set. Jie Jiang 0016, Jing Liu 0001, Jun Fu 0005, Zechao Li, Hanqing Lu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Causal Inference with Knowledge Distilling and Curriculum Learning for Unbiased VQAabstractRecently, many Visual Question Answering (VQA) models rely on the correlations between questions and answers yet neglect those between the visual information and the textual information. They would perform badly if the handled data distribute differently from the training data (i.e., out-of-distribution (OOD) data). Towards this end, we propose a two-stage unbiased VQA approach that addresses the unbiased issue from a causal perspective. In the causal inference stage, we mark the spurious correlation on the causal graph, explore the counterfactual causality, and devise a causal target based on the inherent correlations between the conventional and counterfactual VQA models. In the distillation stage, we introduce the causal target into the training process and leverages distilling as well as curriculum learning to capture the unbiased model. Since Causal Inference with Knowledge Distilling and Curriculum Learning (CKCL) reinforces the contribution of the visual information and eliminates the impact of the spurious correlation by distilling the knowledge in causal inference to the VQA model, it contributes to the good performance on both the standard data and out-of-distribution data. The extensive experimental results on VQA-CP v2 dataset demonstrate the superior performance of the proposed method compared to the state-of-the-art (SotA) methods. Yonghua Pan, Zechao Li, Liyan Zhang 0001, Jinhui Tang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Extracting Useful Knowledge from Noisy Web Images via Data Purification for Fine-Grained RecognitionabstractFine-grained visual recognition tasks typically require training data with reliable acquisition and annotation processes. Acquiring such datasets with precise fine-grained annotations is very expensive and time-consuming. Conversely, a vast amount of web data is relatively easy to obtain with nearly no human effort. Nevertheless, the presence of label noise in web images becomes a huge obstacle for training robust fine-grained recognition models. In this work, we investigate the noisy label problem and propose a method that can specifically distinguish in- and out-of-distribution noisy samples. It can purify the web training data by discarding out-of-distribution noisy images and relabeling in-distribution ones. After purification, we can train the model on a less noisy web training set to achieve better robustness and performance. Extensive experiments on three real-world web datasets for fine-grained visual recognition demonstrate the superiority of our approach. Chuanyi Zhang, Yazhou Yao, Xing Xu 0001, Jie Shao 0001, Jingkuan Song, Zechao Li, Zhenmin Tang |
ACM Multimedia | 6 |
| 2021 | Semi-supervised local feature selection for data classification
Zechao Li, Jinhui Tang 0001 |
Sci. China Inf. Sci. | 1 |
| 2021 | Label Distribution Learning with Label Correlations on Local SamplesabstractLabel distribution learning (LDL) is proposed for solving the label ambiguity problem in recent years, which can be seen as an extension of multi-label learning. To improve the performance of label distribution learning, some existing algorithms exploit label correlations in a global manner that assumes the label correlations are shared by all instances. However, the instances in different groups may share different label correlations, and few label correlations are globally applicable in real-world tasks. In this paper, two novel label distribution learning algorithms are proposed by exploiting label correlations on local samples, which are called GD-LDL-SCL and Adam-LDL-SCL, respectively. To utilize the label correlations on local samples, the influence of local samples is encoded, and a local correlation vector is designed as the additional features for each instance, which is based on the different clustered local samples. Then, the label distribution for an unseen instance can be predicted by exploiting the original features and the additional features simultaneously. Extensive experiments on some real-world data sets validate that our proposed methods can address the label distribution problems effectively and outperform state-of-the-art methods. Xiuyi Jia, Zechao Li, Weiwei Li 0001, Sheng-Jun Huang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Face Super-Resolution Guided by 3D Facial Priors
Xiaobin Hu, Wenqi Ren, John LaMaster, Xiaochun Cao, Xiaoming Li 0002, Zechao Li, Bjoern Menze, Wei Liu 0005 |
ECCV (4) | 6 |
| 2020 | How to Learn Item Representation for Cold-Start Multimedia Recommendation?abstractThe ability of recommending cold items (that have no behavior history) is a core strength of multimedia recommendation compared with behavior-only collaborative filtering. To learn effective item representation, a key challenge lies in the discrepancy between training and testing, since the cold items only exist in the testing data. This means that the signal used to represent an item varies during training and testing --- in the training stage, we can represent an item with both collaborative embedding and content embedding; whereas in the testing stage, we represent a cold item with content embedding only. Nevertheless, existing learning frameworks omit this critical discrepancy, resulting in suboptimal item representation for multimedia recommendation. Xiaoyu Du 0002, Xiang Wang 0010, Xiangnan He 0001, Zechao Li, Jinhui Tang 0001, Tat-Seng Chua |
ACM Multimedia | 4 |
| 2020 | Weakly-Supervised Image Hashing through Masked Visual-Semantic Graph-based ReasoningabstractWith the popularization of social websites, many methods have been proposed to explore the noisy tags for weakly-supervised image hashing.The main challenge lies in learning appropriate and sufficient information from those noisy tags. To address this issue, this work proposes a novel Masked visual-semantic Graph-based Reasoning Network, termed as MGRN, to learn joint visual-semantic representations for image hashing. Specifically, for each image, MGRN constructs a relation graph to capture the interactions among its associated tags and performs reasoning with Graph Attention Networks (GAT). MGRN randomly masks out one tag and then make GAT to predict this masked tag. This forces the GAT model to capture the dependence between the image and its associated tags, which can well address the problem of noisy tags. Thus it can capture key tags and visual structures from images to learn well-aligned visual-semantic representations. Finally, the auto-encoders is leveraged to learn hash codes that can preserve the local structure of the joint space. Meanwhile, the joint visual-semantic representations are reconstructed from those hash codes by using a decoder. Experimental results on two widely-used benchmark datasets demonstrate the superiority of the proposed method for image retrieval compared with several state-of-the-art methods. Lu Jin 0001, Zechao Li, Yonghua Pan, Jinhui Tang 0001 |
ACM Multimedia | 2 |
| 2020 | NuI-Go: Recursive Non-Local Encoder-Decoder Network for Retinal Image Non-Uniform Illumination RemovalabstractRetinal images have been widely used by clinicians for early diagnosis of ocular diseases. However, the quality of retinal images is often clinically unsatisfactory due to eye lesions and imperfect imaging process. One of the most challenging quality degradation issues in retinal images is non-uniform which hinders the pathological information and further impairs the diagnosis of ophthalmologists and computer-aided analysis. To address this issue, we propose a non-uniform illumination removal network for retinal image, called NuI-Go, which consists of three Recursive Non-local Encoder-Decoder Residual Blocks (NEDRBs) for enhancing the degraded retinal images in a progressive manner. Each NEDRB contains a feature encoder module that captures the hierarchical feature representations, a non-local context module that models the context information, and a feature decoder module that recovers the details and spatial dimension. Additionally, the symmetric skip-connections between the encoder module and the decoder module provide long-range information compensation and reuse. Extensive experiments demonstrate that the proposed method can effectively remove the non-uniform illumination on retinal images while well preserving the image details and color. We further demonstrate the advantages of the proposed method for improving the accuracy of retinal vessel segmentation. Chongyi Li, Huazhu Fu, Runmin Cong, Zechao Li, Qianqian Xu 0001 |
ACM Multimedia | 4 |
| 2020 | BlockMix: Meta Regularization and Self-Calibrated Inference for Metric-Based Meta-LearningabstractMost metric-based meta-learning methods learn only the sophisticated similarity metric for few-shot classification, which may lead to the feature deterioration and unreliable prediction. Toward this end, we propose new mechanisms to learn generalized and discriminative feature embeddings as well as improve the robustness of classifiers against prediction corruptions for meta-learning. For this purpose, a new generation operator BlockMix is proposed by integrating interpolation on the images and labels within metric learning. Based on the above BlockMix, we propose a novel regularization method Meta Regularization as an auxiliary task branch with its own classifier to better constraint the feature embedding module and stabilize the meta-learning process. Furthermore, a novel inference scheme Self-Calibrated Inference is proposed to alleviate the unreliable prediction problem by calibrating the prototype of each category with the confidence-weighted average of the support and generated samples. The proposed mechanisms can be used as supplementary techniques alongside standard metric-based meta-learning algorithms without any pre-training. Experimental results demonstrate the insights and the efficiency of the proposed mechanisms respectively, compared with the state-of-the-art methods on the prevalent few-shot benchmarks. Hao Tang 0007, Zechao Li, Zhimao Peng, Jinhui Tang 0001 |
ACM Multimedia | 2 |
| 2020 | Data-driven Meta-set Based Fine-Grained Visual RecognitionabstractConstructing fine-grained image datasets typically requires domain-specific expert knowledge, which is not always available for crowd-sourcing platform annotators. Accordingly, learning directly from web images becomes an alternative method for fine-grained visual recognition. However, label noise in the web training set can severely degrade the model performance. To this end, we propose a data-driven meta-set based approach to deal with noisy web images for fine-grained recognition. Specifically, guided by a small amount of clean meta-set, we train a selection net in a meta-learning manner to distinguish in- and out-of-distribution noisy images. To further boost the robustness of the model, we also learn a labeling net to correct the labels of in-distribution noisy data. In this way, our proposed method can alleviate the harmful effects caused by out-of-distribution noise and properly exploit the in-distribution noisy samples for training. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to state-of-the-art noise-robust methods. Chuanyi Zhang, Yazhou Yao, Xiangbo Shu, Zechao Li, Zhenmin Tang, Qi Wu 0001 |
ACM Multimedia | 4 |
| 2020 | Distilling knowledge in causal inference for unbiased visual question answeringabstractCurrent Visual Question Answering (VQA) models mainly explore the statistical correlations between answers and questions, which fail to capture the relationship between the visual information and answers. The performance dramatically decreases when the distribution of handled data is different from the training data. Towards this end, this paper proposes a novel unbiased VQA model by exploring the Casual Inference with Knowledge Distillation (CIKD) to reduce the influence of bias. Specifically, the causal graph is first constructed to explore the counterfactual causality and infer the casual target based on the causal effect, which well reduces the bias from questions and obtain answers without training. Then knowledge distillation is leveraged to transfer the knowledge of the inferred casual target to the conventional VQA model. It makes the proposed method enable to handle both the biased data and standard data. To address the problem of the bad bias from the knowledge distillation, the ensemble learning is introduced based on the hypothetical bias reason. Experiments are conducted to show the performance of the proposed method. The significant improvements over the state-of-the-art methods on the VQA-CP v2 dataset well validate the contributions of this work. Yonghua Pan, Zechao Li, Liyan Zhang 0001, Jinhui Tang 0001 |
MMAsia | 2 |
| 2020 | Weakly-supervised Semantic Guided Hashing for Social Image Retrieval
Zechao Li, Jinhui Tang 0001, Liyan Zhang 0001, Jian Yang 0003 |
Int. J. Comput. Vis. | 1 |
| 2020 | Discriminative supplementary representation learning for novel-category classification
Qiuli Liu, Zechao Li, Jinhui Tang 0001 |
Neurocomputing | 2 |
| 2020 | Task-Oriented Network for Image DehazingabstractHaze interferes the transmission of scene radiation and significantly degrades color and details of outdoor images. Existing deep neural networks-based image dehazing algorithms usually use some common networks. The network design does not model the image formation of haze process well, which accordingly leads to dehazed images containing artifacts and haze residuals in some special scenes. In this paper, we propose a task-oriented network for image dehazing, where the network design is motivated by the image formation of haze process. The task-oriented network involves a hybrid network containing an encoder and decoder network and a spatially variant recurrent neural network which is derived from the hazy process. In addition, we develop a multi-stage dehazing algorithm to further improve the accuracy by filtering haze residuals in a step-bystep fashion. To constrain the proposed network, we develop a dual composition loss, content-based pixel-wise loss and total variation constraint. We train the proposed network in an end-to-end manner and analyze its effect on image dehazing. Experimental results demonstrate that the proposed algorithm achieves favorable performance against state-of-the-art dehazing methods. Runde Li, Jinshan Pan, Zechao Li, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | RGBD Based Gaze Estimation via Multi-Task CNNabstractThis paper tackles RGBD based gaze estimation with Convolutional Neural Networks (CNNs). Specifically, we propose to decompose gaze point estimation into eyeball pose, head pose, and 3D eye position estimation. Compared with RGB image-based gaze tracking, having depth modality helps to facilitate head pose estimation and 3D eye position estimation. The captured depth image, however, usually contains noise and black holes which noticeably hamper gaze tracking. Thus we propose a CNN-based multi-task learning framework to simultaneously refine depth images and predict gaze points. We utilize a generator network for depth image generation with a Generative Neural Network (GAN), where the generator network is partially shared by both the gaze tracking network and GAN-based depth synthesizing. By optimizing the whole network simultaneously, depth image synthesis improves gaze point estimation and vice versa. Since the only existing RGBD dataset (EYEDIAP) is too small, we build a large-scale RGBD gaze tracking dataset for performance evaluation. As far as we know, it is the largest RGBD gaze dataset in terms of the number of participants. Comprehensive experiments demonstrate that our method outperforms existing methods by a large margin on both our dataset and the EYEDIAP dataset. Dongze Lian, Weixin Luo, Lina Hu, Minye Wu, Zechao Li, Jingyi Yu 0001, Shenghua Gao |
AAAI | 6 |
| 2019 | Facial Emotion Distribution Learning by Exploiting Low-Rank Label Correlations LocallyabstractEmotion recognition from facial expressions is an interesting and challenging problem and has attracted much attention in recent years. Substantial previous research has only been able to address the ambiguity of “what describes the expression”, which assumes that each facial expression is associated with one or more predefined affective labels while ignoring the fact that multiple emotions always have different intensities in a single picture. Therefore, to depict facial expressions more accurately, this paper adopts a label distribution learning approach for emotion recognition that can address the ambiguity of “how to describe the expression” and proposes an emotion distribution learning method that exploits label correlations locally. Moreover, a local low-rank structure is employed to capture the local label correlations implicitly. Experiments on benchmark facial expression datasets demonstrate that our method can better address the emotion distribution recognition problem than state-of-the-art methods. Xiuyi Jia, Weiwei Li 0001, Changqing Zhang 0002, Zechao Li |
CVPR | 5 |
| 2019 | Few-Shot Image Recognition With Knowledge TransferabstractHuman can well recognize images of novel categories just after browsing few examples of these categories. One possible reason is that they have some external discriminative visual information about these categories from their prior knowledge. Inspired from this, we propose a novel Knowledge Transfer Network architecture (KTN) for few-shot image recognition. The proposed KTN model jointly incorporates visual feature learning, knowledge inferring and classifier learning into one unified framework for their optimal compatibility. First, the visual classifiers for novel categories are learned based on the convolutional neural network with the cosine similarity optimization. To fully explore the prior knowledge, a semantic-visual mapping network is then developed to conduct knowledge inference, which enables to infer the classifiers for novel categories from base categories. Finally, we design an adaptive fusion scheme to infer the desired classifiers by effectively integrating the above knowledge and visual information. Extensive experiments are conducted on two widely-used Mini-ImageNet and ImageNet Few-Shot benchmarks to evaluate the effectiveness of the proposed method. The results compared with the state-of-the-art approaches show the encouraging performance of the proposed method, especially on 1-shot and 2-shot tasks. Zhimao Peng, Zechao Li, Junge Zhang, Guo-Jun Qi, Jinhui Tang 0001 |
ICCV | 2 |
| 2019 | Label distribution learning with label-specific featuresabstractLabel distribution learning (LDL) is a novel machine learning paradigm to deal with label ambiguity issues by placing more emphasis on how relevant each label is to a particular instance. Many LDL algorithms have been proposed and most of them concentrate on the learning models, while few of them focus on the feature selection problem. All existing LDL models are built on a simple feature space in which all features are shared by all the class labels. However, this kind of traditional data representation strategy tends to select features that are distinguishable for all labels, but ignores label-specific features that are pertinent and discriminative for each class label. In this paper, we propose a novel LDL algorithm by leveraging label-specific features. The common features for all labels and specific features for each label are simultaneously learned to enhance the LDL model. Moreover, we also exploit the label correlations in the proposed LDL model. The experimental results on several real-world data sets validate the effectiveness of our method. Tingting Ren, Xiuyi Jia, Weiwei Li 0001, Zechao Li |
IJCAI | 5 |
| 2019 | Attention-Aware Feature Pyramid Ordinal Hashing for Image RetrievalabstractDue to the effectiveness of representation learning, deep hashing methods have attracted increasing attention in image retrieval. However, most existing deep hashing methods merely encode the raw information of the last layer for hash learning, which result in the following deficiencies: (1) the useful information from the preceding-layer is not fully exploited; (2) the local salient information of the image is neglected. To this end, we propose a novel deep hashing method, called Attention-Aware Feature Pyramid Ordinal Hashing (AFPH), which explores both the visual structure information and semantic information from different convolutional layers. Specifically, two feature pyramids based on spatial and channel attention are well constructed to capture the local salient structure from multiple scales. Moreover, a multi-scale feature fusion strategy is proposed to aggregate the feature maps from multi-level pyramidal layers to generate the discriminative feature for ranking-based hashing. The experimental results conducted on two widely-used image retrieval datasets demonstrate the superiority of our method. Xie Sun, Lu Jin 0001, Zechao Li |
MMAsia | 3 |
| 2019 | Multimedia retrieval by deep hashing with multilevel similarity learning
Qiuli Liu, Lu Jin 0001, Zechao Li, Jinhui Tang 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Deep networks with non-static activation function
Huajun Zhou, Zechao Li |
Multim. Tools Appl. | 2 |
| 2019 | Deep Collaborative Embedding for Social Image UnderstandingabstractIn this work, we investigate the problem of learning knowledge from the massive community-contributed images with rich weakly-supervised context information, which can benefit multiple image understanding tasks simultaneously, such as social image tag refinement and assignment, content-based image retrieval, tag-based image retrieval and tag expansion. Towards this end, we propose a Deep Collaborative Embedding (DCE) model to uncover a unified latent space for images and tags. The proposed method incorporates the end-to-end learning and collaborative factor analysis in one unified framework for the optimal compatibility of representation learning and latent space discovery. A nonnegative and discrete refined tagging matrix is learned to guide the end-to-end learning. To collaboratively explore the rich context information of social images, the proposed method integrates the weakly-supervised image-tag correlation, image correlation and tag correlation simultaneously and seamlessly. The proposed model is also extended to embed new tags in the uncovered space. To verify the effectiveness of the proposed method, extensive experiments are conducted on two widely-used social image benchmarks for multiple social image understanding tasks. The encouraging performance of the proposed method over the state-of-the-art approaches demonstrates its superiority. Zechao Li, Jinhui Tang 0001, Tao Mei 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Social Anchor-Unit Graph Regularized Tensor Completion for Large-Scale Image RetaggingabstractImage retagging aims to improve the tag quality of social images by completing the missing tags, rectifying the noise-corrupted tags, and assigning new high-quality tags. Recent approaches simultaneously explore visual, user and tag information to improve the performance of image retagging by mining the tag-image-user associations. However, such methods will become computationally infeasible with the rapidly increasing number of images, tags and users. It has been proven that the anchor graph can significantly accelerate large-scale graph-based learning by exploring only a small number of anchor points. Inspired by this, we propose a novel Social anchor-Unit GrAph Regularized Tensor Completion (SUGAR-TC) method to efficiently refine the tags of social images, which is insensitive to the scale of data. First, we construct an anchor-unit graph across multiple domains (e.g., image and user domains) rather than traditional anchor graph in a single domain. Second, a tensor completion based on Social anchor-Unit GrAph Regularization (SUGAR) is implemented to refine the tags of the anchor images. Finally, we efficiently assign tags to non-anchor images by leveraging the relationship between the non-anchor units and the anchor units. Experimental results on a real-world social image database well demonstrate the effectiveness and efficiency of SUGAR-TC, outperforming the state-of-the-art methods. Jinhui Tang 0001, Xiangbo Shu, Zechao Li, Yu-Gang Jiang 0001, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Deep Ordinal Hashing With Spatial AttentionabstractHashing has attracted increasing research attention in recent years due to its high efficiency of computation and storage in image retrieval. Recent works have demonstrated the superiority of simultaneous feature representations and hash functions learning with deep neural networks. However, most existing deep hashing methods directly learn the hash functions by encoding the global semantic information, while ignoring the local spatial information of images. The loss of local spatial structure makes the performance bottleneck of hash functions, therefore limiting its application for accurate similarity retrieval. In this paper, we propose a novel deep ordinal hashing (DOH) method, which learns ordinal representations to generate ranking-based hash codes by leveraging the ranking structure of feature space from both local and global views. In particular, to effectively build the ranking structure, we propose to learn the rank correlation space by exploiting the local spatial information from fully convolutional network and the global semantic information from the convolutional neural network simultaneously. More specifically, an effective spatial attention model is designed to capture the local spatial information by selectively learning well-specified locations closely related to target objects. In such hashing framework, the local spatial and global semantic nature of images is captured in an end-to-end ranking-to-hashing manner. Experimental results conducted on three widely used datasets demonstrate that the proposed DOH method significantly outperforms the state-of-the-art hashing methods. Lu Jin 0001, Xiangbo Shu, Kai Li 0005, Zechao Li, Guo-Jun Qi, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Deep Semantic-Preserving Ordinal Hashing for Cross-Modal Similarity SearchabstractCross-modal hashing has attracted increasing research attention due to its efficiency for large-scale multimedia retrieval. With simultaneous feature representation and hash function learning, deep cross-modal hashing (DCMH) methods have shown superior performance. However, most existing methods on DCMH adopt binary quantization functions (e.g., [Formula: see text]) to generate hash codes, which limit the retrieval performance since binary quantization functions are sensitive to the variations of numeric values. Toward this end, we propose a novel end-to-end ranking-based hashing framework, in this paper, termed as deep semantic-preserving ordinal hashing (DSPOH), to learn hash functions with deep neural networks by exploring the ranking structure of feature dimensions. In DSPOH, the ordinal representation, which encodes the relative rank ordering of feature dimensions, is explored to generate hash codes. Such ordinal embedding benefits from the numeric stability of rank correlation measures. To make the hash codes discriminative, the ordinal representation is expected to well predict the class labels so that the ranking-based hash function learning is optimally compatible with the label predicting. Meanwhile, the intermodality similarity is preserved to guarantee that the hash codes of different modalities are consistent. Importantly, DSPOH can be effectively integrated with different types of network architectures, which demonstrates the flexibility and scalability of our proposed hashing framework. Extensive experiments on three widely used multimodal data sets show that DSPOH outperforms state of the art for cross-modal retrieval tasks. Lu Jin 0001, Kai Li 0005, Zechao Li, Fu Xiao 0001, Guo-Jun Qi, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Nonpeaked Discriminant Analysis for Data RepresentationabstractOf late, there are many studies on the robust discriminant analysis, which adopt L1-norm as the distance metric, but their results are not robust enough to gain universal acceptance. To overcome this problem, the authors of this article present a nonpeaked discriminant analysis (NPDA) technique, in which cutting L1-norm is adopted as the distance metric. As this kind of norm can better eliminate heavy outliers in learning models, the proposed algorithm is expected to be stronger in performing feature extraction tasks for data representation than the existing robust discriminant analysis techniques, which are based on the L1-norm distance metric. The authors also present a comprehensive analysis to show that cutting L1-norm distance can be computed equally well, using the difference between two special convex functions. Against this background, an efficient iterative algorithm is designed for the optimization of the proposed objective. Theoretical proofs on the convergence of the algorithm are also presented. Theoretical insights and effectiveness of the proposed method are validated by experimental tests on several real data sets. Qiaolin Ye, Zechao Li, Liyong Fu, Zhao Zhang 0001, Wankou Yang, Guowei Yang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Show, Reward, and Tell: Adversarial Visual Story GenerationabstractDespite the promising progress made in visual captioning and paragraphing, visual storytelling is still largely unexplored. This task is more challenging due to the difficulty in modeling an ordered photo sequence and in generating a relevant paragraph with expressive language style for storytelling. To deal with these challenges, we propose an Attribute-based Hierarchical Generative model with Reinforcement Learning and adversarial training (AHGRL). First, to model the ordered photo sequence and the complex story structure, we propose an attribute-based hierarchical generator. The generator incorporates semantic attributes to create more accurate and relevant descriptions. The hierarchical framework enables the generator to learn from the complex paragraph structure. Second, to generate story-style paragraphs, we design a language-style discriminator, which provides word-level rewards to optimize the generator by policy gradient. Third, we further consider the story generator and the reward critic as adversaries. The generator aims to create indistinguishable paragraphs to human-level stories, whereas the critic aims at distinguishing them and further improving the generator. Extensive experiments on the widely used dataset well demonstrate the advantages of the proposed method over state-of-the-art methods. Jinhui Tang 0001, Jing Wang 0221, Zechao Li, Jianlong Fu, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2018 | Show, Reward and Tell: Automatic Generation of Narrative Paragraph From Photo Stream by Adversarial TrainingabstractImpressive image captioning results (i.e., an objective description for an image) are achieved with plenty of training pairs. In this paper, we take one step further to investigate the creation of narrative paragraph for a photo stream. This task is even more challenging due to the difficulty in modeling an ordered photo sequence and in generating a relevant paragraph with expressive language style for storytelling. The difficulty can even be exacerbated by the limited training data, so that existing approaches almost focus on search-based solutions. To deal with these challenges, we propose a sequence-to-sequence modeling approach with reinforcement learning and adversarial training. First, to model the ordered photo stream, we propose a hierarchical recurrent neural network as story generator, which is optimized by reinforcement learning with rewards. Second, to generate relevant and story-style paragraphs, we design the rewards with two critic networks, including a multi-modal and a language-style discriminator. Third, we further consider the story generator and reward critics as adversaries. The generator aims to create indistinguishable paragraphs to human-level stories, whereas the critics aim at distinguishing them and further improving the generator by policy gradient. Experiments on three widely-used datasets show the effectiveness, against state-of-the-art methods with relative increase of 20.2% by METEOR. We also show the subjective preference for the proposed approach over the baselines through a user study with 30 human subjects. Jing Wang 0221, Jianlong Fu, Jinhui Tang 0001, Zechao Li, Tao Mei 0001 |
AAAI | 4 |
| 2018 | Single Image Dehazing via Conditional Generative Adversarial NetworkabstractIn this paper, we present an algorithm to directly restore a clear image from a hazy image. This problem is highly ill-posed and most existing algorithms often use hand-crafted features, e.g., dark channel, color disparity, maximum contrast, to estimate transmission maps and then atmospheric lights. In contrast, we solve this problem based on a conditional generative adversarial network (cGAN), where the clear image is estimated by an end-to-end trainable neural network. Different from the generative network in basic cGAN, we propose an encoder and decoder architecture so that it can generate better results. To generate realistic clear images, we further modify the basic cGAN formulation by introducing the VGG features and an L1-regularized gradient prior. We also synthesize a hazy dataset including indoor and outdoor scenes to train and evaluate the proposed algorithm. Extensive experimental results demonstrate that the proposed method performs favorably against the state-of-the-art methods on both synthetic dataset and real world hazy images. Runde Li, Jinshan Pan, Zechao Li, Jinhui Tang 0001 |
CVPR | 3 |
| 2018 | Learning Dual Convolutional Neural Networks for Low-Level VisionabstractIn this paper, we propose a general dual convolutional neural network (DualCNN) for low-level vision problems, e.g., super-resolution, edge-preserving filtering, deraining and dehazing. These problems usually involve the estimation of two components of the target signals: structures and details. Motivated by this, our proposed DualCNN consists of two parallel branches, which respectively recovers the structures and details in an end-to-end manner. The recovered structures and details can generate the target signals according to the formation model for each particular application. The DualCNN is a flexible framework for low-level vision tasks and can be easily incorporated into existing CNNs. Experimental results show that the DualCNN can be effectively applied to numerous low-level vision tasks with favorable performance against the state-of-the-art methods. Jinshan Pan, Sifei Liu, Deqing Sun, Jiawei Zhang 0002, Yang Liu 0119, Jimmy S. J. Ren, Zechao Li, Jinhui Tang 0001, Huchuan Lu, Yu-Wing Tai, Ming-Hsuan Yang 0001 |
CVPR | 7 |
| 2018 | Participation-Contributed Temporal Dynamic Model for Group Activity RecognitionabstractGroup activity recognition, a challenging task that a number of individuals occur in the scene of activity while only a small subset of them participate in, has received increasing attentions. However, most of the previous methods model all the individuals' actions equivalently while ignoring a fact that not all of them are contributed to the discrimination of group activity. That is to say, only a small number of key actors (participants) play important roles in the whole group activity. Inspired by this, we explore a new "One to Key" idea to progressively aggregate temporal dynamics of key actors with different participation degrees over time from each person. Here, we focus on two types of key actors in the whole activity, who steadily move in the whole process (long moving time) or intensely move (but closely related to the group activity) at a significant moment. Based on this, we propose a novel Participation-Contributed Temporal Dynamic Model (PC-TDM) to recognize group activity, which mainly consists of a "One" network and a "One to Key" network. Specifically, "One" network aims at modeling the individual dynamic of each person. "One to Key" network feeds the outputs from the "One" network into a Bidirectional LSTM (Bi-LSTM) according to the order of individual's moving time. Subsequently, each output state of Bi-LSTM weighted by a trainable time-varying attention factor is aggregated by going through LSTM one-by-one. Experimental results on two benchmarks demonstrate that the proposed method improves group activity recognition performance compared to the state-of-the-arts. Rui Yan 0010, Jinhui Tang 0001, Xiangbo Shu, Zechao Li, Qi Tian 0001 |
ACM Multimedia | 4 |
| 2018 | Matrix Entropy Driven Maximum Margin Feature Learning
Jinhui Tang 0001, Zechao Li |
PRICAI (1) | 3 |
| 2018 | Visual understanding by mining social media: recent advances and challenges
Xueming Wang, Zechao Li, Jinhui Tang 0001 |
Frontiers Comput. Sci. | 2 |
| 2018 | Tracking the evolution of overlapping communities in dynamic social networks
Zechao Li, Guan Yuan, Yunlian Sun, Xiaobin Rui, Xinguang Xiang |
Knowl. Based Syst. | 2 |
| 2018 | Video summarization via exploring the global and local importance
Tongling Hu, Zechao Li |
Multim. Tools Appl. | 2 |
| 2018 | Personalized Age Progression with Bi-Level Aging Dictionary LearningabstractAge progression is defined as aesthetically re-rendering the aging face at any future age for an individual face. In this work, we aim to automatically render aging faces in a personalized way. Basically, for each age group, we learn an aging dictionary to reveal its aging characteristics (e.g., wrinkles), where the dictionary bases corresponding to the same index yet from two neighboring aging dictionaries form a particular aging pattern cross these two age groups, and a linear combination of all these patterns expresses a particular personalized aging process. Moreover, two factors are taken into consideration in the dictionary learning process. First, beyond the aging dictionaries, each person may have extra personalized facial characteristics, e.g., mole, which are invariant in the aging process. Second, it is challenging or even impossible to collect faces of all age groups for a particular person, yet much easier and more practical to get face pairs from neighboring age groups. To this end, we propose a novel Bi-level Dictionary Learning based Personalized Age Progression (BDL-PAP) method. Here, bi-level dictionary learning is formulated to learn the aging dictionaries based on face pairs from neighboring age groups. Extensive experiments well demonstrate the advantages of the proposed BDL-PAP over other state-of-the-arts in term of personalized age progression, as well as the performance gain for cross-age face verification by synthesizing aging faces. Xiangbo Shu, Jinhui Tang 0001, Zechao Li, Hanjiang Lai, Liyan Zhang 0001, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Supervised deep hashing for scalable face image retrieval
Jinhui Tang 0001, Zechao Li |
Pattern Recognit. | 2 |
| 2018 | Image Classification With Tailored Fine-Grained DictionariesabstractIn this paper, we propose a novel fine-grained dictionary learning method for image classification. To learn a high-quality discriminative dictionary, three types of multispecific subdictionaries, i.e., class-specific dictionaries (CSDs), universal dictionary (UD), and family-specific dictionaries (FSDs), are simultaneously uncovered. Here, CSDs and UD, respectively, model the patterns for each class and the patterns irrespective of any class. FSDs can help reveal the shared patterns between multiple image classes, by filling the gap between the patterns in CSDs and UD. The dependence among image classes is revealed by the shared FSDs, and a common FSD can be assigned to several classes to represent their residual. Finally, the most discriminative FSD for each class is identified by minimizing the sparse reconstruction error. Extensive experiments are conducted on different widely used data sets for image classification. The results demonstrate the superior performance of the proposed method over some state-of-the-art methods. Xiangbo Shu, Jinhui Tang 0001, Guo-Jun Qi, Zechao Li, Yu-Gang Jiang 0001, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Weakly Supervised Multimodal Hashing for Scalable Social Image RetrievalabstractRecent years have witnessed a dramatic increase in the number of community-contributed images. Hashing-based similarity searches for social images have been attracting considerable interest from computer vision and multimedia communities due to their computational and memory efficiency. In this paper, we propose a novel weakly supervised hashing method named weakly supervised multimodal hashing, for scalable social image retrieval. Semantic-aware hash functions are learned by jointly leveraging the weakly supervised tag information and visual information. Specifically, because user-provided tags associated with social images can describe the semantic information, the hash functions are learned by exploring the semantic structure. Unfortunately, the user-provided tags are imperfect. To avoid overfitting the weakly supervised tags, the local discriminative structure and the geometric structure in the visual space are explored. Besides, to learn compact and non-redundant hash codes, the hash functions are constrained to be orthogonal and an information theoretic regularization based on the maximum entropy principle is introduced to maximize the information provided by each hash code. The learned hash functions are orthogonal, which can avoid redundancy in the learned hash codes as much as possible. The proposed hashing learning problem is formulated as the eigenvalue problem, which can be solved efficiently. Extensive experiments are conducted on two widely used social image data sets and the encouraging performance compared with the state-of-the-art hashing techniques demonstrates the effectiveness of the proposed method. Jinhui Tang 0001, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video ClassificationabstractVideos are inherently multimodal. This paper studies the problem of exploiting the abundant multimodal clues for improved video classification performance. We introduce a novel hybrid deep learning framework that integrates useful clues from multiple modalities, including static spatial appearance information, motion patterns within a short time window, audio information, as well as long-range temporal dynamics. More specifically, we utilize three Convolutional Neural Networks (CNNs) operating on appearance, motion, and audio signals to extract their corresponding features. We then employ a feature fusion network to derive a unified representation with an aim to capture the relationships among features. Furthermore, to exploit the long-range temporal dynamics in videos, we apply two long short-term memory (LSTM) networks with extracted appearance and motion features as inputs. Finally, we also propose refining the prediction scores by leveraging contextual relationships among video semantics. The hybrid deep learning framework is able to exploit a comprehensive set of multimodal features for video classification. Through an extensive set of experiments, we demonstrate that: 1) LSTM networks that model sequences in an explicitly recurrent manner are highly complementary to the CNN models; 2) the feature fusion network that produces a fused representation through modeling feature relationships outperforms a large set of alternative fusion strategies; and 3) the semantic context of video classes can help further refine the predictions for improved performance. Experimental results on two challenging benchmarks-the UCF-101 and the Columbia Consumer Videos (CCV)-provide strong quantitative evidence that our framework can produce promising results: 93.1% on the UCF-101 and 84.5% on the CCV, outperforming several competing methods with clear margins. Yu-Gang Jiang 0001, Zuxuan Wu, Jinhui Tang 0001, Zechao Li, Xiangyang Xue 0001, Shih-Fu Chang |
IEEE Trans. Multim. | 4 |
| 2018 | Robust Structured Nonnegative Matrix Factorization for Image RepresentationabstractDimensionality reduction has attracted increasing attention, because high-dimensional data have arisen naturally in numerous domains in recent years. As one popular dimensionality reduction method, nonnegative matrix factorization (NMF), whose goal is to learn parts-based representations, has been widely studied and applied to various applications. In contrast to the previous approaches, this paper proposes a novel semisupervised NMF learning framework, called robust structured NMF, that learns a robust discriminative representation by leveraging the block-diagonal structure and the -norm (especially when ) loss function. Specifically, the problems of noise and outliers are well addressed by the -norm ( ) loss function, while the discriminative representations of both the labeled and unlabeled data are simultaneously learned by explicitly exploring the block-diagonal structure. The proposed problem is formulated as an optimization problem with a well-defined objective function solved by the proposed iterative algorithm. The convergence of the proposed optimization algorithm is analyzed both theoretically and empirically. In addition, we also discuss the relationships between the proposed method and some previous methods. Extensive experiments on both the synthetic and real-world data sets are conducted, and the experimental results demonstrate the effectiveness of the proposed method in comparison to the state-of-the-art methods. Zechao Li, Jinhui Tang 0001, Xiaofei He 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Discriminative Deep Quantization Hashing for Face Image RetrievalabstractThis paper proposes a new discriminative deep quantization hashing (DDQH) approach for large-scale face image retrieval by learning discriminative and compact binary codes. It jointly explores the discrete code learning, batch normalization quantization (BNQ) module, and end-to-end learning in one unified framework, which can guarantee the optimal compatibility of hash coding and feature learning. To learn multiscale and robust facial features, a deep network properly stacking several convolution-pooling layers and pooling layers is designed, and the facial features are obtained by fusing the outputs of the last convolutional layer and the last pooling layer. Besides, the prediction errors of the learned binary codes are minimized to learn discriminative binary codes of images. To obtain higher retrieval accuracies, a BNQ module is utilized to control quantization at a moderate level. Experiments are conducted on two widely used data sets, and the proposed DDQH method achieves encouraging improvements over some state-of-the-art hashing approaches. Jinhui Tang 0001, Zechao Li, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | L1-Norm Distance Minimization-Based Fast Robust Twin Support Vector k-Plane ClusteringabstractTwin support vector clustering (TWSVC) is a recently proposed powerful k-plane clustering method. It, however, is prone to outliers due to the utilization of squared L2-norm distance. Besides, TWSVC is computationally expensive, attributing to the need of solving a series of constrained quadratic programming problems (CQPPs) in learning each clustering plane. To address these problems, this brief first develops a new k-plane clustering method called L1-norm distance minimization-based robust TWSVC by using robust L1-norm distance. To achieve this objective, we propose a novel iterative algorithm. In each iteration of the algorithm, one CQPP is solved. To speed up the computation of TWSVC and simultaneously inherit the merit of robustness, we further propose Fast RTWSVC and design an effective iterative algorithm to optimize it. Only a system of linear equations needs to be computed in each iteration. These characteristics make our methods more powerful and efficient than TWSVC. We also conduct some insightful analysis on the existence of local minimum and the convergence of the proposed algorithms. Theoretical insights and effectiveness of our methods are further supported by promising experimental results. Qiaolin Ye, Henghao Zhao, Zechao Li, Xubing Yang, Shangbing Gao, Tongming Yin, Ning Ye 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Weakly-Supervised Deep Nonnegative Low-Rank Model for Social Image Tag Refinement and AssignmentabstractIt has been well known that the user-provided tags of social images are imperfect, i.e., there exist noisy, irrelevant or incomplete tags. It heavily degrades the performance of many multimedia tasks. To alleviate this problem, we propose a Weakly-supervised Deep Nonnegative Low-rank model (WDNL) to improve the quality of tags by integrating the low-rank model with deep feature learning. A nonnegative low-rank model is introduced to uncover the intrinsic relationships between images and tags by simultaneously removing noisy or irrelevant tags and complementing missing tags. The deep architecture is leveraged to seamlessly connect the visual content and the semantic tag. That is, the proposed model can well handle the scalability by assigning tags to new images. Extensive experiments conducted on two real-world datasets demonstrate the effectiveness of the proposed method compared with some state-of-the-art methods. Zechao Li, Jinhui Tang 0001 |
AAAI | 1 |
| 2017 | Hardware-Efficient Guided Image Filtering for Multi-label ProblemabstractThe Guided Filter (GF) is well-known for its linear complexity. However, when filtering an image with an n-channel guidance, GF needs to invert an n × n matrix for each pixel. To the best of our knowledge existing matrix inverse algorithms are inefficient on current hardwares. This shortcoming limits applications of multichannel guidance in computation intensive system such as multi-label system. We need a new GF-like filter that can perform fast multichannel image guided filtering. Since the optimal linear complexity of GF cannot be minimized further, the only way thus is to bring all potentialities of current parallel computing hardwares into full play. In this paper we propose a hardware-efficient Guided Filter (HGF), which solves the efficiency problem of multichannel guided image filtering and yields competent results when applying it to multi-label problems with synthesized polynomial multichannel guidance. Specifically, in order to boost the filtering performance, HGF takes a new matrix inverse algorithm which only involves two hardware-efficient operations: element-wise arithmetic calculations and box filtering. In order to break the linear model restriction, HGF synthesizes a polynomial multichannel guidance to introduce nonlinearity. Benefiting from our polynomial guidance and hardware-efficient matrix inverse algorithm, HGF not only is more sensitive to the underlying structure of guidance but also achieves the fastest computing speed. Due to these merits, HGF obtains state-of-the-art results in terms of accuracy and efficiency in the computation intensive multi-label systems. Longquan Dai, Mengke Yuan, Zechao Li, Xiaopeng Zhang 0001, Jinhui Tang 0001 |
CVPR | 3 |
| 2017 | Discriminative Deep Hashing for Scalable Face Image RetrievalabstractWith the explosive growth of images containing faces, scalable face image retrieval has attracted increasing attention. Due to the amazing effectiveness, deep hashing has become a popular hashing method recently. In this work, we propose a new Discriminative Deep Hashing (DDH) network to learn discriminative and compact hash codes for large-scale face image retrieval. The proposed network incorporates the end-to-end learning, the divide-and-encode module and the desired discrete code learning into a unified framework. Specifically, a network with a stack of convolution-pooling layers is proposed to extract multi-scale and robust features by merging the outputs of the third max pooling layer and the fourth convolutional layer. To reduce the redundancy among hash codes and the network parameters simultaneously, a divide-and-encode module to generate compact hash codes. Moreover, a loss function is introduced to minimize the prediction errors of the learned hash codes, which can lead to discriminative hash codes. Extensive experiments on two datasets demonstrate that the proposed method achieves superior performance compared with some state-of-the-art hashing methods. Zechao Li, Jinhui Tang 0001 |
IJCAI | 2 |
| 2017 | Wheel: Accelerating CNNs with Distributed GPUs via Hybrid Parallelism and Alternate StrategyabstractConvolutional Neural Networks (CNNs) have been widely used and achieve amazing performance, typically at the cost of very expensive computation. Some methods accelerate the CNN training by distributed GPUs those deploying GPUs on multiple servers. Unfortunately, they need to transmit a large amount of data among servers, which leads to long data transmitting time and long GPU idle time. Towards this end, we propose a novel hybrid parallelism architecture named "Wheel" to accelerate the CNN training by reducing the transmitted data and fully using GPUs simultaneously. Specifically, Wheel first partitions the layers of a CNN into two kinds of modules: convolutional module and fully-connected module, and deploys them following the proposed hybrid parallelism. In this way, Wheel transmits only a few parameters of CNNs among different servers, and transmits most of the parameters within the same server. The time to transmit data is significantly reduced. Second, to fully run each GPU and reduce the idle time, Wheel devises an alternate strategy deploying multiple workers on each GPU. Once one worker is suspended for receiving data, another one in the same GPU starts to execute the computing task. The workers in each GPU run concurrently and repeatedly like Wheels. Experiments are conducted to show the outperformance of the proposed scheme over the state-of-the-art parallel approaches. Xiaoyu Du 0002, Jinhui Tang 0001, Zechao Li, Zhiguang Qin |
ACM Multimedia | 3 |
| 2017 | Learning discriminative supplementary features to attributes for novel-category classificationabstractSemantic attributes have been introduced as an effective representation for image classification especially in zero-shot learning. However, most of the existing semantic attributes are previously defined by people, thus the size of the attribute is restricted in practice and these attributes are not necessarily discriminative. Therefore, the classification accuracy is often relatively low using a fixed incomplete semantic attribute set for image representation. One intuitive solution is to expand the semantic attribute representation with some non-semantic features. However, how to make the supplementary features more effective and discriminative is still an open problem. In this paper, we propose a Discriminative Supplementary Feature Learning (DSFL) method to implement semantic attribute augmentation. In DSFL, the non-semantic supplementary features are learned simultaneously with the classifiers for the novel-categories. Extensive experiments are conducted on two public datasets and the results show that our approach achieves encouraging performance. Qiuli Liu, Zechao Li, Jinhui Tang 0001 |
VCIP | 2 |
| 2017 | Multimedia news QA: Extraction and visualization integration with multiple-source information
Xueming Wang, Zechao Li, Jinhui Tang 0001 |
Image Vis. Comput. | 2 |
| 2017 | Tri-Clustered Tensor Completion for Social-Aware Image Tag RefinementabstractSocial image tag refinement, which aims to improve tag quality by automatically completing the missing tags and rectifying the noise-corrupted ones, is an essential component for social image search. Conventional approaches mainly focus on exploring the visual and tag information, without considering the user information, which often reveals important hints on the (in)correct tags of social images. Towards this end, we propose a novel tri-clustered tensor completion framework to collaboratively explore these three kinds of information to improve the performance of social image tag refinement. Specifically, the inter-relations among users, images and tags are modeled by a tensor, and the intra-relations between users, images and tags are explored by three regularizations respectively. To address the challenges of the super-sparse and large-scale tensor factorization that demands expensive computing and memory cost, we propose a novel tri-clustering method to divide the tensor into a certain number of sub-tensors by simultaneously clustering users, images and tags into a bunch of tri-clusters. And then we investigate two strategies to complete these sub-tensors by considering (in)dependence between the sub-tensors. Experimental results on a real-world social image database demonstrate the superiority of the proposed method compared with the state-of-the-art methods. Jinhui Tang 0001, Xiangbo Shu, Guo-Jun Qi, Zechao Li, Meng Wang 0001, Shuicheng Yan, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Weakly Supervised Deep Matrix Factorization for Social Image UnderstandingabstractThe number of images associated with weakly supervised user-provided tags has increased dramatically in recent years. User-provided tags are incomplete, subjective and noisy. In this paper, we focus on the problem of social image understanding, i.e., tag refinement, tag assignment, and image retrieval. Different from previous work, we propose a novel weakly supervised deep matrix factorization algorithm, which uncovers the latent image representations and tag representations embedded in the latent subspace by collaboratively exploring the weakly supervised tagging information, the visual structure, and the semantic structure. Due to the well-known semantic gap, the hidden representations of images are learned by a hierarchical model, which are progressively transformed from the visual feature space. It can naturally embed new images into the subspace using the learned deep architecture. The semantic and visual structures are jointly incorporated to learn a semantic subspace without overfitting the noisy, incomplete, or subjective tags. Besides, to remove the noisy or redundant visual features, a sparse model is imposed on the transformation matrix of the first layer in the deep architecture. Finally, a unified optimization problem with a well-defined objective function is developed to formulate the proposed problem and solved by a gradient descent procedure with curvilinear search. Extensive experiments on real-world social image databases are conducted on the tasks of image understanding: image tag refinement, assignment, and retrieval. Encouraging results are achieved with comparison with the state-of-the-art algorithms, which demonstrates the effectiveness of the proposed method. Zechao Li, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Domain-sensitive Recommendation with user-item subgroup analysisabstractIn this paper, we propose a Domain-sensitive Recommendation (DsRec) algorithm, to make the rating prediction by exploring the user-item subgroup analysis simultaneously, in which a user-item subgroup is deemed as a domain consisting of a subset of items with similar attributes and a subset of users who have interests in these items. The proposed framework of DsRec includes three components: a matrix factorization model for the observed rating reconstruction, a bi-clustering model for the user-item subgroup analysis, and two regularization terms to connect the above two components into a unified formulation. Extensive experiments on three real-world datasets show that our method achieves the better performance over some state-of-the-art methods. Jing Liu 0001, Zechao Li, Xi Zhang 0018, Hanqing Lu |
ICDE | 3 |
| 2016 | Object co-segmentation via salient and common regions discovery
Yong Li 0034, Jing Liu 0001, Zechao Li, Hanqing Lu, Songde Ma |
Neurocomputing | 3 |
| 2016 | Projective nonnegative matrix factorization for social image retrieval
Qiuli Liu, Zechao Li |
Neurocomputing | 2 |
| 2016 | Age progression: Current technologies and applications
Xiangbo Shu, Guosen Xie, Zechao Li, Jinhui Tang 0001 |
Neurocomputing | 3 |
| 2016 | Overlapping community detection based on node location analysis
Zhi-Xiao Wang, Zechao Li, Xiao-fang Ding, Jinhui Tang 0001 |
Knowl. Based Syst. | 2 |
| 2016 | Multimedia News Summarization in SearchabstractIt is a necessary but challenging task to relieve users from the proliferative news information and allow them to quickly and comprehensively master the information of the whats and hows that are happening in the world every day. In this article, we develop a novel approach of multimedia news summarization for searching results on the Internet, which uncovers the underlying topics among query-related news information and threads the news events within each topic to generate a query-related brief overview. First, the hierarchical latent Dirichlet allocation (hLDA) model is introduced to discover the hierarchical topic structure from query-related news documents, and a new approach based on the weighted aggregation and max pooling is proposed to identify one representative news article for each topic. One representative image is also selected to visualize each topic as a complement to the text information. Given the representative documents selected for each topic, a time-bias maximum spanning tree (MST) algorithm is proposed to thread them into a coherent and compact summary of their parent topic. Finally, we design a friendly interface to present users with the hierarchical summarization of their required news information. Extensive experiments conducted on a large-scale news dataset collected from multiple news Web sites demonstrate the encouraging performance of the proposed solution for news summarization in news retrieval. Zechao Li, Jinhui Tang 0001, Xueming Wang, Jing Liu 0001, Hanqing Lu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2016 | Domain-Sensitive Recommendation with User-Item Subgroup AnalysisabstractCollaborative Filtering (CF) is one of the most successful recommendation approaches to cope with information overload in the real world. However, typical CF methods equally treat every user and item, and cannot distinguish the variation of user's interests across different domains. This violates the reality that user's interests always center on some specific domains, and the users having similar tastes on one domain may have totally different tastes on another domain. Motivated by the observation, in this paper, we propose a novel Domain-sensitive Recommendation (DsRec) algorithm, to make the rating prediction by exploring the user-item subgroup analysis simultaneously, in which a user-item subgroup is deemed as a domain consisting of a subset of items with similar attributes and a subset of users who have interests in these items. The proposed framework of DsRec includes three components: a matrix factorization model for the observed rating reconstruction, a bi-clustering model for the user-item subgroup analysis, and two regularization terms to connect the above two components into a unified formulation. Extensive experiments on Movielens-100K and two real-world product review datasets show that our method achieves the better performance in terms of prediction accuracy criterion over the state-of-the-art methods. Jing Liu 0001, Zechao Li, Xi Zhang 0018, Hanqing Lu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Generalized Deep Transfer Networks for Knowledge Propagation in Heterogeneous DomainsabstractIn recent years, deep neural networks have been successfully applied to model visual concepts and have achieved competitive performance on many tasks. Despite their impressive performance, traditional deep networks are subjected to the decayed performance under the condition of lacking sufficient training data. This problem becomes extremely severe for deep networks trained on a very small dataset, making them overfitting by capturing nonessential or noisy information in the training set. Toward this end, we propose a novel generalized deep transfer networks (DTNs), capable of transferring label information across heterogeneous domains, textual domain to visual domain. The proposed framework has the ability to adequately mitigate the problem of insufficient training images by bringing in rich labels from the textual domain. Specifically, to share the labels between two domains, we build parameter- and representation-shared layers. They are able to generate domain-specific and shared interdomain features, making this architecture flexible and powerful in capturing complex information from different domains jointly. To evaluate the proposed method, we release a new dataset extended from NUS-WIDE at http://imag.njust.edu.cn/NUS-WIDE-128.html. Experimental results on this dataset show the superior performance of the proposed DTNs compared to existing state-of-the-art methods. Jinhui Tang 0001, Xiangbo Shu, Zechao Li, Guo-Jun Qi, Jingdong Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2015 | Face Clustering in Videos with Proportion Prior
Yifan Zhang 0001, Zechao Li, Hanqing Lu |
IJCAI | 3 |
| 2015 | Semantic-aware Hashing for Social Image RetrievalabstractWith the proliferation of large-scale social images, recent years have witnessed the increasing amount of images with user-provided tags, which leads to considerable effort made on hashing based approximate nearest neighbor (ANN) search in huge databases. In this work, we propose a novel Semantic-aware Hashing method (SaH) by discovering knowledge from these social media resources to implement approximate similarity search. Different from the previous work, the proposed method learns semantic hashing codes by exploiting heterogeneous information from the textual and visual domains. The semantic structure in the textual domain is well preserved to learn the binary codes. To handle the noisy, incomplete, or subjective user-provided tags, the visual structure is also leveraged. On the other hand, an information theoretic regularization is exploited by using maximum entropy principle and a row-wise sparse model with l2,p (0 < p ≤ 1) mixed norm is introduced to filter certain noisy or redundant visual features. Experiments are conducted on a widely-used social image dataset and the comparison results demonstrate the outperforming performance of the proposed SaH method over state-of-the-art hashing techniques. Jinhui Tang 0001, Zechao Li, Liyan Zhang 0001, Qingming Huang |
ICMR | 2 |
| 2015 | Partially Common-Semantic Pursuit for RGB-D Object RecognitionabstractFor the RGB-D object recognition task, the robust and rich representations can boost the performance. Most works employ feature learning approaches to learn specific representation for the RGB and depth modalities independently, while some directly learn common property. Different from them, this paper proposes a novel supervised feature learning method for RGB-D object recognition, named Partially Common-Semantic Learning (PCSL), which jointly captures the complementary and consistency semantic information from RGB and depth modalities. The complementary information is revealed by the individual modality, while the consistency is exploited by both modalities simultaneously. In PCSL, Reconstruction Independent Component Analysis (RICA) is extended to integrate the supervised information and learn both of the complementary and partially shared common semantic information. The proposed approach is evaluated on two public RGB-D datasets and achieves better performance than several state-of-the-art methods. Lu Jin 0001, Zechao Li, Xiangbo Shu, Shenghua Gao, Jinhui Tang 0001 |
ACM Multimedia | 2 |
| 2015 | Deep Matrix Factorization for social image tag refinement and assignmentabstractThe number of images associated with user-provided tags has increased dramatically in recent years. User-provided tags are incomplete, subjective and noisy. In this work, we focus on the problem of image tag refinement and assignment. Different from previous work, we propose a novel Deep Matrix Factorization (DMF) algorithm, which uncovers the latent image representations and tag representations embedded in the latent subspace by exploiting the weakly-supervised tagging information and visual information. Due to the well-known semantic gap, the hidden representations of images are learned by a hierarchical model, which are progressively transformed from the visual feature space. It can naturally embed new images into the subspace using the learned deep architecture. Besides, to remove the noisy or redundant visual features, a sparse model is imposed on the transformation matrix of the first layer in the deep architecture. Finally, a unified optimization problem with a well-defined objective function is developed to formulate the proposed problem. Extensive experiments on real-world social image databases are conducted on the tasks of image tag refinement and assignment. Encouraging results are achieved with comparison to the state-of-the-art algorithms, which demonstrates the effectiveness of the proposed method. Zechao Li, Jinhui Tang 0001 |
MMSP | 1 |
| 2015 | Deep kinship verificationabstractTo improve the performance of kinship verification, we propose a novel deep kinship verification (DKV) model by integrating excellent deep learning architecture into metric learning. Unlike most existing shallow models based on metric learning for kinship verification, we employ a deep learning model followed by a metric learning formulation to select nonlinear features, which can find the appropriate project space to ensure the margin of negative sample pairs (i.e. parent and child without kinship relation) as large as possible and the margin of positive sample pairs (i.e. parent and child with kinship relation) as small as possible. Experimental results show that our method achieves satisfactory performance on two widely-used benchmarks, i.e. KFW-I and KFW-II. Mengyin Wang, Zechao Li, Xiangbo Shu, Jingdong Wang 0001, Jinhui Tang 0001 |
MMSP | 2 |
| 2015 | Tag ranking based on salient region graph propagation
Jinhui Tang 0001, Minxian Li, Zechao Li, Chunxia Zhao |
Multim. Syst. | 3 |
| 2015 | Boosted MIML method for weakly-supervised image semantic segmentation
Yang Liu 0021, Zechao Li, Jing Liu 0001, Hanqing Lu |
Multim. Tools Appl. | 2 |
| 2015 | Robust Structured Subspace Learning for Data RepresentationabstractTo uncover an appropriate latent subspace for data representation, in this paper we propose a novel Robust Structured Subspace Learning (RSSL) algorithm by integrating image understanding and feature learning into a joint learning framework. The learned subspace is adopted as an intermediate space to reduce the semantic gap between the low-level visual features and the high-level semantics. To guarantee the subspace to be compact and discriminative, the intrinsic geometric structure of data, and the local and global structural consistencies over labels are exploited simultaneously in the proposed algorithm. Besides, we adopt the l2,1-norm for the formulations of loss function and regularization respectively to make our algorithm robust to the outliers and noise. An efficient algorithm is designed to solve the proposed optimization problem. It is noted that the proposed framework is a general one which can leverage several well-known algorithms as special cases and elucidate their intrinsic relationships. To validate the effectiveness of the proposed method, extensive experiments are conducted on diversity datasets for different image understanding tasks, i.e., image tagging, clustering, and classification, and the more encouraging results are achieved compared with some state-of-the-art approaches. Zechao Li, Jing Liu 0001, Jinhui Tang 0001, Hanqing Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Unsupervised Feature Selection via Nonnegative Spectral Analysis and Redundancy ControlabstractIn many image processing and pattern recognition problems, visual contents of images are currently described by high-dimensional features, which are often redundant and noisy. Toward this end, we propose a novel unsupervised feature selection scheme, namely, nonnegative spectral analysis with constrained redundancy, by jointly leveraging nonnegative spectral clustering and redundancy analysis. The proposed method can directly identify a discriminative subset of the most useful and redundancy-constrained features. Nonnegative spectral analysis is developed to learn more accurate cluster labels of the input images, during which the feature selection is performed simultaneously. The joint learning of the cluster labels and feature selection matrix enables to select the most discriminative features. Row-wise sparse models with a general ℓ(2, p)-norm (0 < p ≤ 1) are leveraged to make the proposed model suitable for feature selection and robust to noise. Besides, the redundancy between features is explicitly exploited to control the redundancy of the selected subset. The proposed problem is formulated as an optimization problem with a well-defined objective function solved by the developed simple yet efficient iterative algorithm. Finally, we conduct extensive experiments on nine diverse image benchmarks, including face data, handwritten digit data, and object image data. The proposed method achieves encouraging the experimental results in comparison with several representative algorithms, which demonstrates the effectiveness of the proposed algorithm for unsupervised feature selection. Zechao Li, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Neighborhood Discriminant Hashing for Large-Scale Image RetrievalabstractWith the proliferation of large-scale community-contributed images, hashing-based approximate nearest neighbor search in huge databases has aroused considerable interest from the fields of computer vision and multimedia in recent years because of its computational and memory efficiency. In this paper, we propose a novel hashing method named neighborhood discriminant hashing (NDH) (for short) to implement approximate similarity search. Different from the previous work, we propose to learn a discriminant hashing function by exploiting local discriminative information, i.e., the labels of a sample can be inherited from the neighbor samples it selects. The hashing function is expected to be orthogonal to avoid redundancy in the learned hashing bits as much as possible, while an information theoretic regularization is jointly exploited using maximum entropy principle. As a consequence, the learned hashing function is compact and nonredundant among bits, while each bit is highly informative. Extensive experiments are carried out on four publicly available data sets and the comparison results demonstrate the outperforming performance of the proposed NDH method over state-of-the-art hashing techniques. Jinhui Tang 0001, Zechao Li, Meng Wang 0001, Ruizhen Zhao |
IEEE Trans. Image Process. | 2 |
| 2015 | Weakly Supervised Deep Metric Learning for Community-Contributed Image RetrievalabstractRecent years have witnessed the explosive growth of community-contributed images with rich context information, which is beneficial to the task of image retrieval. It can help us to learn a suitable metric to alleviate the semantic gap. In this paper, we propose a new distance metric learning algorithm, namely weakly-supervised deep metric learning (WDML), under the deep learning framework. It utilizes a progressive learning manner to discover knowledge by jointly exploiting the heterogeneous data structures from visual contents and user-provided tags of social images. The semantic structure in the textual space is expected to be well preserved while the problem of the noisy, incomplete or subjective tags is addressed by leveraging the visual structure in the original visual space. Besides, a sparse model with the l2,1 mixed norm is imposed on the transformation matrix of the first layer in the deep architecture to compress the noisy or redundant visual features. The proposed problem is formulated as an optimization problem with a well-defined objective function and a simple yet efficient iterative algorithm is proposed to solve it. Extensive experiments on real-world social image datasets are conducted to verify the effectiveness of the proposed method for image retrieval. Encouraging experimental results are achieved compared with several representative metric learning methods. Zechao Li, Jinhui Tang 0001 |
IEEE Trans. Multim. | 1 |
| 2015 | RGB-D Object Recognition via Incorporating Latent Data Structure and Prior KnowledgeabstractFor the task of RGB-D object recognition, it is important to identify suitable representations of images, which can boost the performance of object recognition. In this work, we propose a novel representation learning method for RGB-D images by jointly incorporating the underlying data structure and the prior knowledge of the data. Specifically, the convolutional neural networks (CNN) are employed to learn image representation by exploiting the underlying data structure. To handle the problem of the limited RGB and depth images for object recognition, the multi-level hierarchies of features trained on ImageNet from the CNN are transferred to learn rich generic feature representation for RGB and depth images while the labeled images are leveraged. On the other hand, we propose a novel deep auto-encoders (DAE) to exploit the prior knowledge, which can overcome the expensive computational cost of optimization in feature encoding. The expected representations of images are obtained by integrating the two types of image representations. To verify the effectiveness of the proposed method, we thoroughly conduct extensive experiments on two publicly available RGB-D datasets. The encouraging experimental results compared with the state-of-the-art approaches demonstrate the advantages of the proposed method. Jinhui Tang 0001, Lu Jin 0001, Zechao Li, Shenghua Gao |
IEEE Trans. Multim. | 3 |
| 2015 | Partially Shared Latent Factor Learning With Multiview DataabstractMultiview representations reveal the fundamental attributes of the studied instances from different perspectives. Some common perspectives are reviewed by multiple views simultaneously, while some specific ones are reflected by individual views. That is, there are two kinds of properties embedded in the multiview data: 1) consistency and 2) complementarity. Different from most multiview learning approaches only focusing on either consistency or complementarity, this paper proposes a novel semisupervised multiview learning algorithm, called partially shared latent factor (PSLF) learning, which jointly exploits both consistent and complementary information among multiple views. In PSLF, a nonnegative matrix factorization (NMF)-based formulation is adopted to learn a compact and comprehensive partially shared latent representation, which is composed of common latent factors shared by multiple views and some specific latent factors to each view. With the learned representations of multiview data, we introduce a robust sparse regression model to predict the cluster labels of labeled data. By integrating the NMF-based model and the regression model, we obtain a unified formulation and propose a multiplicative-based alternative algorithm for optimization. In addition, PSLF can learn the weights of different views adaptively according to the reconstruction precisions of data matrices. Our experimental study indicates different multiview data that contains consistent and complementary information in different degrees. In addition, the encouraging results of the proposed algorithm are achieved in comparison with the state-of-the-art algorithms on real-world data sets. Jing Liu 0001, Zechao Li, Zhi-Hua Zhou, Hanqing Lu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | Learning Low-Rank Representations with Classwise Block-Diagonal Structure for Robust Face RecognitionabstractFace recognition has been widely studied due to its importance in various applications. However, the case that both training images and testing images are corrupted is not well addressed. Motivated by the success of low-rank matrix recovery, we propose a novel semi-supervised low-rank matrix recovery algorithm for robust face recognition. The proposed method can learn robust discriminative representations for both training images and testing images simultaneously by exploiting the classwise block-diagonal structure. Specifically, low-rank matrix approximation can handle the possible contamination of data. Moreover, the classwise block-diagonal structure is exploited to promote discrimination of representations for robust recognition. The above issues are formulated into a unified objective function and we design an efficient optimization procedure based on augmented Lagrange multiplier method to solve it. Extensive experiments on three public databases are performed to validate the effectiveness of our approach. The strong identification capability of representations with block-diagonal structure is verified. Yong Li 0034, Jing Liu 0001, Zechao Li, Yangmuzi Zhang, Hanqing Lu, Songde Ma |
AAAI | 3 |
| 2014 | Image Representation Learning by Deep Appearance and Spatial Coding
Bingyuan Liu, Jing Liu 0001, Zechao Li, Hanqing Lu |
ACCV (1) | 3 |
| 2014 | Hand-Crafted Features or Machine Learnt Features? Together They Improve RGB-D Object RecognitionabstractRGB-D object recognition is an important research topic in computer version, and seeking a robust image representation is the most important sub problem for RGB-D object recognition. On the one hand, the recently emerging deep learning methods, which learns image representations automatically by capturing the data structure, have demonstrated the impressive performance for object recognition. On the other hand, the previously commonly used hand-crafted features also encodes the prior knowledge about the data. By realizing that the hand-crafted features and machine learnt features actually characterize the different aspects of image data, rather than only using one type of feature, we propose to jointly use the machine learnt features and hand-crafted features for RGB-D object recognition. Specifically, we use the Convolution Neural Networks (CNNs) to extract the machine learnt representation, and use Locality-constrained Linear Coding (LLC) based spatial pyramid matching for hand-crafted features. We evaluated our proposed approach on three publicly available RGB-D datasets. Experimental results show that our method achieves the best performance under all the cases, which demonstrates the effectiveness of our method. Lu Jin 0001, Shenghua Gao, Zechao Li, Jinhui Tang 0001 |
ISM | 3 |
| 2014 | Projective Matrix Factorization with unified embedding for social image tagging
Zechao Li, Jing Liu 0001, Jinhui Tang 0001, Hanqing Lu |
Comput. Vis. Image Underst. | 1 |
| 2014 | Sparse semantic metric learning for image retrieval
Jing Liu 0001, Zechao Li, Hanqing Lu |
Multim. Syst. | 2 |
| 2014 | Semi-supervised Unified Latent Factor learning with multi-view data
Jing Liu 0001, Zechao Li, Hanqing Lu |
Mach. Vis. Appl. | 3 |
| 2014 | Clustering-Guided Sparse Structural Learning for Unsupervised Feature SelectionabstractMany pattern analysis and data mining problems have witnessed high-dimensional data represented by a large number of features, which are often redundant and noisy. Feature selection is one main technique for dimensionality reduction that involves identifying a subset of the most useful features. In this paper, a novel unsupervised feature selection algorithm, named clustering-guided sparse structural learning (CGSSL), is proposed by integrating cluster analysis and sparse structural analysis into a joint framework and experimentally evaluated. Nonnegative spectral clustering is developed to learn more accurate cluster labels of the input samples, which guide feature selection simultaneously. Meanwhile, the cluster labels are also predicted by exploiting the hidden structure shared by different features, which can uncover feature correlations to make the results more reliable. Row-wise sparse models are leveraged to make the proposed model suitable for feature selection. To optimize the proposed formulation, we propose an efficient iterative algorithm. Finally, extensive experiments are conducted on 12 diverse benchmarks, including face data, handwritten digit data, document data, and biomedical data. The encouraging experimental results in comparison with several representative algorithms and the theoretical analysis demonstrate the efficiency and effectiveness of the proposed algorithm for feature selection. Zechao Li, Jing Liu 0001, Yi Yang 0001, Xiaofang Zhou 0001, Hanqing Lu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | Personalized Geo-Specific Tag Recommendation for Photos on Social WebsitesabstractSocial tagging becomes increasingly important to organize and search large-scale community-contributed photos on social websites. To facilitate generating high-quality social tags, tag recommendation by automatically assigning relevant tags to photos draws particular research interest. In this paper, we focus on the personalized tag recommendation task and try to identify user-preferred, geo-location-specific as well as semantically relevant tags for a photo by leveraging rich contexts of the freely available community-contributed photos. For users and geo-locations, we assume they have different preferred tags assigned to a photo, and propose a subspace learning method to individually uncover the both types of preferences. The goal of our work is to learn a unified subspace shared by the visual and textual domains to make visual features and textual information of photos comparable. Considering the visual feature is a lower level representation on semantics than the textual information, we adopt a progressive learning strategy by additionally introducing an intermediate subspace for the visual domain, and expect it to have consistent local structure with the textual space. Accordingly, the unified subspace is mapped from the intermediate subspace and the textual space respectively. We formulate the above learning problems into a united form, and present an iterative optimization with its convergence proof. Given an untagged photo with its geo-location to a user, the user-preferred and the geo-location-specific tags are found by the nearest neighbor search in the corresponding unified spaces. Then we combine the obtained tags and the visual appearance of the photo to discover the semantically and visually related photos, among which the most frequent tags are used as the recommended tags. Experiments on a large-scale data set collected from Flickr verify the effectivity of the proposed solution. Jing Liu 0001, Zechao Li, Jinhui Tang 0001, Hanqing Lu |
IEEE Trans. Multim. | 2 |
| 2013 | Weakly-Supervised Dual Clustering for Image Semantic SegmentationabstractIn this paper, we propose a novel Weakly-Supervised Dual Clustering (WSDC) approach for image semantic segmentation with image-level labels, i.e., collaboratively performing image segmentation and tag alignment with those regions. The proposed approach is motivated from the observation that super pixels belonging to an object class usually exist across multiple images and hence can be gathered via the idea of clustering. In WSDC, spectral clustering is adopted to cluster the super pixels obtained from a set of over-segmented images. At the same time, a linear transformation between features and labels as a kind of discriminative clustering is learned to select the discriminative features among different classes. The both clustering outputs should be consistent as much as possible. Besides, weakly-supervised constraints from image-level labels are imposed to restrict the labeling of super pixels. Finally, the non-convex and non-smooth objective function are efficiently optimized using an iterative CCCP procedure. Extensive experiments conducted on MSRC and Label Me datasets demonstrate the encouraging performance of our method in comparison with some state-of-the-arts. Yang Liu 0021, Jing Liu 0001, Zechao Li, Jinhui Tang 0001, Hanqing Lu |
CVPR | 3 |
| 2013 | Object co-segmentation via discriminative low rank matrix recoveryabstractThe goal of this paper is to simultaneously segment the object regions appearing in a set of images of the same object class, known as object co-segmentation. Different from typical methods, simply assuming that the regions common among images are the object regions, we additionally consider the disturbance from consistent backgrounds, and indicate not only common regions but salient ones among images to be the object regions. To this end, we propose a Discriminative Low Rank matrix Recovery (DLRR) algorithm to divide the over-completely segmented regions (i.e.,superpixels) of a given image set into object and non-object ones. In DLRR, a low-rank matrix recovery term is adopted to detect salient regions in an image, while a discriminative learning term is used to distinguish the object regions from all the super-pixels. An additional regularized term is imported to jointly measure the disagreement between the predicted saliency and the objectiveness probability corresponding to each super-pixel of the image set. For the unified learning problem by connecting the above three terms, we design an efficient optimization procedure based on block-coordinate descent. Extensive experiments are conducted on two public datasets, i.e., MSRC and iCoseg, and the comparisons with some state-of-the-arts demonstrate the effectiveness of our work. Yong Li 0034, Jing Liu 0001, Zechao Li, Yang Liu 0021, Hanqing Lu |
ACM Multimedia | 3 |
| 2013 | Structure preserving non-negative matrix factorization for dimensionality reduction
Zechao Li, Jing Liu 0001, Hanqing Lu |
Comput. Vis. Image Underst. | 1 |
| 2013 | Nonlinear matrix factorization with unified embedding for social tag relevance learning
Zechao Li, Jing Liu 0001, Hanqing Lu |
Neurocomputing | 1 |
| 2013 | Correlation consistency constrained probabilistic matrix factorization for social tag refinement
Jing Liu 0001, Yifan Zhang 0001, Zechao Li, Hanqing Lu |
Neurocomputing | 3 |
| 2013 | MLRank: Multi-correlation Learning to Rank for image annotation
Zechao Li, Jing Liu 0001, Changsheng Xu, Hanqing Lu |
Pattern Recognit. | 1 |
| 2013 | Enhancing news organization for convenient retrieval and browsingabstractTo facilitate users to access news quickly and comprehensively, we design a news search and browsing system named GeoVisNews, in which the news elements of “Where”, “Who”, “What” and “When” are enhanced via news geo-localization, image enrichment and joint ranking, respectively. For news geo-localization, an Ordinal Correlation Consistent Matrix Factorization (OCCMF) model is proposed to maintain the relevance rankings of locations to a specific news document and simultaneously capture intra-relations among locations and documents. To visualize news, we develop a novel method to enrich news documents with appropriate web images. Specifically, multiple queries are first generated from news documents for image search, and then the appropriate images are selected from the collected web images by an intelligent fusion approach based on multiple features. Obtaining the geo-localized and image enriched news resources, we further employ a joint ranking strategy to provide relevant, timely and popular news items as the answer of user searching queries. Extensive experiments on a large-scale news dataset collected from the web demonstrate the superior performance of the proposed approaches over related methods. Zechao Li, Jing Liu 0001, Meng Wang 0001, Changsheng Xu, Hanqing Lu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2012 | Unsupervised Feature Selection Using Nonnegative Spectral AnalysisabstractIn this paper, a new unsupervised learning algorithm, namely Nonnegative Discriminative Feature Selection (NDFS), is proposed. To exploit the discriminative information in unsupervised scenarios, we perform spectral clustering to learn the cluster labels of the input samples, during which the feature selection is performed simultaneously. The joint learning of the cluster labels and feature selection matrix enables NDFS to select the most discriminative features. To learn more accurate cluster labels, a nonnegative constraint is explicitly imposed to the class indicators. To reduce the redundant or even noisy features, l2,1-norm minimization constraint is added into the objective function, which guarantees the feature selection matrix sparse in rows. Our algorithm exploits the discriminative information and feature correlation simultaneously to select a better feature subset. A simple yet efficient iterative algorithm is designed to optimize the proposed objective function. Experimental results on different real world datasets demonstrate the encouraging performance of our algorithm over the state-of-the-arts. Zechao Li, Yi Yang 0001, Jing Liu 0001, Xiaofang Zhou 0001, Hanqing Lu |
AAAI | 1 |
| 2012 | Efficient Clothing Retrieval with Semantic-Preserving Visual Phrases
Jianlong Fu, Jinqiao Wang, Zechao Li, Min Xu 0001, Hanqing Lu |
ACCV (2) | 3 |
| 2012 | Co-regularized PLSA for Multi-view Clustering
Jing Liu 0001, Zechao Li, Hanqing Lu |
ACCV (2) | 3 |
| 2012 | Learning Semantic Motion Patterns for Dynamic Scenes by Improved Sparse Topical CodingabstractWith the proliferation of cameras in public areas, it becomes increasingly desirable to develop fully automated surveillance and monitoring systems. In this paper, we propose a novel unsupervised approach to automatically explore motion patterns occurring in dynamic scenes under an improved sparse topical coding (STC) framework. Given an input video with a fixed camera, we first segment the whole video into a sequence of clips (documents) without overlapping. Optical flow features are extracted from each pair of consecutive frames, and quantized into discrete visual words. Then the video is represented by a word-document hierarchical topic model through a generative process. Finally, an improved sparse topical coding approach is proposed for model learning. The semantic motion patterns (latent topics) are learned automatically and each video clip is represented as a weighted summation of these patterns with only a few nonzero coefficients. The proposed approach is purely data-driven and scene independent (not an object-class specific), which make it suitable for very large range of scenarios. Experiments demonstrate that our approach outperforms the state-of-the art technologies in dynamic scene analysis. Jinqiao Wang, Zechao Li, Hanqing Lu, Songde Ma |
ICME | 3 |
| 2012 | Noisy Tag Alignment with Image RegionsabstractWith the permeation of Web 2.0, large-scale user contributed images with tags are easily available on social websites. How to align these social tags with image regions is a challenging task while no additional human intervention is considered, but a valuable one since the alignment can provide more detailed image semantic information and improve the accuracy of image retrieval. To this end, we propose a large margin discriminative model for automatically locating unaligned and possibly noisy image-level tags to the corresponding regions, and the model is optimized using concave-convex procedure (CCCP). In the model, each image is considered as a bag of segmented regions, associated with a set of candidate labeling vectors. Each labeling vector encodes a possible label arrangement for the regions of an image. To make the size of admissible labels tractable, we adopt an effective strategy based on the consistency between visual similarity and semantic correlation to generate a more compact set of labeling vectors. Extensive experiments on MSRC and SAIAPR TC-12 databases have been conducted to demonstrate the encouraging performance of our method comparing with other baseline methods. Yang Liu 0021, Jing Liu 0001, Zechao Li, Hanqing Lu |
ICME | 3 |
| 2012 | Collaborative PLSA for multi-view clustering
Jing Liu 0001, Zechao Li, Hanqing Lu |
ICPR | 3 |
| 2012 | Low rank metric learning for social image retrievalabstractWith the popularity of social media applications, large amounts of social images associated with rich context are available, which is helpful for many applications. In this paper, we propose a Low Rank distance Metric Learning (LRML) algorithm by discovering knowledge from these rich contextual data, to boost the performance of CBIR. Different from traditional approaches that often use the must-links and cannot-links between images, the proposed method exploits information from the visual and textual domains. We assume that the visual similarity estimated by the learned metric is expected to be consistent with the semantic similarity in the textual domain. Since tags are usually noisy, misspelling or meaningless, we also leverage the preservation of visual structure to prevent overfitting those noisy tags. On the other hand, the metric is straightforward constrained to be low rank. We formulate it as a convex optimization problem with nuclear norm minimization and propose an effective optimization algorithm based on proximal gradient method. With the learned metric for image retrieval, some experimental evaluations on a real-world dataset demonstrate the outperformance of our approach over other related work. Zechao Li, Jing Liu 0001, Jinhui Tang 0001, Hanqing Lu |
ACM Multimedia | 1 |
| 2012 | Social tag alignment with image regions by sparse reconstructionsabstractHow to align social tags with image regions without additional human intervention is a challenging but a valuable task since it can provide more detailed image semantic information and improve the accuracy of image retrieval. To this end, we propose a novel tag-to-region method with two phases of sparse reconstructions by exploring the large-scale user contributed resources. Given an image with social tags, we first explore the tagging information of large-scale social images to sparsely reconstruct the label vector of the given image, and then use the reconstructing weights as the semantic relevance to the image. With the top $T$ semantically relevant images, we further employ a group sparse coding algorithm to reconstruct each region of the given image, in which the regions from the social images with a common label are deemed as a label group. The group sparsity works on the assumption that one image region corresponds to tags as few as possible. Finally, the region-level tags can be predicted based on the reconstruction error in the corresponding label groups. Extensive experiments on MSRC and SAIAPR TC-12 datasets demonstrate the encouraging performance of our method in comparison with other baselines. Yang Liu 0021, Jing Liu 0001, Zechao Li, Biao Niu, Hanqing Lu |
ACM Multimedia | 3 |
| 2011 | News contextualization with geographic and visual informationabstractIn this paper, we investigate the contextualization of news documents with geographic and visual information. We propose a matrix factorization approach to analyze the location relevance for each news document. We also propose a method to enrich the document with a set of web images. For location relevance analysis, we first perform toponym extraction and expansion to obtain a toponym list from news documents. We then propose a matrix factorization method to estimate the location-document relevance scores while simultaneously capturing the correlation of locations and documents. For image enrichment, we propose a method to generate multiple queries from each news document for image search and then employ an intelligent fusion approach to collect a set of images from the search results. Based on the location relevance analysis and image enrichment, we introduce a news browsing system named NewsMap which can support users in reading news via browsing a map and retrieving news with location queries. The news documents with the corresponding enriched images are presented to help users quickly get information. Extensive experiments demonstrate the effectiveness of our approaches. Zechao Li, Meng Wang 0001, Jing Liu 0001, Changsheng Xu, Hanqing Lu |
ACM Multimedia | 1 |
| 2011 | Correlated PLSA for Image Clustering
Jian Cheng 0001, Zechao Li, Hanqing Lu |
MMM (1) | 3 |
| 2010 | Multi-modal multi-correlation person-centric news retrievalabstractIn this paper, we propose a framework of multi-modal multi-correlation person-centric news retrieval, which integrates news event correlations, news entity correlations, and event-entity correlations simultaneously by exploring both text and image information. The proposed framework is confined to a person-name query and enables a more vivid and informative person-centric news retrieval by providing two views of result presentation, namely a query-oriented multi-correlation map and a ranking list of news items with necessary descriptions including news image, news title and summary, central entities and relevant news events. First, we pre-process news articles using natural language techniques, and initialize the three correlations by statistical analysis about events and entities in news articles and face images. Second, a Multi-correlation Probabilistic Matrix Factorization (MPMF) algorithm is proposed to complete and refine the three correlations. Different from traditional Probabilistic Matrix Factorization (PMF), the proposed MPFM additionally considers the event correlations and the entity correlations as well as the event-entity correlations during the factor analysis. Third, the result ranking and visualization are conducted to present search results relevant to a target news topic. Experimental results on a news dataset collected from multiple news websites demonstrate the attractive performance of the proposed solution for news retrieval. Zechao Li, Jing Liu 0001, Xiaobin Zhu 0003, Hanqing Lu |
CIKM | 1 |
| 2010 | Sparse constraint nearest neighbour selection in cross-media retrievalabstractWith the rapid increasing multimedia documents including videos, images or text, the cross-media retrieval is being focused on. Currently, most of state-of-art methods belonging to the retrieval methods are developed within the scope of the transductive learning. And as soon as the query samples are outside the database, k-nearest-neighbor method is always adopted. However, under such circumstances the fixed global parameter k is not robust for all queries with diverse semantics. In this paper, we propose an alternative method based on the sparse representation. The query sample is considered as a sparse linear combination of all training samples, and the number of nearest neighbors is determined automatically according to the sparse coefficients to the query. Then we import the selection of nearest neighbors into a cross-media ranking model with Local Regression and Global Alignment (LRGA) to get the relevant documents to the query. We conduct extensive experiments for cross-media retrieval to demonstrate the efficiency and effectiveness of our methods. Zechao Li, Jing Liu 0001, Hanqing Lu |
ICIP | 1 |
| 2010 | Image annotation using multi-correlation probabilistic matrix factorizationabstractThe image-word correlation estimation is an essential issue in image annotation. In this paper, we propose a multi-correlation probabilistic matrix factorization (MPMF) algorithm for the correlation estimation. Different from the traditional solutions which treat the image-word correlation, image similarity and word relation independently or sequentially, in the proposed MPMF, these three elements are integrated together simultaneously and seamlessly. Specifically, we have derived two low-dimensional sets by conducting a joint factorization upon the word-to-image relation matrix, the image similarity matrix, and the word relation matrix to derive two low-dimensional sets of latent word factors and latent image factors. Finally, the annotation words of each untagged or noisily tagged image can be predicted by reconstructing the image-word correlations with the both derived latent factors. Experimental results on the Corel dataset and a Flickr image dataset show the superior performance of our proposed algorithm over the state-of-the-arts. Zechao Li, Jing Liu 0001, Xiaobin Zhu 0003, Tinglin Liu, Hanqing Lu |
ACM Multimedia | 1 |