VLDB 2026 Research / reviewers in the wild / expert
Hongtao Lu 0001
dblp:56/2480-1
· DBLP profile ↗
142ranked-venue papers
7as first author
75since 2021 · last 2026
0000-0003-2300-3039ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 77 · 4 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 73 · 1 first-author · 41 since 2021Databases, data management, data science and information retrieval · 15 · 11 since 2021Systems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HACMatch: Semi-supervised rotation regression with hardness-aware curriculum pseudo labeling
Huayi Zhou 0001, Suizhi Huang, Yue Ding 0001, Hongtao Lu 0001 |
Comput. Vis. Image Underst. | 6 |
| 2026 | Semi-Supervised Unconstrained Head Pose Estimation in the WildabstractExisting research on unconstrained in-the-wild head pose estimation suffers from the flaws of its datasets, which consist of either numerous samples by non-realistic synthesis or constrained collection, or small-scale natural images yet with plausible manual annotations. This makes fully-supervised solutions compromised due to the reliance on generous labels. To alleviate it, we propose the first semi-supervised unconstrained head pose estimation method SemiUHPE, which can leverage abundant easily available unlabeled head images. Technically, we choose semi-supervised rotation regression and adapt it to the error-sensitive and label-scarce problem of unconstrained head pose. Our method is based on the observation that the aspect-ratio invariant cropping of wild heads is superior to previous landmark-based affine alignment given that landmarks of unconstrained human heads are usually unavailable, especially for underexplored non-frontal heads. Instead of using a pre-fixed threshold to filter out pseudo labeled heads, we propose dynamic entropy based filtering to adaptively remove unlabeled outliers as training progresses by updating the threshold in multiple stages. We then revisit the design of weak-strong augmentations and improve it by devising two novel head-oriented strong augmentations, termed pose-irrelevant cut-occlusion and pose-altering rotation consistency respectively. Extensive experiments and ablation studies show that SemiUHPE outperforms its counterparts greatly on public benchmarks under both the front-range and full-range settings. Furthermore, our proposed method is also beneficial for solving other closely related problems, including generic object rotation regression and 3D head reconstruction, demonstrating good versatility and extensibility. Huayi Zhou 0001, Fei Jiang 0006, Yong Rui, Hongtao Lu 0001, Kui Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | From Anchors to Answers: A Novel Node Tokenizer for Integrating Graph Structure into Large Language ModelsabstractEnabling large language models (LLMs) to effectively process and reason with graph-structured data remains a significant challenge despite their remarkable success in natural language tasks. Current approaches either convert graph structures into verbose textual descriptions, consuming substantial computational resources, or employ complex graph neural networks as tokenizers, which introduce significant training overhead. To bridge this gap, we present NT-LLM, a novel framework with an anchor-based positional encoding scheme for graph representation. Our approach strategically selects reference nodes as anchors and encodes each node's position relative to these anchors, capturing essential topological information without the computational burden of existing methods. Notably, we identify and address a fundamental issue: the inherent misalignment between discrete hop-based distances in graphs and continuous distances in embedding spaces. By implementing a rank-preserving objective for positional encoding pretraining, NT-LLM achieves superior performance across diverse graph tasks ranging from basic structural analysis to complex reasoning scenarios. Our comprehensive evaluation demonstrates that this lightweight yet powerful approach effectively enhances LLMs' ability to understand and reason with graph-structured information, offering an efficient solution for graph-based applications of language models. Yanbiao Ji, Chang Liu 0078, Xin Chen 0077, Dan Luo 0004, Yue Ding 0001, Wenqing Lin, Hongtao Lu 0001 |
CIKM | 8 |
| 2025 | Few-shot Implicit Function Generation via EquivarianceabstractImplicit Neural Representations (INRs) have emerged as a powerful framework for representing continuous signals. However, generating diverse INR weights remains challenging due to limited training data. We introduce Few-Shot Implicit Function Generation, a new problem setup that aims to generate diverse yet functionally consistent INR weights from only a few examples. This is challenging because even for the same signal, the optimal INRs can vary significantly depending on their initializations. To tackle this, we propose EquiGen, a framework that can generate new INRs from limited data. The core idea is that functionally similar networks can be transformed into one another through weight permutations, forming an equivariance group. By projecting these weights into an equivariant latent space, we enable diverse generation within these groups, even with few examples. EquiGen implements this through an equivariant encoder trained via contrastive learning and smooth augmentation, an equivariance-guided diffusion process, and controlled perturbations in the equivariant subspace. Experiments on 2D image and 3D shape INR datasets demonstrate that our approach effectively generates diverse INR weights while preserving their functional properties in few-shot scenarios. Code is available at https://github.com/JeanDiable/EquiGen. Suizhi Huang, Xingyi Yang, Hongtao Lu 0001, Xinchao Wang |
CVPR | 3 |
| 2025 | Unsupervised Domain Adaptive Hand Mesh Reconstruction of 2D Images in the Wild
Xinyi Hou, Huayi Zhou 0001, Yue Ding 0001, Hongtao Lu 0001 |
ICANN (2) | 4 |
| 2025 | Discretized Gaussian Representation for Tomographic Reconstruction
Shaokai Wu, Yapan Guo, Suizhi Huang, Shalayiding Sirejiding, Qichen He, Jing Tong, Yanbiao Ji, Yue Ding 0001, Hongtao Lu 0001 |
ICCV | 12 |
| 2025 | BECAME: Bayesian Continual Learning with Adaptive Model MergingabstractContinual Learning (CL) strives to learn incrementally across tasks while mitigating catastrophic forgetting. A key challenge in CL is balancing stability (retaining prior knowledge) and plasticity (learning new tasks). While representative gradient projection methods ensure stability, they often limit plasticity. Model merging techniques offer promising solutions, but prior methods typically rely on empirical assumptions and carefully selected hyperparameters. In this paper, we explore the potential of model merging to enhance the stability-plasticity trade-off, providing theoretical insights that underscore its benefits. Specifically, we reformulate the merging mechanism using Bayesian continual learning principles and derive a closed-form solution for the optimal merging coefficient that adapts to the diverse characteristics of tasks. To validate our approach, we introduce a two-stage framework named BECAME, which synergizes the expertise of gradient projection and adaptive merging. Extensive experiments show that our approach outperforms state-of-the-art CL methods and existing merging strategies https://github.com/limei0818/BECAME. Qinyan Dai, Suizhi Huang, Yue Ding 0001, Hongtao Lu 0001 |
ICML | 6 |
| 2025 | Generating Negative Samples for Multi-Modal RecommendationabstractMulti-modal recommender systems (MMRS) have gained significant attention due to their ability to leverage information from various modalities to enhance recommendation quality. However, existing negative sampling techniques often struggle to effectively utilize the multi-modal data, leading to suboptimal performance. In this paper, we identify two key challenges in negative sampling for MMRS: (1) producing cohesive negative samples contrasting with positive samples and (2) maintaining a balanced influence across different modalities. To address these challenges, we propose NegGen, a novel framework that utilizes multi-modal large language models (MLLMs) to generate balanced and contrastive negative samples. We design three different prompt templates to enable NegGen to analyze and manipulate item attributes across multiple modalities, and then generate negative samples that introduce better supervision signals and ensure modality balance. Furthermore, NegGen employs a causal learning module to disentangle the effect of intervened key features and irrelevant item attributes, enabling fine-grained learning of user preferences. Extensive experiments on real-world datasets demonstrate the superior performance of NegGen compared to state-of-the-art methods in both negative sampling and multi-modal recommendation. Yanbiao Ji, Dan Luo 0004, Chang Liu 0078, Shaokai Wu, Jing Tong, Qichen He, Deyi Ji, Hongtao Lu 0001, Yue Ding 0001 |
ACM Multimedia | 8 |
| 2025 | CLIP-MT: Multi-Modal Knowledge-Driven Adaptive Scale Feature Allocation for Multi-Task Dense PredictionabstractRecent advancements in visual multi-task learning (MTL) have sparked significant interest. However, existing dense prediction MTL methods predominantly rely on single-modality image data, limiting their performance due to the absence of complementary knowledge from other modalities. Additionally, different dense tasks exhibit heterogeneous preferences during information decoding, posing a critical challenge in effectively allocating multi-scale encoded features. To address these limitations, we propose CLIP-MT, a Multi-Modal Knowledge-Driven Adaptive Scale Feature Allocation for Multi-Task Dense Prediction. Specifically, to enrich task-shared image features with multi-modal knowledge, we introduce a novel CLIP-Guided Global Feature Enhancer (CGGF), which leverages aligned text-image information to augment object-level representations through a dual-path feature fusion architecture. Furthermore, to tackle the task-specific scale preference problem, we design an Adaptive Scale Selection Gate (ASSG), a learnable gating mechanism that dynamically selects high- or low-scale features based on task-specific demands. Finally, we integrate multi-modal and multi-scale information through a Task-Aware Feature Fusion Module (TAFF). Extensive experiments on the NYUDv2 and PASCAL-Context datasets demonstrate that CLIP-MT achieves state-of-the-art performance, outperforming existing methods across multiple dense prediction tasks. Shalayiding Sirejiding, Yue Ding 0001, Xinyi Hou, Shaokai Wu, Qichen He, Hongtao Lu 0001 |
ACM Multimedia | 9 |
| 2025 | MORE: Multi-Organ Medical Image REconstruction DatasetabstractCT reconstruction provides radiologists with images for diagnosis and treatment, yet current deep learning methods are typically limited to specific anatomies and datasets, hindering generalization ability to unseen anatomies and lesions. To address this, we introduce the Multi-Organ medical image REconstruction (MORE) dataset, comprising CT scans across 9 diverse anatomies with 15 lesion types. This dataset serves two key purposes: (1) enabling robust training of deep learning models on extensive, heterogeneous data, and (2) facilitating rigorous evaluation of model generalization for CT reconstruction. We further establish a strong baseline solution that outperforms prior approaches under these challenging conditions. Our results demonstrate that: (1) a comprehensive dataset helps improve the generalization capability of models, and (2) optimization-based methods offer enhanced robustness for unseen anatomies. The MORE dataset is freely accessible under CC-BY-NC 4.0 at our project page https://more-med.github.io/. Shaokai Wu, Yapan Guo, Yanbiao Ji, Jing Tong, Suizhi Huang, Yue Ding 0001, Hongtao Lu 0001 |
ACM Multimedia | 9 |
| 2025 | How Does Topology Bias Distort Message Passing in Graph Recommender? A Dirichlet Energy PerspectiveabstractGraph-based recommender systems have achieved remarkable effectiveness by modeling high-order interactions between users and items. However, such approaches are significantly undermined by popularity bias, which distorts the interaction graph’s structure—referred to as topology bias. This leads to overrepresentation of popular items, thereby reinforcing biases and fairness issues through the user-system feedback loop. Despite attempts to study this effect, most prior work focuses on the embedding or gradient level bias, overlooking how topology bias fundamentally distorts the message passing process itself. We bridge this gap by providing an empirical and theoretical analysis from a Dirichlet energy perspective, revealing that graph message passing inherently amplifies topology bias and consistently benefits highly connected nodes. To address these limitations, we propose Test-time Simplicial Propagation (TSP), which extends message passing to higher-order simplicial complexes. By incorporating richer structures beyond pairwise connections, TSP mitigates harmful topology bias and substantially improves the representation and recommendation of long-tail items during inference. Extensive experiments across five real-world datasets demonstrate the superiority of our approach in mitigating topology bias and enhancing recommendation quality. The implementation code is available at https://github.com/sotaagi/TSP. Yanbiao Ji, Yue Ding 0001, Dan Luo 0004, Chang Liu 0078, Xin Xin 0003, Hongtao Lu 0001 |
NeurIPS | 7 |
| 2025 | Towards Personalized Federated Multi-Scenario Multi-Task RecommendationabstractIn modern recommender systems, especially in e-commerce, predicting multiple targets such as click-through rate (CTR) and post-view conversion rate (CTCVR) is common. Multi-task recommender systems are increasingly popular in both research and practice, as they leverage shared knowledge across diverse business scenarios to enhance performance. However, emerging real-world scenarios and data privacy concerns complicate the development of a unified multi-task recommendation model. Yue Ding 0001, Yanbiao Ji, Xin Xin 0003, Suizhi Huang, Chang Liu 0078, Xiaofeng Gao 0001, Tsuyoshi Murata, Hongtao Lu 0001 |
WSDM | 10 |
| 2025 | Structural and Statistical Texture Knowledge Distillation and Learning for SegmentationabstractLow-level texture feature/knowledge is also of vital importance for characterizing the local structural pattern and global statistical properties, such as boundary, smoothness, regularity, and color contrast, which may not be well addressed by high-level deep features. In this paper, we aim to re-emphasize the low-level texture information in deep networks for semantic segmentation and related knowledge distillation tasks. To this end, we take full advantage of both structural and statistical texture knowledge and propose a novel Structural and Statistical Texture Knowledge Distillation (SSTKD) framework for semantic segmentation. Specifically, Contourlet Decomposition Module (CDM) is introduced to decompose the low-level features with iterative Laplacian pyramid and directional filter bank to mine the structural texture knowledge, and Texture Intensity Equalization Module (TIEM) is designed to extract and enhance the statistical texture knowledge with the corresponding Quantization Congruence Loss (QDL). Moreover, we propose the Co-occurrence TIEM (C-TIEM) and generic segmentation frameworks, namely STLNet++ and U-SSNet, to enable existing segmentation networks to harvest the structural and statistical texture information more effectively. Extensive experimental results on three segmentation tasks demonstrate the effectiveness of the proposed methods and their state-of-the-art performance on seven popular benchmark datasets, respectively. Deyi Ji, Feng Zhao 0004, Hongtao Lu 0001, Feng Wu 0005, Jieping Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Learning Auxiliary Representations With Inconsistency-Guided Detail Regularization for Mask-Guided MattingabstractMask-guided matting networks have achieved significant improvements and have shown great potential in practical applications in recent years. However, simply learning matting representation from synthetic and lack-of-real-world-diversity matting data, these approaches tend to overfit low-level details in wrong regions, lack generalization to objects with complex structures and real-world scenes such as shadows, as well as suffer from interference of background lines or textures. To address these challenges, in this paper, we propose a novel auxiliary learning framework for mask-guided matting models, incorporating three auxiliary tasks: semantic segmentation, edge detection, and background line detection besides matting, to learn different and effective auxiliary representations from different types of data and annotations. Our framework and model introduce the following key aspects: 1) to learn real-world adaptive semantic representation for objects with diverse and complex structures under real-world scenes, we introduce extra semantic segmentation and edge detection tasks on more diverse real-world data with segmentation annotations; 2) to avoid overfitting on low-level details, we propose a module to utilize the inconsistency between learned segmentation and matting representations to regularize detail refinement; 3) we propose a novel background line detection task into our auxiliary learning framework, to suppress interference of background lines or textures. In addition, we propose a high-quality matting benchmark, Plant-Mat, to evaluate matting methods on complex structures. Extensively quantitative and qualitative results show that our approach outperforms state-of-the-art mask-guided methods. Zhaozhi Xie, Longjie Qi, Jingyong Cai, Hiroyuki Uchiyama, Yue Ding 0001, Hongtao Lu 0001 |
IEEE Trans. Multim. | 9 |
| 2024 | Fedhca2: Towards Hetero-Client Federated Multi-Task LearningabstractFederated Learning (FL) enables joint training across distributed clients using their local data privately. Federated Multi-Task Learning (FMTL) builds on FL to handle multiple tasks, assuming model congruity that identical model architecture is deployed in each client. To relax this assumption and thus extend real-world applicability, we introduce a novel problem setting, Hetero-Client Fed-erated Multi-Task Learning (HC-FMTL), to accommodate diverse task setups. The main challenge of HC-FMTL is the model incongruity issue that invalidates conventional aggregation methods. It also escalates the difficulties in model aggregation to deal with data and task heterogeneity inherent in FMTL. To address these challenges, we pro-pose the$FedHCA^{2}$framework, which allows for federated training of personalized models by modeling relationships among heterogeneous clients. Drawing on our theoretical insights into the difference between multi-task and federated optimization, we propose the Hyper Conflict-Averse Aggregation scheme to mitigate conflicts during encoder updates. Additionally, inspired by task interaction in MTL, the Hyper Cross Attention Aggregation scheme uses layer-wise cross attention to enhance decoder interactions while alleviating model incongruity. Moreover, we employ learnable Hyper Aggregation Weights for each client to customize personalized parameter updates. Extensive experiments demon-strate the superior performance of$FedHCA^{2}$in various HC-FMTL scenarios compared to representative methods. Code is available at https://github.com/innovator-zero/FedHCA2. Suizhi Huang, Yuwen Yang, Shalayiding Sirejiding, Yue Ding 0001, Hongtao Lu 0001 |
CVPR | 6 |
| 2024 | Towards Mixture of Task-Intensive Experts for Multi-task Recommendation
Hongtao Lu 0001, Yue Ding 0001 |
DASFAA (3) | 3 |
| 2024 | Mediate: Mixture Domain Model-Agnostic Federated Learning
Chang Liu 0078, Yuwen Yang, Yue Ding 0001, Hongtao Lu 0001 |
DASFAA (1) | 5 |
| 2024 | YOLO-Med : Multi-Task Interaction Network for Biomedical ImagesabstractObject detection and semantic segmentation are pivotal components in biomedical image analysis. Current single-task networks exhibit promising outcomes in both detection and segmentation tasks. Multi-task networks have gained prominence due to their capability to simultaneously tackle segmentation and detection tasks, while also accelerating the segmentation inference. Nevertheless, recent multi-task networks confront distinct limitations such as the difficulty in striking a balance between accuracy and inference speed. Additionally, they often overlook the integration of cross-scale features, which is especially important for biomedical image analysis. In this study, we propose an efficient end-to-end multi-task network capable of concurrently performing object detection and semantic segmentation called YOLO-Med. Our model employs a backbone and a neck for multi-scale feature extraction, complemented by the inclusion of two task-specific decoders. A cross-scale task-interaction module is employed in order to facilitate information fusion between various tasks. Our model exhibits promising results in balancing accuracy and speed when evaluated on the Kvasir-seg dataset and a private biomedical image dataset. Suizhi Huang, Shalayiding Sirejiding, Yue Ding 0001, Leheng Liu, Hongtao Lu 0001 |
ICASSP | 7 |
| 2024 | Changenet: Multi-Temporal Asymmetric Change Detection DatasetabstractChange Detection (CD) has been attracting extensive interests with the availability of bi-temporal datasets. However, due to the huge cost of multi-temporal images acquisition and labeling, existing change detection datasets are small in quantity, short in temporal, and low in practicability. Therefore, a large-scale practical-oriented dataset covering wide temporal phases is urgently needed to facilitate the community. To this end, the ChangeNet dataset is presented especially for multi-temporal change detection, along with the new task of "Asymmetric Change Detection". Specifically, ChangeNet consists of 31,000 multi-temporal images pairs, a wide range of complex scenes from 100 cities, and 6 pixel-level annotated categories, which is far superior to all the existing change detection datasets including LEVIR-CD, WHU Building CD, etc.. In addition, ChangeNet contains amounts of real-world perspective distortions in different temporal phases on the same areas, which is able to promote the practical application of change detection algorithms. The ChangeNet dataset is suitable for both binary change detection (BCD) and semantic change detection (SCD) tasks. Accordingly, we benchmark the ChangeNet dataset on six BCD methods and two SCD methods, and extensive experiments demonstrate its challenges and great significance. The dataset is available at https://github.com/jankyee/ChangeNet. Deyi Ji, Mingyuan Tao, Hongtao Lu 0001, Feng Zhao 0004 |
ICASSP | 4 |
| 2024 | Task Indicating Transformer for Task-Conditional Dense PredictionsabstractThe task-conditional model is a distinctive stream for efficient multi-task learning. Existing works encounter a critical limitation in learning task-agnostic and task-specific representations, primarily due to shortcomings in global context modeling arising from CNN-based architectures, as well as a deficiency in multi-scale feature interaction within the decoder. In this paper, we introduce a novel task-conditional framework called Task Indicating Transformer (TIT) to tackle this challenge. Our approach designs a Mix Task Adapter module within the transformer block, which incorporates a Task Indicating Matrix through matrix decomposition, thereby enhancing long-range dependency modeling and parameter-efficient feature adaptation by capturing intra- and inter-task features. Moreover, we propose a Task Gate Decoder module that harnesses a Task Indicating Vector and gating mechanism to facilitate adaptive multi-scale feature refinement guided by task embeddings. Experiments on two public multi-task dense prediction benchmarks, NYUD-v2 and PASCAL-Context, demonstrate that our approach surpasses state-of-the-art task-conditional methods. Shalayiding Sirejiding, Bayram Bayramli, Suizhi Huang, Yue Ding 0001, Hongtao Lu 0001 |
ICASSP | 6 |
| 2024 | UNIDEAL: Curriculum Knowledge Distillation Federated LearningabstractFederated Learning (FL) has emerged as a promising approach to enable collaborative learning among multiple clients while preserving data privacy. However, cross-domain FL tasks, where clients possess data from different domains or distributions, remain a challenging problem due to the inherent heterogeneity. In this paper, we present UNIDEAL, a novel FL algorithm specifically designed to tackle the challenges of cross-domain scenarios and heterogeneous model architectures. The proposed method introduces Adjustable Teacher-Student Mutual Evaluation Curriculum Learning, which significantly enhances the effectiveness of knowledge distillation in FL settings. We conduct extensive experiments on various datasets, comparing UNIDEAL with state-of-the-art baselines. Our results demonstrate that UNIDEAL achieves superior performance in terms of both model accuracy and communication efficiency. Additionally, we provide a convergence analysis of the algorithm, showing a convergence rate of $O\left( {\frac{1}{T}} \right)$ under non-convex conditions. Yuwen Yang, Chang Liu 0078, Suizhi Huang, Hongtao Lu 0001, Yue Ding 0001 |
ICASSP | 5 |
| 2024 | Enhancing Cross-Domain Detection: Adaptive Class-Aware Contrastive TransformerabstractRecently, the detection transformer has gained substantial attention for its inherent minimal post-processing requirement. However, this paradigm relies on abundant training data, yet in the context of the cross-domain adaptation, insufficient labels in the target domain exacerbate issues of class imbalance and model performance degradation. To address these challenges, we propose a novel class-aware cross domain detection transformer based on the adversarial learning and mean-teacher framework. First, considering the inconsistencies between the classification and regression tasks, we introduce an IoU-aware prediction branch and exploit the consistency of classification and location scores to filter and reweight pseudo labels. Second, we devise a dynamic category threshold refinement to adaptively manage model confidence. Third, to alleviate the class imbalance, an instance-level class-aware contrastive learning module is presented to encourage the generation of discriminative features for each class, particularly benefiting minority classes. Experimental results across diverse domain-adaptive scenarios validate our method’s effectiveness in improving performance and alleviating class imbalance issues, which outperforms the state-of-the-art transformer based methods. Ziru Zeng, Yue Ding 0001, Hongtao Lu 0001 |
ICASSP | 3 |
| 2024 | CGCUT: Unpaired Image-to-Image Translation via Cluster-Guided Contrastive LearningabstractMaintaining content consistency is a critical point in unpaired image-to-image translation. Contrastive learning is an essential framework in unpaired image-to-image translation to maximize the mutual information between corresponding patches of input and generated images. However, ignoring the importance of background information, the previous CL-based methods suffer from the problems that the background of generated images, as well as the regions where the object and background intersect, appear to be blurred. To tackle these problems, we propose cluster-guided contrastive learning in the I2I task. We leverage the cluster information of features to mine hard negative samples and distinguish the concept and the detail without additional parameters. We further provide rigorous proof of mutual information of our proposed method, which proves that our method preserves the lower bound of previous work and has a lower variance. Our proposed method, CGCUT, achieves state-of-the-art performance on most metrics on three benchmark datasets. Longjie Qi, Yue Ding 0001, Hongtao Lu 0001 |
ICME | 3 |
| 2024 | PA-SAM: Prompt Adapter SAM for High-Quality Image SegmentationabstractThe Segment Anything Model (SAM) has exhibited outstanding performance in various image segmentation tasks. Despite being trained with over a billion masks, SAM faces challenges in mask prediction quality in numerous scenarios, especially in real-world contexts. In this paper, we introduce a novel prompt-driven adapter into SAM, namely Prompt Adapter Segment Anything Model (PA-SAM), aiming to enhance the segmentation mask quality of the original SAM. By exclusively training the prompt adapter, PA-SAM extracts detailed information from images and optimizes the mask decoder feature at both sparse and dense prompt levels, improving the segmentation performance of SAM to produce high-quality masks. Experimental results demonstrate that our PA-SAM outperforms other SAM-based methods in high-quality, zero-shot, and open-set segmentation. We’re making the source code and models available at https://github.com/xzz2/pa-sam. Zhaozhi Xie, Bochen Guan, Muyang Yi, Yue Ding 0001, Hongtao Lu 0001, Lei Zhang 0006 |
ICME | 6 |
| 2024 | BARTENDER: A simple baseline model for task-level heterogeneous federated learningabstractThis study presents the Task-level Heterogeneous Federated Learning (TH-FL), a novel paradigm that fuses the principles of Federated Learning (FL) and Multi-Task Learning (MTL). In the TH-FL scenario, each client can learn an indefinite number of tasks, which may vary in type and originate from distinct domains. We introduce a unique baseline model, BARTENDER, that integrates a Conditional Prompt (CP) module. This module encodes task-specific and domain-specific information, enabling the model to generate tailored outputs based on the encoding inputs. This innovative strategy not only minimizes the communication costs associated with FL but also enhances model generalization across a variety of task types. Through extensive experiments, we establish that the BARTENDER model surpasses traditional multi-decoder architecture models across diverse scenarios. We also explore the influence of the parameter decoupling strategy on model training and outline the assumptions necessary for achieving a $O\left( {1/\sqrt T } \right)$ convergence speed in the TH-FL scenario. Yuwen Yang, Suizhi Huang, Shalayiding Sirejiding, Chang Liu 0078, Muyang Yi, Zhaozhi Xie, Yue Ding 0001, Hongtao Lu 0001 |
ICME | 9 |
| 2024 | Discrete Latent Perspective Learning for Segmentation and DetectionabstractIn this paper, we address the challenge of Perspective-Invariant Learning in machine learning and computer vision, which involves enabling a network to understand images from varying perspectives to achieve consistent semantic interpretation. While standard approaches rely on the labor-intensive collection of multi-view images or limited data augmentation techniques, we propose a novel framework, Discrete Latent Perspective Learning (DLPL), for latent multi-perspective fusion learning using conventional single-view images. DLPL comprises three main modules: Perspective Discrete Decomposition (PDD), Perspective Homography Transformation (PHT), and Perspective Invariant Attention (PIA), which work together to discretize visual features, transform perspectives, and fuse multi-perspective semantic information, respectively. DLPL is a universal perspective learning framework applicable to a variety of scenarios and vision tasks. Extensive experiments demonstrate that DLPL significantly enhances the network’s capacity to depict images across diverse scenarios (daily photos, UAV, auto-driving) and tasks (detection, segmentation). Deyi Ji, Feng Zhao 0004, Lanyun Zhu, Wenwei Jin, Hongtao Lu 0001, Jieping Ye |
ICML | 5 |
| 2024 | PPTFormer: Pseudo Multi-Perspective Transformer for UAV Segmentation
Deyi Ji, Wenwei Jin, Hongtao Lu 0001, Feng Zhao 0004 |
IJCAI | 3 |
| 2024 | Integrating Semantic Segmentation Model for Self-Supervised Scene Flow Estimation via Cross Task DistillationabstractThis paper introduces a self-supervised learning approach for monocular scene flow estimation, addressing challenges posed by the reliance on expensive 3D sensing technologies such as Lidar and RGB-D cameras and the scarcity of datasets with ground truth labels. Our method utilizes cross-task distillation, where a semantic segmentation network serves as a teacher to impart valuable information to the scene flow network. To facilitate effective information exchange between tasks with inherent differences, we incorporate self-attention blocks within each network. Specifically, we transfer self-attention weights from the semantic segmentation network to the scene flow network, aligning the attention probabilities of both networks. The integration of self-attention mechanisms enhances the adaptability of our framework to complex scene structures, contributing to robust and accurate scene flow estimation. Quantitative and qualitative experiments validate the efficacy of our approach, demonstrating a significant reduction in the error metric from 30.97% to 28.50%, representing an approximate 8% improvement compared to the best-performing existing self-supervised scene flow method. Bayram Bayramli, Yue Ding 0001, Hongtao Lu 0001 |
IJCNN | 3 |
| 2024 | Beyond Binary Preference: Leveraging Bayesian Approaches for Joint Optimization of Ranking and CalibrationabstractPredicting click-through rate (CTR) is a critical task in recommendation systems, where the models are optimized with pointwise loss to infer the probability of items being clicked. In industrial practice, applications also require ranking items based on these probabilities. Existing solutions primarily combine the ranking-based loss, i.e., pairwise and listwise loss, with CTR prediction. However, they can hardly calibrate or generalize well in CTR scenarios where the clicks reflect the binary preference. This is because the binary click feedback leads to a large number of ties, which renders high data sparsity. In this paper, we propose an effective data augmentation strategy, named Beyond Binary Preference (BBP) training framework, to address this problem. Our key idea is to break the ties by leveraging Bayesian approaches, where the beta distribution models click behavior as probability distributions in the training data that naturally break ties. Therefore, we can obtain an auxiliary training label that generates more comparable pairs and improves the ranking performance. Besides, BBP formulates ranking and calibration as a multi-task framework to optimize both objectives simultaneously. Through extensive offline experiments and online tests on various datasets, we demonstrate that BBP significantly outperforms state-of-the-art methods in both ranking and calibration capabilities, showcasing its effectiveness in addressing the limitations of existing methods. Our code is available at https://github.com/AlvinIsonomia/BBP. Chang Liu 0078, Wenqing Lin, Yue Ding 0001, Hongtao Lu 0001 |
KDD | 5 |
| 2024 | DAG: Deep Adaptive and Generative K-Free Community Detection on Attributed GraphsabstractCommunity detection on attributed graphs with rich semantic and topological information offers great potential for real-world network analysis, especially user matching in online games. Graph Neural Networks (GNNs) have recently enabled Deep Graph Clustering (DGC) methods to learn cluster assignments from semantic and topological information. However, their success depends on the prior knowledge related to the number of communities K, which is unrealistic due to the high costs and privacy issues of acquisition. In this paper, we investigate the community detection problem without prior K, referred to as K-Free Community Detection problem. To address this problem, we propose a novel Deep Adaptive and Generative model~(DAG) for community detection without specifying the prior K. DAG consists of three key components, i.e., a node representation learning module with masked attribute reconstruction, a community affiliation readout module, and a community number search module with group sparsity. These components enable DAG to convert the process of non-differentiable grid search for the community number, i.e., a discrete hyperparameter in existing DGC methods, into a differentiable learning process. In such a way, DAG can simultaneously perform community detection and community number search end-to-end. To alleviate the cost of acquiring community labels in real-world applications, we design a new metric, EDGE, to evaluate community detection methods even when the labels are not feasible. Extensive offline experiments on five public datasets and a real-world online mobile game dataset demonstrate the superiority of our DAG over the existing state-of-the-art (SOTA) methods. DAG has a relative increase of 7.35% in teams in a Tencent online game compared with the best competitor. Chang Liu 0078, Yuwen Yang, Yue Ding 0001, Hongtao Lu 0001, Wenqing Lin, Ziming Wu, Wendong Bi |
KDD | 4 |
| 2024 | Federated Multi-Task Learning on Non-IID Data Silos: An Experimental StudyabstractThe innovative Federated Multi-Task Learning (FMTL) approach consolidates the benefits of Federated Learning (FL) and Multi-Task Learning (MTL), enabling collaborative model training on multi-task learning datasets. However, a comprehensive evaluation method, integrating the unique features of both FL and MTL, is currently absent in the field. This paper fills this void by introducing a novel framework, FMTL-Bench, for systematic evaluation of the FMTL paradigm. This benchmark covers various aspects at the data, model, and optimization algorithm levels, and comprises seven sets of comparative experiments, encapsulating a wide array of non-independent and identically distributed (Non-IID) data partitioning scenarios. We propose a systematic process for comparing baselines of diverse indicators and conduct a case study on communication expenditure, time, and energy consumption. Through our exhaustive experiments, we aim to provide valuable insights into the strengths and limitations of existing baseline methods, contributing to the ongoing discourse on optimal FMTL application in practical scenarios. The source code can be found at https://github.com/youngfish42/FMTL-Benchmark. Yuwen Yang, Suizhi Huang, Shalayiding Sirejiding, Hongtao Lu 0001, Yue Ding 0001 |
ICMR | 5 |
| 2024 | Task-Interaction-Free Multi-Task Learning with Efficient Hierarchical Feature RepresentationabstractTraditional multi-task learning often relies on explicit task interaction mechanisms to enhance multi-task performance. However, these approaches encounter challenges such as negative transfer when jointly learning multiple weakly correlated tasks. Additionally, these methods handle encoded features at a large scale, which escalates computational complexity to ensure dense prediction task performance. In this study, we introduce a Task-Interaction-Free Network (TIF) for multi-task learning, which diverges from explicitly designed task interaction mechanisms. Firstly, we present a Scale Attentive-Feature Fusion Module (SAFF) to enhance each scale in the shared encoder to have rich task-agnostic encoded features. Subsequently, our proposed task and scale-specific decoders efficiently decode the enhanced features shared across tasks without necessitating task-interaction modules. Concretely, we utilize a Self-Feature Distillation Module (SFD) to explore task-specific features at lower scales and the Low-To-High Scale Feature Diffusion Module (LTHD) to diffuse global pixel relationships from low-level to high-level scales. Experiments on publicly available multi-task learning datasets validate that our TIF attains state-of-the-art performance. Shalayiding Sirejiding, Bayram Bayramli, Yuwen Yang, Tamam Alsarhan, Hongtao Lu 0001, Yue Ding 0001 |
ACM Multimedia | 6 |
| 2024 | MMAT: Multi-scale Multi-attention Transformer for Fine-Grained Wild Fungi Visual Classification
Qinyan Dai, Hongtao Lu 0001 |
PRICAI (3) | 4 |
| 2024 | TFUT: Task fusion upward transformer model for multi-task learning on dense prediction
Zewei Xin, Shalayiding Sirejiding, Yue Ding 0001, Tamam Alsarhan, Hongtao Lu 0001 |
Comput. Vis. Image Underst. | 7 |
| 2024 | T-Skeleton: Accurate scene text detection via instance-aware skeleton embeddingabstractAbstract Existing segmentation‐based methods have made considerable progress in arbitrarily shaped text detection due to the advantage of dealing with shape variation. However, there still exist challenges to detecting accurate text instances with dense layouts, inaccurate annotations, and complex backgrounds. Many recent works have focused on improving arbitrary boundary prediction, but it may be difficult to accurately distinguish each instance of dense layouts because their boundary pixels may be mistakenly classified to produce inaccurate results (i.e., adhesive texts) with inaccurate annotation and complex backgrounds. Considering the local and long‐range dependencies, this paper proposes an efficient text detector, namely T‐Skeleton, to obtain more reliable segmentation detections. In the spirit of object skeletonization, we introduce the text instance skeleton highlighting the semantically significant structure (similar to the skeleton of a fish) to explicitly capture the long‐range dependencies of text instances. The key idea of T‐Skeleton is to calibrate the coarse text proposals by embedding text instance skeletons to separate crowd texts accurately and robustly. We further design a channel attention module to enlarge the performance margin between T‐Skeleton and the segmentation baseline. Experimental results on four publicly available datasets show the superiority of T‐Skeleton in handling long and curved texts. Xingfei Hu, Hongtao Lu 0001 |
IET Image Process. | 3 |
| 2024 | BPJDet: Extended Object Representation for Generic Body-Part Joint DetectionabstractDetection of human body and its parts has been intensively studied. However, most of CNNs-based detectors are trained independently, making it difficult to associate detected parts with body. In this paper, we focus on the joint detection of human body and its parts. Specifically, we propose a novel extended object representation integrating center-offsets of body parts, and construct an end-to-end generic Body-Part Joint Detector (BPJDet). In this way, body-part associations are neatly embedded in a unified representation containing both semantic and geometric contents. Therefore, we can optimize multi-loss to tackle multi-tasks synergistically. Moreover, this representation is suitable for anchor-based and anchor-free detectors. BPJDet does not suffer from error-prone post matching, and keeps a better trade-off between speed and accuracy. Furthermore, BPJDet can be generalized to detect body-part or body-parts of either human or quadruped animals. To verify the superiority of BPJDet, we conduct experiments on datasets of body-part (CityPersons, CrowdHuman and BodyHands) and body-parts (COCOHumanParts and Animals5C). While keeping high detection accuracy, BPJDet achieves state-of-the-art association performance on all datasets. Besides, we show benefits of advanced body-part association capability by improving performance of two representative downstream applications: accurate crowd head detection and hand contact estimation. Huayi Zhou 0001, Fei Jiang 0006, Jiaxin Si, Yue Ding 0001, Hongtao Lu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Spammer detection on short video applications
Muyang Yi, Yue Ding 0001, Hongtao Lu 0001 |
Pattern Recognit. Lett. | 5 |
| 2024 | Superpixel Guided Network for Weakly Supervised Semantic SegmentationabstractImage-level weakly supervised semantic segmentation faces challenges in accurately capturing boundaries and representing intricate details due to the absence of pixel-level supervision. Constrained by the enormous number of pixels, pixel-level propagation has difficulty in capturing the long-range dependency, particularly in small, isolated regions. To this end, we introduce a novel approach of self-supervised segmentation integrated with superpixel, and develop a network called superpixel guided network (SPGNet) to simultaneously perform superpixel generation and segmentation mask prediction. Significantly, our framework facilitates mutual supervised learning between the segmentation branch and the superpixel branch. The superpixel guides the predicted mask for improved boundary location, while the latter provides supervision on superpixel through superpixel center generation (SCG) and union boundary extraction (UBE). Furthermore, we propose superpixel context fusion (SCF) to generate compact pseudo masks and capture long-range dependency. Experimental results demonstrate that the proposed SPGNet achieves outstanding performance on the PASCAL VOC 2012 segmentation benchmark Zhaozhi Xie, Yuwen Yang, Hongtao Lu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Adaptive Task-Wise Message Passing for Multi-Task Learning: A Spatial Interaction PerspectiveabstractRecent advancements have facilitated the simultaneous processing of multiple dense prediction tasks, utilizing diverse correlations between these tasks. However, many of these advances predominantly focus on a singular or fixed task interaction, leading to negative transfer effects. In this paper, we introduce an end-to-end model called the Adaptive Task-Wise Message Passing Network (ATMPNet) for multi-task learning. Our proposed model focuses on excavating comprehensive spatial messages among tasks in an adaptive manner. To achieve this, ATMPNet incorporates the Adaptive Spatial Message Interaction (ASMI) module, which models various local spatial message interactions and global interactions among tasks. ASMI explores potential spatial relationships by generating a task-specific message pool for each target task. Furthermore, we propose an Adaptive Task Message Passing (ATMP) module, a novel method for aggregating messages. The ATMP module generates refined global-local messages from each message pool and adaptively transfers them to the corresponding target tasks through a well-designed message passing scheme. We conduct extensive experiments on the NYUD-v2 and PASCAL-Context datasets to evaluate the effectiveness of ATMPNet. The results demonstrate the state-of-the-art performance of our proposed model in handling multi-task learning scenarios. Code will be publicly available in here. Shalayiding Sirejiding, Bayram Bayramli, Suizhi Huang, Hongtao Lu 0001, Yue Ding 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Prompt Guided Transformer for Multi-Task Dense PredictionabstractTask-conditional architecture offers advantage in parameter efficiency but falls short in performance compared to state-of-the-art multi-decoder methods. How to trade off performance and model parameters is an important and difficult problem. In this paper, we introduce a simple and lightweight task-conditional model called Prompt Guided Transformer (PGT) to optimize this challenge. Our approach designs a Prompt-conditioned Transformer block, which incorporates task-specific prompts in the self-attention mechanism to achieve global dependency modeling and parameter-efficient feature adaptation across multiple tasks. This block is integrated into both the shared encoder and decoder, enhancing the capture of intra- and inter-task features. Moreover, we design a lightweight decoder to further reduce parameter usage, which accounts for only 2.7% of the total model parameters. Extensive experiments on two multi-task dense prediction benchmarks, PASCAL-Context and NYUD-v2, demonstrate that our approach achieves state-of-the-art results among task-conditional methods while using fewer parameters, and maintains a significant balance between performance and parameter size. Shalayiding Sirejiding, Yue Ding 0001, Hongtao Lu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Ultra-High Resolution Segmentation with Ultra-Rich Context: A Novel BenchmarkabstractWith the increasing interest and rapid development of methods for Ultra-High Resolution (UHR) segmentation, a large-scale benchmark covering a wide range of scenes with full fine-grained dense annotations is urgently needed to facilitate the field. To this end, the URUR dataset is introduced, in the meaning of Ultra-High Resolution dataset with Ultra-Rich Context. As the name suggests, URUR contains amounts of images with high enough resolution (3,008 images of size 5,120 × 5,120), a wide range of complex scenes (from 63 cities), rich-enough context (1 million instances with 8 categories) and fine-grained annotations (about 80 billion manually annotated pixels), which is far superior to all the existing UHR datasets including DeepGlobe, Inria Aerial, UDD, etc.. Moreover, we also propose WSDNet, a more efficient and effective framework for UHR segmentation especially with ultra-rich context. Specifically, multi-level Discrete Wavelet Transform (DWT) is naturally integrated to release computation burden while preserve more spatial details, along with a Wavelet Smooth Loss (WSL) to reconstruct original structured context and texture with a smooth constrain. Experiments on several UHR datasets demonstrate its state-of-the-art performance. The dataset is available at https://github.com/jankyee/URUR. Deyi Ji, Feng Zhao 0004, Hongtao Lu 0001, Mingyuan Tao, Jieping Ye |
CVPR | 3 |
| 2023 | A Simple Framework for Text-Supervised Semantic SegmentationabstractText-supervised semantic segmentation is a novel research topic that allows semantic segments to emerge with image-text contrasting. However, pioneering methods could be subject to specifically designed network architectures. This paper shows that a vanilla contrastive language-image pretraining (CLIP) model is an effective text-supervised semantic segmentor by itself. First, we reveal that a vanilla CLIP is inferior to localization and segmentation due to its optimization being driven by densely aligning visual and language representations. Second, we propose the locality-driven alignment (LoDA) to address the problem, where CLIP optimization is driven by sparsely aligning local representations. Third, we propose a simple segmentation (SimSeg) framework. LoDA and SimSeg jointly amelio-rate a vanilla CLIP to produce impressive semantic segmentation results. Our method outperforms previous state-of-the-art methods on PASCAL VOC 2012, PASCAL Context and COCO datasets by large margins. Code and models are available at github.com/muyangyi/SimSeg. Muyang Yi, Quan Cui, Osamu Yoshie, Hongtao Lu 0001 |
CVPR | 6 |
| 2023 | Generating Human Motion from Textual Descriptions with Discrete RepresentationsabstractIn this work, we investigate a simple and must-known conditional generative framework based on Vector Quantised-Variational AutoEncoder (VQ-VAE) and Generative Pre-trained Transformer (GPT) for human motion generation from textural descriptions. We show that a simple CNN-based VQ-VAE with commonly used training recipes (EMA and Code Reset) allows us to obtain high-quality discrete representations. For GPT, we incorporate a simple corruption strategy during the training to alleviate training-testing discrepancy. Despite its simplicity, our T2M-GPT shows better performance than competitive approaches, including recent diffusion-based approaches. For example, on HumanML3D, which is currently the largest dataset, we achieve comparable performance on the consistency between text and generated motion (R-Precision), but with FID 0.116 largely outperforming MotionDiffuse of 0.630. Additionally, we conduct analyses on HumanML3D and observe that the dataset size is a limitation of our approach. Our work suggests that VQ-VAE still remains a competitive approach for human motion generation. Our implementation is available on the project page: https://mael-zys.github.io/T2M-GPT/. Jianrong Zhang, Yangsong Zhang 0002, Xiaodong Cun, Yong Zhang 0034, Hongtao Lu 0001, Xi Shen 0001, Shan Ying |
CVPR | 6 |
| 2023 | Spammer Detection on Short Video Applications: A new Challenge and BaselinesabstractUsers can interact with the advertisements and share their impressions through the review system on short video applications. However, spammers may post false or malicious comments to mislead normal users due to profit-driven reasons, damaging the community’s positive atmosphere. In this paper, we introduce a new challenge of spammer detection on short video applications, where the multi-modal information of videos and reviews plays a more critical role than the spam relation graph. Then we propose SPAM-3, a novel baseline to detect SPAM reviews with Multi-Modal representation using attentive heterogeneous graph convolution. Our approach balances multi-modal representation fusion and graph relation extraction, enabling fine-grained interaction and generating discriminative features for the spammer classification task. We demonstrate SPAM-3 via experiments and analysis. Muyang Yi, Yue Ding 0001, Hongtao Lu 0001 |
ICASSP | 5 |
| 2023 | Stuart: Individualized Classroom Observation of Students with Automatic Behavior Recognition And TrackingabstractEach student matters, but it is hardly for instructors to observe all the students during the courses and provide helps to the needed ones immediately. In this paper, we present StuArt, a novel automatic system designed for the individualized classroom observation, which empowers instructors to concern the learning status of each student. StuArt can recognize five representative student behaviors (hand-raising, standing, sleeping, yawning, and smiling) that are highly related to the engagement and track their variation trends during the course. To protect the privacy of students, all the variation trends are indexed by the seat numbers without any personal identification information. Furthermore, StuArt adopts various user-friendly visualization designs to help instructors quickly understand the individual and whole learning status. Experimental results on real classroom videos have demonstrated the superiority and robustness of the embedded algorithms. We expect our system promoting the development of large-scale individualized guidance of students. More information is in https://github.com/hnuzhy/StuArt. Huayi Zhou 0001, Fei Jiang 0006, Jiaxin Si, Lili Xiong, Hongtao Lu 0001 |
ICASSP | 5 |
| 2023 | Scale-Aware Task Message Transferring for Multi-Task LearningabstractExploring cross-task interaction has been the mainstream in recent multi-task learning for dense predictions. However, existing works that focus on excavating cross-task contextual information are briefly based on hierarchical all-scale features, ignoring the complementary information from task relationships in multi-scale features. Meanwhile, recent research advances rarely pay attention to efficient information transformation between tasks. In this paper, we propose a novel end-to-end multi-task learning network termed as Scale-Aware Task Message Transferring Network (SATMTN) to explore the task relationships in multi-scale features to enrich the contextual cross-task information in the all-scale features. Specifically, we design Multi-Scale Message Passing Decoder (MSMP) to model interaction information between tasks in the multi-scale features, which is characterized by stacked bi-directional fully connected graph neural networks. Further, we devise another All-Scale Message Passing decoder (ASMP) to extract rich contextual information between tasks in the all-scale features, which is based on adaptive message transformation in the graph. Extensive experiments are implemented to show that our method surpasses the current state-of-the-art works on public multi-task benchmarks. Shalayiding Sirejiding, Hongtao Lu 0001, Yue Ding 0001 |
ICME | 3 |
| 2023 | Body-Part Joint Detection and Association via Extended Object RepresentationabstractThe detection of human body and its related parts (e.g., face, head or hands) have been intensively studied and greatly improved since the breakthrough of deep CNNs. However, most of these detectors are trained independently, making it a challenging task to associate detected body parts with people. This paper focuses on the problem of joint detection of human body and its corresponding parts. Specifically, we propose a novel extended object representation that integrates the center location offsets of body or its parts, and construct a dense single-stage anchor-based Body-Part Joint Detector (BPJDet). Body-part associations in BPJDet are embedded into the unified representation which contains both the semantic and geometric information. Therefore, BPJDet does not suffer from error-prone association post-matching, and has a better accuracy-speed trade-off. Furthermore, BPJDet can be seamlessly generalized to jointly detect any body part. To verify the effectiveness and superiority of our method, we conduct extensive experiments on the CityPersons, CrowdHuman and BodyHands datasets. The proposed BPJDet detector achieves state-of-the-art association performance on these three benchmarks while maintains high accuracy of detection. Code is released in https://github.com/hnuzhy/BPJDet. Huayi Zhou 0001, Fei Jiang 0006, Hongtao Lu 0001 |
ICME | 3 |
| 2023 | Landmark-Assisted Facial Action Unit Detection with Optimal Attention and Contrastive Learning
Qiaoping Hu, Hongtao Lu 0001, Fei Jiang 0006, Yaoyi Li |
ICONIP (12) | 3 |
| 2023 | Multi-scale Local Region-Based Facial Action Unit Detection with Graph Convolutional Network
Zhenchang Zhang, Hongtao Lu 0001, Fei Jiang 0006 |
ICONIP (12) | 3 |
| 2023 | Guided Patch-Grouping Wavelet Transformer with Spatial Congruence for Ultra-High Resolution SegmentationabstractMost existing ultra-high resolution (UHR) segmentation methods always struggle in the dilemma of balancing memory cost and local characterization accuracy, which are both taken into account in our proposed Guided Patch-Grouping Wavelet Transformer (GPWFormer) that achieves impressive performances. In this work, GPWFormer is a Transformer (T)-CNN (C) mutual leaning framework, where T takes the whole UHR image as input and harvests both local details and fine-grained long-range contextual dependencies, while C takes downsampled image as input for learning the category-wise deep context. For the sake of high inference speed and low computation complexity, T partitions the original UHR image into patches and groups them dynamically, then learns the low-level local details with the lightweight multi-head Wavelet Transformer (WFormer) network. Meanwhile, the fine-grained long-range contextual dependencies are also captured during this process, since patches that are far away in the spatial domain can also be assigned to the same group. In addition, masks produced by C are utilized to guide the patch grouping process, providing a heuristics decision. Moreover, the congruence constraints between the two branches are also exploited to maintain the spatial consistency among the patches. Overall, we stack the multi-stage process in a pyramid way. Experiments show that GPWFormer outperforms the existing methods with significant improvements on five benchmark datasets. Deyi Ji, Feng Zhao 0004, Hongtao Lu 0001 |
IJCAI | 3 |
| 2023 | Cooperative Self-Training for Multi-Target Adaptive Semantic SegmentationabstractIn this work we address multi-target domain adaptation (MTDA) in semantic segmentation, which consists in adapting a single model from an annotated source dataset to multiple unannotated target datasets that differ in their underlying data distributions. To address MTDA, we propose a self-training strategy that employs pseudo-labels to induce cooperation among multiple domain-specific classifiers. We employ feature stylization as an efficient way to generate image views that forms an integral part of self-training. Additionally, to prevent the network from overfitting to noisy pseudo-labels, we devise a rectification strategy that leverages the predictions from different classifiers to estimate the quality of pseudo-labels. Our extensive experiments on numerous settings, based on four different semantic segmentation datasets, validates the effectiveness of the proposed self-training strategy and shows that our method outperforms state-of-the-art MTDA approaches. https://github.com/Mael-zys/CoaST. Yangsong Zhang 0002, Subhankar Roy, Hongtao Lu 0001, Elisa Ricci 0001, Stéphane Lathuilière |
WACV | 3 |
| 2023 | Position-Aware Subgraph Neural Networks with Data-Efficient LearningabstractData-efficient learning on graphs (GEL) is essential in real-world applications. Existing GEL methods focus on learning useful representations for nodes, edges, or entire graphs with "small" labeled data. But the problem of data-efficient learning for subgraph prediction has not been explored. The challenges of this problem lie in the following aspects: 1) It is crucial for subgraphs to learn positional features to acquire structural information in the base graph in which they exist. Although the existing subgraph neural network method is capable of learning disentangled position encodings, the overall computational complexity is very high. 2) Prevailing graph augmentation methods for GEL, including rule-based, sample-based, adaptive, and automated methods, are not suitable for augmenting subgraphs because a subgraph contains fewer nodes but richer information such as position, neighbor, and structure. Subgraph augmentation is more susceptible to undesirable perturbations. 3) Only a small number of nodes in the base graph are contained in subgraphs, which leads to a potential "bias" problem that the subgraph representation learning is dominated by these "hot" nodes. By contrast, the remaining nodes fail to be fully learned, which reduces the generalization ability of subgraph representation learning. In this paper, we aim to address the challenges above and propose a Position-Aware Data-Efficient Learning framework for subgraph neural networks called PADEL. Specifically, we propose a novel node position encoding method that is anchor-free, and design a new generative subgraph augmentation method based on a diffused variational subgraph autoencoder, and we propose exploratory and exploitable views for subgraph contrastive learning. Extensive experiment results on three real-world datasets show the superiority of our proposed method over state-of-the-art baselines. Chang Liu 0078, Yuwen Yang, Zhe Xie, Hongtao Lu 0001, Yue Ding 0001 |
WSDM | 4 |
| 2023 | Trimap-guided feature mining and fusion network for natural image matting
Dongdong Yu, Zhaozhi Xie, Yaoyi Li, Zehuan Yuan, Hongtao Lu 0001 |
Comput. Vis. Image Underst. | 6 |
| 2023 | SSDA-YOLO: Semi-supervised domain adaptive YOLO for cross-domain object detection
Huayi Zhou 0001, Fei Jiang 0006, Hongtao Lu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2023 | RAFT-MSF: Self-Supervised Monocular Scene Flow Using Recurrent Optimizer
Bayram Bayramli, Junhwa Hur, Hongtao Lu 0001 |
Int. J. Comput. Vis. | 3 |
| 2023 | USMLP: U-shaped Sparse-MLP network for mass segmentation in mammograms
Jiaming Luo, Yongzhe Tang, Hongtao Lu 0001 |
Image Vis. Comput. | 4 |
| 2023 | Arbitrary shape scene text detector with accurate text instance generation based on instance-relevant contexts
Yangsong Zhang 0002, Bayram Bayramli, Hongtao Lu 0001 |
Multim. Tools Appl. | 4 |
| 2023 | Single Person Dense Pose Estimation via Geometric Equivariance ConsistencyabstractWe study the task of single person dense pose estimation. Specifically, given a human-centric image, we learn to map all human pixels onto a 3D, surface-based human body model. Existing methods approach this problem by fitting deep convolutional networks on sparse annotated points where the regression on both surface coordinate components for each body part is uncorrelated and optimized separately. In this work, we devise a novel, unified loss function that explicitly characterizes the correlation for surface coordinates regression, achieving significant improvements in both accuracy and efficiency. Furthermore, based on an observation that the image-to-surface correspondence is intrinsically invariant to geometric transformations from input images, we propose to enforce a geometric equivariance consistency on the target mapping, thereby allowing us to enable reliable supervision on large amounts of unlabeled pixels. We conduct comprehensive studies on the effectiveness of our approach using a quite simple network. Extensive experiments on the DensePose-COCO dataset show that our model achieves superior performance against previous state-of-the-art methods with much less computation complexity. We hope that our work would serve as a solid baseline for future study in the field. The code will be available athttps://github.com/Johnqczhang/densepose.pytorch. Qinchuan Zhang, Qin Zhou 0002, Yiru Zhao, Yao Liu 0014, Hongtao Lu 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 6 |
| 2022 | Structural and Statistical Texture Knowledge Distillation for Semantic SegmentationabstractExisting knowledge distillation works for semantic seg-mentation mainly focus on transfering high-level contextual knowledge from teacher to student. However, low-level texture knowledge is also of vital importance for characterizing the local structural pattern and global statistical prop-erty, such as boundary, smoothness, regularity and color contrast, which may not be well addressed by high-level deep features. In this paper, we are intended to take full advantage of both structural and statistical texture knowledge and propose a novel Structural and Statistical Texture Knowledge Distillation (SSTKD) framework for Semantic Segmentation. Specifically, for structural texture knowledge, we introduce a Contourlet Decomposition Module (CDM) that decomposes low-level features with iterative laplacian pyramid and directional filter bank to mine the structural texture knowledge. For statistical knowledge, we propose a Denoised Texture Intensity Equalization Module (DTIEM) to adaptively extract and enhance statistical texture knowledge through heuristics iterative quantization and denoised operation. Finally, each knowledge learning is supervised by an individual loss function, forcing the student network to mimic the teacher better from a broader perspective. Experiments show that the proposed method achieves state-of-the-art performance on Cityscapes, Pascal VOC 2012 and ADE20K datasets. Deyi Ji, Mingyuan Tao, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Hongtao Lu 0001 |
CVPR | 6 |
| 2022 | Exploring Category Consistency for Weakly Supervised Semantic SegmentationabstractSelf-supervised framework has been widely used in weakly supervised semantic segmentation. Generating a reliable and detailed pseudo mask label is the main challenge for improving the quality of predicted mask. In this paper, we propose Category Consistency Mask Refinement (CCMR) to explore the category consistency cued with the input image, and inject such information to mask refinement, guaranteeing the completeness of the refined mask. Moreover, we exploit Selective Weighted Pooling (SWP) to restrict the backward propagation of background, limiting the update of the background. Experimental results demonstrate that our methods can boost the performance on the PASCAL VOC 2012 segmentation benchmark, outperforming the state-of-the-art weakly supervised semantic segmentation methods. Zhaozhi Xie, Hongtao Lu 0001 |
ICASSP | 2 |
| 2022 | Coneface: Approximate Pairwise Loss for Face RecognitionabstractThe discriminability of learned face features is the key to a successful face recognition algorithm under the open-set protocol. Recent research exploits well-designed loss functions to penalize the angles between the deep features and their class centers for reaching the purpose of minimizing the intra-class variance and achieves a significant increase in recognition accuracy. In this paper, we proposed an approximate pairwise loss (APL) to encourage inter-class separability as well as intra-class compactness. More specifically, we use cones to approximate the location of hard examples in the feature space which replaces the thorny hard-mining step, and the APL is obtained by calculating the angular distance between the features of training samples and the cone and we named our method ConeFace. Moreover, ConeFace can be easily used together with other SOTA methods to improve the performance with negligible computational overhead. Extensive experiments on Labeled Face in the Wild (LFW), Celebrities in Frontal-Profile in the Wild (CFP), AgeDB-30, and MegaFace datasets show the effectiveness of the proposed ConeFace. Zijun Zhuang, Hongtao Lu 0001 |
ICASSP | 2 |
| 2022 | Merged U-Net for Bone Tumors X-Ray Images SegmentationabstractBone tumors X-ray image segmentation is crucial for many medical image processing applications such as X-ray image enhancement, lesion diagnosis, etc. In this paper, we propose a new topological structure of u-net, named merged u-net, for bone tumors X-ray image segmentation. We attach a top-down merged branch to the vanilla encoder-decoder network, enhancing the hierarchical feature aggregation. In the merged path, a multi-feature aggregation block, named merged gate, is proposed for better localization of lesion area. According to the attention matrix, the proposed merged gate can generate a reconstructed feature that contains both low-level information and high-level information. Then we incorporate the reconstructed feature with the small scare feature cued with a local feature fusion method, achieving multi-scare feature aggregation. Additionally, we propose a new X-ray image dataset for bone tumors segmentation, which consists of 88 benign images and 217 malignant images, provided with corresponding segmentation masks. Experimental results demonstrate that the proposed merged u-net outperforms other u-net based medical segmentation methods on the proposed X-ray image dataset. Zhaozhi Xie, Keyang Zhao, Sheng-Hui Wu, Jiong Mei, Hongtao Lu 0001 |
ICIP | 6 |
| 2022 | Enhanced discriminative graph convolutional network with adaptive temporal modelling for skeleton-based action recognition
Tamam Alsarhan, Usman Ali 0009, Hongtao Lu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2022 | Efficient dual attention SlowFast networks for video action recognition
Dafeng Wei, Liqing Wei, Siqian Chen, Shiliang Pu, Hongtao Lu 0001 |
Comput. Vis. Image Underst. | 7 |
| 2022 | ULN: An efficient face recognition method for person wearing a mask
Hongtao Lu 0001, Zijun Zhuang |
Multim. Tools Appl. | 1 |
| 2021 | Towards Dynamic and Scalable Active Learning with Neural Architecture Adaption for Object Detection
Fuhui Tang, Chenhan Jiang, Dafeng Wei, Hang Xu 0004, Andi Zhang 0001, Wei Zhang 0196, Hongtao Lu 0001, Chunjing Xu |
BMVC | 7 |
| 2021 | Semi-deterministic and Contrastive Variational Graph Autoencoder for RecommendationabstractVariational AutoEncoder (VAE) is a popular deep generative framework with a solid theoretical basis. There are many research efforts on improving VAE. Among the existing works, a recently proposed deterministic Regularized AutoEncoder (RAE) provides a new scheme for generative modeling. RAE fixes the variance of the inferred Gaussian approximate posterior distribution as a hyperparameter, and substitutes the stochastic encoder by injecting noise into the input of a deterministic decoder. However, the deterministic RAE has three limitations: 1) RAE needs to fit the variance; 2) RAE requires ex-post density estimation to ensure sample quality; 3) RAE employs an additional gradient regularization to ensure training smoothness. Thus, it raises an interesting research question: Can we maintain the flexibility of variational inference while simplifying VAE, and at the same time ensuring a smooth training process to obtain good generative performance? Based on the above motivation, in this paper, we propose a novel Semi-deterministic and Contrastive Variational Graph autoencoder (SCVG) for item recommendation. The core design of SCVG is to learn the variance of the approximate Gaussian posterior distribution in a semi-deterministic manner by aggregating inferred mean vectors from other connected nodes via graph convolution operation. We analyze the expressive power of SCVG for the Weisfeiler-Lehman graph isomorphism test, and we deduce the simplified form of the evidence lower bound of SCVG. Besides, we introduce an efficient contrastive regularization instead of gradient regularization. We empirically show that the contrastive regularization makes learned user/item latent representation more personalized and helps to smooth the training process. We conduct extensive experiments on three real-world datasets to show the superiority of our model over state-of-the-art methods for the item recommendation task. Codes are available at https://github.com/syxkason/SCVG. Yue Ding 0001, Yuxiang Shi, Bo Chen 0023, Chenghua Lin 0002, Hongtao Lu 0001, Jie Li 0002, Ruiming Tang, Dong Wang 0024 |
CIKM | 5 |
| 2021 | Laplacian Regularized Tensor Low-Rank Minimization for Hyperspectral Snapshot Compressive ImagingabstractSnapshot Compressive Imaging (SCI) systems, including hyperspectral compressive imaging and video compressive imaging, are designed to depict high-dimensional signals with limited data by mapping multiple images into one. One key module of SCI systems is a high quality reconstruction algorithm for original frames. However, most existing decoding algorithms are based on vectorization representation and fail to capture the intrinsic structural information of high dimensional signals. In this paper, we propose a tensor-based low-rank reconstruction algorithm with hyper-Laplacian constraint for hyperspectral SCI systems. First, we integrate the non-local self-similarity and tensor low-rank minimization approach to explore the intrinsic structural correlations along spatial and spectral domains. Then, we introduce a hyper-Laplacian constraint to model the global spectral structures, alleviating the ringing artifacts in the spatial domain. Experimental results on hyperspectral image corpus demonstrate the proposed algorithm achieves average 0.8~2.9 dB improvement in PSNR over state-of-the-art work. Fei Jiang 0006, Hongtao Lu 0001 |
ICASSP | 3 |
| 2021 | Text Detection by Jointly Learning Character and Word Regions
Deyang Wu, Xingfei Hu, Zhaozhi Xie, Usman Ali 0009, Hongtao Lu 0001 |
ICDAR (1) | 6 |
| 2021 | ShallowNet: An Efficient Lightweight Text Detection Network Based on Instance Count-Aware Supervision Information
Xingfei Hu, Deyang Wu, Fei Jiang 0006, Hongtao Lu 0001 |
ICONIP (1) | 5 |
| 2021 | Spatial Gradient Guided Learning and Semantic Relation Transfer for Facial Landmark Detection
Jian Wang 0137, Yaoyi Li, Hongtao Lu 0001 |
MMM (1) | 3 |
| 2021 | Adversarial and Contrastive Variational Autoencoder for Sequential RecommendationabstractSequential recommendation as an emerging topic has attracted increasing attention due to its important practical significance. Models based on deep learning and attention mechanism have achieved good performance in sequential recommendation. Recently, the generative models based on Variational Autoencoder (VAE) have shown the unique advantage in collaborative filtering. In particular, the sequential VAE model as a recurrent version of VAE can effectively capture temporal dependencies among items in user sequence and perform sequential recommendation. However, VAE-based models suffer from a common limitation that the representational ability of the obtained approximate posterior distribution is limited, resulting in lower quality of generated samples. This is especially true for generating sequences. To solve the above problem, in this work, we propose a novel method called Adversarial and Contrastive Variational Autoencoder (ACVAE) for sequential recommendation. Specifically, we first introduce the adversarial training for sequence generation under the Adversarial Variational Bayes (AVB) framework, which enables our model to generate high-quality latent variables. Then, we employ the contrastive loss. The latent variables will be able to learn more personalized and salient characteristics by minimizing the contrastive loss. Besides, when encoding the sequence, we apply a recurrent and convolutional structure to capture global and local relationships in the sequence. Finally, we conduct extensive experiments on four real-world datasets. The experimental results show that our proposed ACVAE model outperforms other state-of-the-art methods. Zhe Xie, Chengxuan Liu, Hongtao Lu 0001, Dong Wang 0024, Yue Ding 0001 |
WWW | 4 |
| 2021 | Small and accurate heatmap-based face alignment via distillation strategy and cascaded architecture
Jiaxin Si, Fei Jiang 0006, Ruimin Shen, Hongtao Lu 0001 |
Comput. Vis. Image Underst. | 4 |
| 2021 | Multi-modal constraint propagation via compatible conditional distribution reconstruction
Yaoyi Li, Hongtao Lu 0001 |
Neurocomputing | 2 |
| 2021 | A lightweight network for monocular depth estimation with decoupled body and edge supervision
Usman Ali 0009, Bayram Bayramli, Tamam Alsarhan, Hongtao Lu 0001 |
Image Vis. Comput. | 4 |
| 2020 | Natural Image Matting via Guided Contextual AttentionabstractOver the last few years, deep learning based approaches have achieved outstanding improvements in natural image matting. Many of these methods can generate visually plausible alpha estimations, but typically yield blurry structures or textures in the semitransparent area. This is due to the local ambiguity of transparent objects. One possible solution is to leverage the far-surrounding information to estimate the local opacity. Traditional affinity-based methods often suffer from the high computational complexity, which are not suitable for high resolution alpha estimation. Inspired by affinity-based method and the successes of contextual attention in inpainting, we develop a novel end-to-end approach for natural image matting with a guided contextual attention module, which is specifically designed for image matting. Guided contextual attention module directly propagates high-level opacity information globally based on the learned low-level affinity. The proposed method can mimic information flow of affinity-based methods and utilize rich features learned by deep neural networks simultaneously. Experiment results on Composition-1k testing set and alphamatting.com benchmark dataset demonstrate that our method outperforms state-of-the-art approaches in natural image matting. Code and models are available at https://github.com/Yaoyi-Li/GCA-Matting. Yaoyi Li, Hongtao Lu 0001 |
AAAI | 2 |
| 2020 | Ensemble-Based Deep Metric Learning for Few-Shot Learning
Yaoyi Li, Hongtao Lu 0001 |
ICANN (1) | 3 |
| 2020 | Inductive Guided Filter: Real-Time Deep Matting with Weakly Annotated Masks on Mobile DevicesabstractRecently, significant progress has been achieved in deep image matting. Most of the classical image matting methods are time-consuming and require an ideal trimap which is difficult to attain in practice. An efficient image matting method based on a weakly annotated mask is in demand for mobile applications. In this paper, we propose a novel method called Inductive Guided Filter, which tackles the real-time general image matting task with weakly annotated masks on mobile devices. The Inductive Guided Filter exploits the gradient prior implicit in Guided Filter to reduce the computational burden tremendously in a deep learning manner. The use of Gabor loss is also proposed for complicated textures in image matting. Moreover, we create an image matting dataset MAT-2793 with a variety of foreground objects. Experimental results demonstrate that our proposed method massively reduces running time with robust accuracy. Yaoyi Li, Jianfu Zhang 0003, Weijie Zhao 0003, Hongtao Lu 0001 |
ICME | 5 |
| 2020 | Image-based Table Cell Detection: a Novel Table Structure Decomposition Method with New DatasetabstractRecently deep learning has been applied to decompose table structure with the main ideas of detecting table lines and then forming table cells. However, the existing methods face problems in dealing with tables with rotation or no internal table lines. To tackle these problems, we propose a novel table structure decomposition method, which directly detects table cells as objects and creates table structure. Extensions to the existing object detection models including effective table projection module are proposed to adapt to the table cell detection. To support the training of the enhanced models, we create a large image-based table dataset TableCell with cell level annotations. A novel and efficient semi-supervised method is proposed to annotate this new dataset. Experiments demonstrate that our proposed table structure decomposition method is simple, effective and robust to the tables without table lines or with rotation. Our dataset and code will be made available11https://github.com/weidafeng/TableCell. Dafeng Wei, Hongtao Lu 0001, Yi Zhou 0003, Kai Chen 0006 |
ICPR | 2 |
| 2019 | Attribute-Driven Feature Disentangling and Temporal Aggregation for Video Person Re-IdentificationabstractVideo-based person re-identification plays an important role in surveillance video analysis, expanding image-based methods by learning features of multiple frames. Most existing methods fuse features by temporal average-pooling, without exploring the different frame weights caused by various viewpoints, poses, and occlusions. In this paper, we propose an attribute-driven method for feature disentangling and frame re-weighting. The features of single frames are disentangled into groups of sub-features, each corresponds to specific semantic attributes. The sub-features are re-weighted by the confidence of attribute recognition and then aggregated at the temporal dimension as the final representation. By means of this strategy, the most informative regions of each frame are enhanced and contributes to a more discriminative sequence representation. Extensive ablation studies demonstrate the effectiveness of feature disentangling as well as temporal re-weighting. The experimental results on the iLIDS-VID, PRID-2011 and MARS datasets demonstrate that our proposed method outperforms existing state-of-the-art approaches. Yiru Zhao, Xu Shen 0001, Zhongming Jin 0001, Hongtao Lu 0001, Xian-Sheng Hua 0001 |
CVPR | 4 |
| 2019 | Temporal Continuity Based Unsupervised Learning for Person Re-identification
Usman Ali 0009, Bayram Bayramli, Hongtao Lu 0001 |
ICONIP (5) | 3 |
| 2019 | FH-GAN: Face Hallucination and Recognition Using Generative Adversarial Network
Bayram Bayramli, Usman Ali 0009, Te Qi, Hongtao Lu 0001 |
ICONIP (1) | 4 |
| 2019 | Improve Image Captioning by Self-attention
Zhenru Li, Yaoyi Li, Hongtao Lu 0001 |
ICONIP (5) | 3 |
| 2019 | Community detection in scientific collaborative network with bayesian matrix learning
Xiaohua Shi, Hongtao Lu 0001 |
Frontiers Comput. Sci. | 2 |
| 2018 | Efficient Multi-Dimensional Tensor Sparse Coding Using t-Linear CombinationabstractIn this paper, we propose two novel multi-dimensional tensor sparse coding (MDTSC) schemes using the t-linear combination. Based on the t-linear combination, the shifted versions of the bases are used for the data approximation, but without need to store them. Therefore, the dictionaries of the proposed schemes are more concise and the coefficients have richer physical explanations. Moreover, we propose an efficient alternating minimization algorithm, including the tensor coefficient learning and the tensor dictionary learning, to solve the proposed problems. For the tensor coefficient learning, we design a tensor-based fast iterative shrinkage algorithm. For the tensor dictionary learning, we first divide the problem into several nearly-independent subproblems in the frequency domain, and then utilize the Lagrange dual to further reduce the number of optimization variables. Experimental results on multi-dimensional signals denoising and reconstruction (3DTSC, 4DTSC, 5DTSC) show that the proposed algorithms are more efficient and outperform the state-of-the-art tensor-based sparse coding models. Fei Jiang 0006, Xiao-Yang Liu, Hongtao Lu 0001, Ruimin Shen |
AAAI | 3 |
| 2018 | Multilevel Collaborative Attention Network for Person Search
Zhenyong Fu, Hongtao Lu 0001 |
ACCV (1) | 4 |
| 2018 | An Adversarial Approach to Hard Triplet Generation
Yiru Zhao, Zhongming Jin 0001, Guo-Jun Qi, Hongtao Lu 0001, Xian-Sheng Hua 0001 |
ECCV (9) | 4 |
| 2018 | Anisotropic Total Variation Regularized Low-Rank Tensor Completion Based On Tensor Nuclear Norm for Color Image InpaintingabstractIn this paper, we propose a novel low-rank tensor completion (LRTC) model under the circulant algebra for color image inpainting, which simultaneously preserves the low-rank structures of images, and also explore the local smooth and piecewise priors of the images in the spatial domain. First, color images are naturally represented by 3-order tensors which preserve the intrinsic structures of color images. Second, we preserve the low-rank structures of these tensors with tensor nuclear norm, which can simultaneously exploit the correlations among the spatial and channel domains. Third, we integrate an anisotropic total variation into our low-rank tensor completion model, which preserve the local smooth and piecewise priors of color images. Then, an efficient alternating direction method of multipliers (ADMM) is proposed to solve the resulting optimization problem. Experimental results on eight widely used color images demonstrate the effectiveness and superiority of the proposed algorithm. Fei Jiang 0006, Xiao-Yang Liu, Hongtao Lu 0001, Ruimin Shen |
ICASSP | 3 |
| 2018 | End to end multi-scale convolutional neural network for crowd countingabstractCrowd counting is a challenging task in computer vison field and haven’t been well addressed until now. In this paper, we intend to develop an end to end multi-scale deep convolutional neural network(CNN) model that can accurately estimate the crowd count from an individual image with arbitrary crowd density and perspective. The proposed model extract multi-scale deep CNN features from the input image and regress the crwod count directly, without any post-processing . Hence our model could handle muti-scale targets well in various crowd scene. We evaluate our model on several benchmark datasets and the performance outperforms some state-of-the-art methods. What’s more, due to the end-to-end characteristics, our model demonstrates good practical application performance. Deyi Ji, Hongtao Lu 0001, Tongzhen Zhang |
ICMV | 2 |
| 2018 | On Multi-modal Fusion Learning in constraint propagation
Yaoyi Li, Hongtao Lu 0001 |
Inf. Sci. | 2 |
| 2017 | Adaptive Overlapping Community Detection with Bayesian NonNegative Matrix Factorization
Xiaohua Shi, Hongtao Lu 0001, Guanbo Jia |
DASFAA (2) | 2 |
| 2017 | Graph regularized tensor sparse coding for image representationabstractSparse coding (SC) is a unsupervised learning scheme that has received an increasing amount of interests in recent years. However, conventional SC vectorizes the input images, which destructs the intrinsic spatial structures of the images. In this paper, we propose a novel graph regularized tensor sparse coding (GTSC) for image representation. GTSC preserves the local proximity of elementary structures in the image by adopting the newly proposed tubal-tensor representation. Simultaneously, it considers the intrinsic geometric properties by imposing graph regularization that has been successfully applied to uncover the geometric distribution for the image data. Moreover, the learned sparse representations by GTSC have better physical explanations as the key operation (i.e., circular convolution) in the tubal-tensor model preserves the shifting invariance property. Experimental results on image clustering demonstrate the effectiveness of the proposed scheme. Fei Jiang 0006, Xiao-Yang Liu, Hongtao Lu 0001, Ruimin Shen |
ICME | 3 |
| 2017 | Quantizable deep representation learning with gradient snapping layer for large scale searchabstractRecent advance of large scale similarity search requires to learn deep representations that both strongly preserve similarities between data pairs and can be accurately quantized via vector quantization. Existing methods simply leverage quantization loss and similarity loss, which result in unexpectedly biased back-propagating gradients and affect the search performances. To this end, we propose a novel gradient snapping layer (GSL) to regularize the back-propagating gradient towards a neighboring codeword, the generated gradients works better on reducing similarity loss and also propel the learned representations to be accurately quantized. Joint deep representation and vector quantization learning can be easily performed by alternatively optimizing the quantization codebook and the deep neural network. The proposed framework is compatible with various existing vector quantization approaches. Experimental results on various standard benchmark datasets demonstrate that the proposed framework is effective, flexible and outperforms the state-of-the-art large scale similarity search methods. Shicong Liu, Hongtao Lu 0001 |
ICME | 2 |
| 2017 | Space shuttle model: A physics inspired method for learning quantizable deep representationsabstractRecent advance of large scale similarity search involves using deeply learned representations to improve the search accuracy and use vector quantization methods to increase the search speed. However, how to learn deep representations that both strongly preserve similarities between data pairs and can be accurately quantized via vector quantization remains a challenging task. In this paper, we propose a novel physics based method named space shuttle model (SSM) to learn effective deep representations that can be accurately quantized. It consider network output as a roaming space shuttle “propelled” by similarity loss and subject to “gravitational forces” from quantization codewords. SSM is related to momentum methods commonly used in deep learning but is applied on network outputs instead of network parameters. Experimental results on large scale similarity search demonstrate that the proposed framework outperforms the state-of-the-art. Shicong Liu, Hongtao Lu 0001 |
ICME | 2 |
| 2017 | Learning deep representations with diode loss for quantization-based similarity searchabstractRecent advance of large scale similarity search involves using deeply learned representations to improve the search accuracy and apply vector quantization techniques to accelerate the search speed. However, simultaneous learning of deep representations and vector quantizers still remains ineffective. To this end, we propose to directly optimize the asymmetric distance between a query representation and the quantized database representations. A novel diode loss is proposed, it wraps a commonly used similarity loss function and then it allows effective end-to-end learning of both deep representations and vector quantizers with a siamese network. The proposed learning framework is compatible with various existing vector quantization approaches, and is compatible with commonly used loss functions for learning representations preserving similarities. Experimental results demonstrate that the proposed framework is effective, flexible and outperforms the state-of-the-art large scale similarity search methods. Shicong Liu, Hongtao Lu 0001 |
IJCNN | 2 |
| 2017 | Stylized Adversarial AutoEncoder for Image GenerationabstractIn this paper, we propose an autoencoder-based generative adversarial network (GAN) for automatic image generation, which is called "stylized adversarial autoencoder". Different from existing generative autoencoders which typically impose a prior distribution over the latent vector, the proposed approach splits the latent variable into two components: style feature and content feature, both encoded from real images. The split of the latent vector enables us adjusting the content and the style of the generated image arbitrarily by choosing different exemplary images. In addition, a multiclass classifier is adopted in the GAN network as the discriminator, which makes the generated images more realistic. We performed experiments on hand-writing digits, scene text and face datasets, in which the stylized adversarial autoencoder achieves superior results for image generation as well as remarkably improves the corresponding supervised recognition task. Yiru Zhao, Bing Deng, Jianqiang Huang 0001, Hongtao Lu 0001, Xian-Sheng Hua 0001 |
ACM Multimedia | 4 |
| 2017 | Spatio-Temporal AutoEncoder for Video Anomaly DetectionabstractAnomalous events detection in real-world video scenes is a challenging problem due to the complexity of "anomaly" as well as the cluttered backgrounds, objects and motions in the scenes. Most existing methods use hand-crafted features in local spatial regions to identify anomalies. In this paper, we propose a novel model called Spatio-Temporal AutoEncoder (ST AutoEncoder or STAE), which utilizes deep neural networks to learn video representation automatically and extracts features from both spatial and temporal dimensions by performing 3-dimensional convolutions. In addition to the reconstruction loss used in existing typical autoencoders, we introduce a weight-decreasing prediction loss for generating future frames, which enhances the motion feature learning in videos. Since most anomaly detection datasets are restricted to appearance anomalies or unnatural motion anomalies, we collected a new challenging dataset comprising a set of real-world traffic surveillance videos. Several experiments are performed on both the public benchmarks and our traffic dataset, which show that our proposed method remarkably outperforms the state-of-the-art approaches. Yiru Zhao, Bing Deng, Chen Shen 0003, Yao Liu 0014, Hongtao Lu 0001, Xian-Sheng Hua 0001 |
ACM Multimedia | 5 |
| 2017 | Generalized Residual Vector Quantization and Aggregating Tree for Large Scale SearchabstractVector quantization is an essential tool for tasks involving large scale data, for example, large scale similarity search, which is crucial for content-based information retrieval and analysis. In this paper, we propose a novel vector quantization framework that iteratively minimizes quantization error. First, we provide a detailed review on a relevant vector quantization method named residual vector quantization (RVQ). Next, we propose generalized residual vector quantization (GRVQ) to further improve over RVQ. Many vector quantization methods can be viewed as special cases of our proposed method. To enable GRVQ on billion scale data, we introduce a nonexhaustive search scheme named aggregating tree (A-Tree) for high dimensional data that uses GRVQ encodings to build a radix tree and perform the nearest neighbor search by beam search. To search accurately and efficiently, VQ-encodings should satisfy locally aggregating encoding criterion: For any node of the corresponding A-Tree, neighboring vectors should aggregate in fewer subtrees to make beam search efficient. We show that the proposed GRVQ encodings best satisfy the suggested criterion, and the joint use of GRVQ and A-Tree shows significantly better performances on billion scale datasets. Our methods are validated on several standard benchmark datasets. Experimental results and empirical analysis show the superior efficiency and effectiveness of our proposed methods compared to the state-of-the-art for large scale search. Shicong Liu, Junru Shao, Hongtao Lu 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | Community Inference with Bayesian Non-negative Matrix Factorization
Xiaohua Shi, Hongtao Lu 0001 |
APWeb (1) | 2 |
| 2016 | Online self-organizing hashingabstractHashing for similarity search in large scale data has become an increasingly popular technique. K-means Hashing (KMH) has been proven effective because of the benefits of adaptive k-means quantization. However, KMH is a batch-based learning model requiring high time and storage complexities, which makes it hard to load large scale data into memory to train and deal with streaming data. To address this problem, in this paper we propose an online hashing method using Self-Organizing Map (SOM) algorithm, named as Online Self-Organizing Hashing (SOH). Specifically, we map the training data to an affinity preserving hyper-cube with each vertex assigned a binary code using a SOM alike algorithm. After training, a new data point is quantized into a vertex of the hyper-cube and encoded into related binary code. Experimental results demonstrate that SOH has better or comparable retrieval performance to various state-of-the-art hashing methods while simultaneously requiring rather low computational complexity and storage space. Junxuan Chen, Yaoyi Li, Hongtao Lu 0001 |
ICME | 3 |
| 2016 | Adaptive affinity matrix for unsupervised metric learningabstractSpectral clustering is one of the most popular clustering approaches with the capability to handle some challenging clustering problems. Only a little work of spectral clustering focuses on the explicit linear map which can be viewed as the distance metric learning. In practice, the selection of the affinity matrix exhibits a tremendous impact on the unsupervised learning. In this paper, we propose a novel method, dubbed Adaptive Affinity Matrix (AdaAM), to learn an adaptive affinity matrix and derive a distance metric. We assume the affinity matrix to be positive semidefinite with ability to quantify the pairwise dissimilarity. Our method is based on posing the optimization of objective function as a spectral decomposition problem. The provided matrix can be regarded as the optimal representation of pairwise relationship on the manifold. Extensive experiments on a number of image data sets show the effectiveness and efficiency of AdaAM. Yaoyi Li, Junxuan Chen, Yiru Zhao, Hongtao Lu 0001 |
ICME | 4 |
| 2016 | Generalized residual vector quantization for large scale dataabstractVector quantization is an essential tool for tasks involving large scale data, for example, large scale similarity search, which is crucial for content-based information retrieval and analysis. In this paper, we propose a novel vector quantization framework that iteratively minimizes quantization error. First, we provide a detailed review on a relevant vector quantization method named residual vector quantization (RVQ). Next, we propose generalized residual vector quantization (GRVQ) to further improve over RVQ. Many vector quantization methods can be viewed as the special cases of our proposed framework. We evaluate GRVQ on several large scale benchmark datasets for large scale search, classification and object retrieval. We compared GRVQ with existing methods in detail. Extensive experiments demonstrate our GRVQ framework substantially outperforms existing methods in term of quantization accuracy and computation efficiency. Shicong Liu, Junru Shao, Hongtao Lu 0001 |
ICME | 3 |
| 2016 | Deep CTR Prediction in Display AdvertisingabstractClick through rate (CTR) prediction of image ads is the core task of online display advertising systems, and logistic regression (LR) has been frequently applied as the prediction model. However, LR model lacks the ability of extracting complex and intrinsic nonlinear features from handcrafted high-dimensional image features, which limits its effectiveness. To solve this issue, in this paper, we introduce a novel deep neural network (DNN) based model that directly predicts the CTR of an image ad based on raw image pixels and other basic features in one step. The DNN model employs convolution layers to automatically extract representative visual features from images, and nonlinear CTR features are then learned from visual features and other contextual features by using fully-connected layers. Empirical evaluations on a real world dataset with over 50 million records demonstrate the effectiveness and efficiency of this method. Junxuan Chen, Baigui Sun, Hao Li 0030, Hongtao Lu 0001, Xian-Sheng Hua 0001 |
ACM Multimedia | 4 |
| 2016 | LSOD: Local Sparse Orthogonal Descriptor for Image MatchingabstractWe propose a novel method for feature description used for image matching in this paper. Our method is inspired by the autoencoder, an artificial neural network designed for learning efficient codings. Sparse and orthogonal constraints are imposed on the autoencoder and make it a highly discriminative descriptor. It is shown that the proposed descriptor is not only invariant to geometric and photometric transformations (such as viewpoint change, intensity change, noise, image blur and JPEG compression), but also highly efficient. We compare it with existing state-of-the-art descriptors on standard benchmark datasets, the experimental results show that our LSOD method yields better performance both in accuracy and efficiency. Yiru Zhao, Yaoyi Li, Zhiwen Shao, Hongtao Lu 0001 |
ACM Multimedia | 4 |
| 2015 | Community Detection in Social Network with Pairwisely Constrained Symmetric Non-Negative Matrix FactorizationabstractNon-negative Matrix Factorization (NMF) aims to find two non-negative matrices whose product approximates the original matrix well, and is widely used in clustering condition with good physical interpretability and universal applicability. Detecting communities with NMF can keep non-negative network physical definition and effectively capture communities-based structure in the low dimensional data space. However some NMF methods in community detection did not concern with more network inner structures or existing ground-truth community information. Xiaohua Shi, Hongtao Lu 0001, Yangcheng He |
ASONAM | 2 |
| 2015 | Graph regularized non-negative local coordinate factorization with pairwise constraints for image representationabstractChen et al. proposed a non-negative local coordinate factorization algorithm for feature extraction (NLCF) [1], which incorporated the local coordinate constraint into non-negative matrix factorization (NMF). However, NLCF is actually a unsupervised method without making use of prior information of problems in hand. In this paper, we propose a novel graph regularized non-negative local coordinate factorization with pairwise constraints algorithm (PCGNLCF) for image representation. PCGNLCF incorporates pairwise constraints and graph Laplacian into NLCF. More specifically, we expect that data points having pairwise must-link constraints will have the similar coordinates as much as possible, while data points with pairwise cannot-link constraints will have distinct coordinates as much as possible. Experimental results show the effectiveness of our proposed method in comparison to the state-of-the-art algorithms on several real-world applications. Yangcheng He, Hongtao Lu 0001, Bao-Liang Lu |
ICME | 2 |
| 2015 | Distance Preserving Marginal Hashing for image retrievalabstractHashing for image retrieval has attracted lots of attentions in recent years due to its fast computational speed and storage efficiency. Many existing hashing methods obtain the hashing functions through mapping neighbor items to similar codes, while ignoring the non-neighbor items. One exception is the Local Linear Spectral Hashing (LLSH), which introduces negative values into the local affinity matrix to map non-neighbor images to non-similar codes. However, setting 10th percentile distance in affinity matrix as a threshold, which is used to judge neighbors and non-neighbors, is not reasonable. In this paper, we propose a novel unsupervised hashing method called Distance Preserving Marginal Hashing (DPMH) which not only makes the average Hamming distance minimized for the intra-cluster pairs and maximized for the inter-cluster pairs, but also preserves the distance of non-neighbor points. Furthermore, we adopt an efficient sequential procedure to learn the hashing functions. The experimental results on two large-scale benchmark datasets demonstrate the effectiveness and efficiency of our method over other state-of-the-art unsupervised methods. Hongtao Lu 0001, Bao-Liang Lu |
ICME | 3 |
| 2015 | Constrained Non-negative Matrix Factorization with Graph Laplacian
Yangcheng He, Hongtao Lu 0001 |
ICONIP (3) | 3 |
| 2015 | Local similarity learning for pairwise constraint propagationabstractPairwise constraint propagation studies the problem of propagating the scarce pairwise constraints across the entire dataset. Effective propagation algorithms have previously been designed based on the graph-based semi-supervised learning framework. Therefore, these previous constraint propagation methods rely critically on a good similarity measure over the data points. Improper or noisy similarity measurements may dramatically degrade the performance of the constraint propagation algorithms. In this paper, we make attempt to exploit the available pairwise constraints to learn a new set of similarities, which are consistent with the supervisory information in the pairwise constraints, before propagating these initial constraints. Our method is a local learning algorithm. More specifically, we compute the similarities at each data point through simultaneously minimizing the local reconstruction error and local constraint error. The proposed method has been tested in the constrained clustering tasks on eight real-life datasets and then shown to achieve significant improvements with respect to the state of the arts. Zhenyong Fu, Zhiwu Lu 0001, Horace Ho-Shing Ip, Hongtao Lu 0001 |
Multim. Tools Appl. | 4 |
| 2015 | Non-negative Matrix Factorization with Pairwise Constraints and Graph Laplacian
Yangcheng He, Hongtao Lu 0001, Lei Huang 0005, Xiaohua Shi |
Neural Process. Lett. | 2 |
| 2014 | Locality Preserving Discriminative HashingabstractHashing for large scale similarity search has become more and more popular because of its improvement in computational speed and storage reduction. Semi-supervised Hashing (SSH) has been proven effective since it integrates both labeled and unlabeled data to leverage semantic similarity while keeping robust to overfitting. However, it ignores the global label information and the local structure of the feature space. In this paper, we concentrate on these two issues and propose a novel semi-supervised hashing method called Locality Preserving Discriminative Hashing which combines two classical dimensionality reduction approaches, Linear Discriminant Analysis (LDA) and Locality Preserving Projection (LPP). The proposed method presents a rigorous formulation in which the supervised term tries to maintain the global information of the labeled data while the unsupervised term provides effective regularization to model local relationships of the unlabeled data. We apply an efficient sequential procedure to learn the hashing functions. Experimental comparisons with other state-of-the-art methods on three large scale datasets demonstrate the effectiveness and efficiency of our method. Hongtao Lu 0001, Yangcheng He, Shaokun Feng |
ACM Multimedia | 2 |
| 2014 | An integrated Gaussian mixture model to estimate vigilance level based on EEG recordings
Jing-Nan Gu, Hongtao Lu 0001, Bao-Liang Lu |
Neurocomputing | 2 |
| 2014 | Semi-supervised non-negative matrix factorization for image clustering with graph Laplacian
Yangcheng He, Hongtao Lu 0001, Saining Xie |
Multim. Tools Appl. | 2 |
| 2014 | Pairwise constrained concept factorization for data representation
Yangcheng He, Hongtao Lu 0001, Lei Huang 0005, Saining Xie |
Neural Networks | 2 |
| 2014 | Non-negative and sparse spectral clustering
Hongtao Lu 0001, Zhenyong Fu |
Pattern Recognit. | 1 |
| 2013 | Image Classification Based on Weight Adjustment before Feature Pooling
Shaokun Feng, Hongtao Lu 0001, Lei Huang 0005 |
ICONIP (3) | 2 |
| 2012 | Multi-task co-clustering via nonnegative matrix factorization
Saining Xie, Hongtao Lu 0001, Yangcheng He |
ICPR | 2 |
| 2012 | Modalities consensus for multi-modal constraint propagationabstractThis paper presents a novel modalities consensus framework for multi-modal pairwise constraint propagation (MCP). We first combine multiple single-modal constraint propagation (SCP) problems together, and then explicitly introduce a new modalities consensus regularizer to force the propagation results on different modalities to be consistent with each other. With a separable consensus regularizer, the proposed approach can be effectively solved using an alternating optimization way. More importantly, based on our modalities consensus framework, two single-modal constraint propagation algorithms can be directly reformulated as two well-defined multi-modal solutions. Experimental results on constrained clustering tasks have shown that the proposed framework can achieve significant improvements with respect to the state of the arts. Zhenyong Fu, Hongtao Lu 0001, Horace Ho-Shing Ip, Zhiwu Lu 0001 |
ACM Multimedia | 2 |
| 2012 | Incremental visual objects clustering with the growing vocabulary tree
Zhenyong Fu, Hongtao Lu 0001 |
Multim. Tools Appl. | 2 |
| 2012 | Novel robust image watermarking based on subsampling and DWT
Wei Lu 0001, Wei Sun 0007, Hongtao Lu 0001 |
Multim. Tools Appl. | 3 |
| 2011 | Symmetric Graph Regularized Constraint PropagationabstractThis paper presents a novel symmetric graph regularization framework for pairwise constraint propagation. We first decompose the challenging problem of pairwise constraint propagation into a series of two-class label propagation subproblems and then deal with these subproblems by quadratic optimization with symmetric graph regularization. More importantly, we clearly show that pairwise constraint propagation is actually equivalent to solving a Lyapunov matrix equation, which is widely used in Control Theory as a standard continuous-time equation. Different from most previous constraint propagation methods that suffer from severe limitations, our method can directly be applied to multi-class problem and also can effectively exploit both must-link and cannot-link constraints. The propagated constraints are further used to adjust the similarity between data points so that they can be incorporated into subsequent clustering. The proposed method has been tested in clustering tasks on six real-life data sets and then shown to achieve significant improvements with respect to the state of the arts. Zhenyong Fu, Zhiwu Lu 0001, Horace Ho-Shing Ip, Yuxin Peng 0001, Hongtao Lu 0001 |
AAAI | 5 |
| 2011 | An Integrated Hierarchical Gaussian Mixture Model to Estimate Vigilance Level Based on EEG Recordings
Jing-Nan Gu, Hongtao Lu 0001, Bao-Liang Lu |
ICONIP (1) | 3 |
| 2011 | Multi-modal constraint propagation for heterogeneous image clusteringabstractThis paper presents a multi-modal constraint propagation approach to exploiting pairwise constraints for constrained clustering tasks on multi-modal datasets. Pairwise constraint propagation methods have previously been designed primarily for single modality data and cannot be directly applied to multi-modal data or a dataset with multiple representations. In this paper, we provide an effective solution to the multi-modal constraint propagation problem by decomposing it into a set of independent multi-graph based two-class label propagation subproblems which are then merged into a unified problem and solved by quadratic optimization. We also show that such a formulation yields a closed-form solution. Our approach allows the initial pairwise constraints to be propagated throughout the entire multi-modal dataset. The propagated constraints are further used to refine the similarities between the objects for subsequent clustering tasks. The proposed method has been tested in constrained clustering tasks on two real-life multi-modal image datasets and shown to achieve significant improvements with respect to the single modality methods. Zhenyong Fu, Horace Ho-Shing Ip, Hongtao Lu 0001, Zhiwu Lu 0001 |
ACM Multimedia | 3 |
| 2011 | Revealing digital fakery using multiresolution decomposition and higher order statistics
Wei Lu 0001, Wei Sun 0007, Korris Fu-Lai Chung, Hongtao Lu 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2011 | Digital image splicing detection based on approximate run length
Zhongwei He, Wei Sun 0007, Wei Lu 0001, Hongtao Lu 0001 |
Pattern Recognit. Lett. | 4 |
| 2008 | Blind Image Watermark Analysis Using Feature Fusion and Neural Network Classifier
Wei Lu 0001, Wei Sun 0007, Hongtao Lu 0001 |
ISNN (2) | 3 |
| 2006 | A Direct Evolutionary Feature Extraction Algorithm for Classifying High Dimensional Data
Qijun Zhao, David Zhang 0001, Hongtao Lu 0001 |
AAAI | 3 |
| 2006 | Image Fakery and Neural Network Based Detection
Wei Lu 0001, Korris Fu-Lai Chung, Hongtao Lu 0001 |
ISNN (2) | 3 |
| 2006 | Robust Image Watermarking Using RBF Neural Network
Wei Lu 0001, Hongtao Lu 0001, Korris Fu-Lai Chung |
ISNN (2) | 2 |
| 2006 | Parsimonious Feature Extraction Based on Genetic Algorithms and Support Vector Machines
Qijun Zhao, Hongtao Lu 0001, David Zhang 0001 |
ISNN (1) | 2 |
| 2006 | A fast evolutionary pursuit algorithm based on linearly combining vectors
Qijun Zhao, Hongtao Lu 0001, David Zhang 0001 |
Pattern Recognit. | 2 |
| 2005 | SVR-Based Oblivious Watermarking Scheme
Yonggang Fu, Ruimin Shen, Hongtao Lu 0001, Xusheng Lei |
ISNN (2) | 3 |
| 2005 | Subsampling-Based Robust Watermarking Using Neural Network Detector
Wei Lu 0001, Hongtao Lu 0001, Korris Fu-Lai Chung |
ISNN (2) | 2 |
| 2005 | A novel image watermarking scheme based on support vector regression
Ruimin Shen, Yonggang Fu, Hongtao Lu 0001 |
J. Syst. Softw. | 3 |
| 2005 | Global exponential stability of delayed competitive neural networks with different time scales
Hongtao Lu 0001, Zhenya He |
Neural Networks | 1 |
| 2005 | Absolute exponential stability of a class of recurrent neural networks with multiple and variable delays
Hongtao Lu 0001, Ruimin Shen, Korris Fu-Lai Chung |
Theor. Comput. Sci. | 1 |
| 2005 | Global exponential convergence of Cohen-Grossberg neural networks with time delaysabstractIn this paper, we derive a general sufficient condition ensuring global exponential convergence of Cohen-Grossberg neural networks with time delays by constructing a novel Lyapunov functional and smartly estimating its derivative. The proposed condition is related to the convex combinations of the column-sum and the row-sum of the connection matrices and also relaxes the constraints on the network coefficients. Therefore, the proposed condition generalizes some previous results in the literature. Hongtao Lu 0001, Ruimin Shen, Korris Fu-Lai Chung |
IEEE Trans. Neural Networks | 1 |
| 2004 | Optimal Watermark Detection Based on Support Vector Machines
Yonggang Fu, Ruimin Shen, Hongtao Lu 0001 |
ISNN (1) | 3 |
| 2004 | Color Image Watermarking Based on Neural Networks
Wei Lu 0001, Hongtao Lu 0001, Ruimin Shen |
ISNN (2) | 2 |
| 2004 | Some sufficient conditions for global exponential stability of delayed Hopfield neural networks
Hongtao Lu 0001, Korris Fu-Lai Chung, Zhenya He |
Neural Networks | 1 |
| 1995 | On the Capacity of Intraconnected Bidirectional Associative MemoryabstractIn this paper, we addressed a theoretical analysis for the capacity of parallel intra-connected bidirectional associative memory (MIBAM) and proved the conclusions: two MIBAM with the equal total number of neurons have the equal recalling probability for m pairs of stored pattern pairs if m is not too large. The results of computer simulation support the conclusions well. Baoyun Wang, Luxi Yang, Hongtao Lu 0001, Zhenya He |
ISCAS | 3 |
| 1995 | A New Type of Chaotic Attractor with Cellular Neural NetworksabstractBy computer simulation, we detect a new type of strange attractor in a three-cell cellular neural network which defers from that found by other researchers. Bifurcation phenomena are analyzed. Hongtao Lu 0001, Luxi Yang, Baoyun Wang, Zhenya He |
ISCAS | 1 |