EDBT 2026 Demo / reviewers in the wild / expert
Hongyuan Zhang 0001
dblp:96/1452-1
· DBLP profile ↗
43ranked-venue papers
8as first author
41since 2021 · last 2026
0000-0003-4274-7332ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 7 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rectified Noise: A Generative Model Using Positive-incentive NoiseabstractRectified Flow (RF) has been widely used as an effective generative model. Although RF is primarily based on probability flow Ordinary Differential Equations (ODE), recent studies have shown that injecting noise through reverse-time Stochastic Differential Equations (SDE) for sampling can achieve superior generative performance. Inspired by Positive-incentive Noise (Pi-noise), we propose an innovative generative algorithm to train Pi-noise generators, namely Rectified Noise (RN), which improves the generative performance by injecting Pi-noise into the velocity field of pre-trained RF models. After introducing the Rectified Noise pipeline, pre-trained RF models can be efficiently transformed into Pi-noise generators. We validate Rectified Noise by conducting extensive experiments across various model architectures on different datasets. Notably, we find that: (1) RF models using Rectified Noise reduce FID from10.16 to 9.05 on ImageNet-1k. (2) The models of Pi-noise generators achieve improved performance with only 0.39% additional training parameters. Yanchen Xu, Sida Huang, Yubin Guo, Hongyuan Zhang 0001 |
AAAI | 5 |
| 2026 | Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion TransformersabstractWith the development of diffusion models, enhancing spatial controllability in text-to-image generation has become a vital challenge. As a representative task for addressing this challenge, layout-to-image generation aims to generate images that are spatially consistent with the given layout condition. Existing layout-to-image methods typically introduce the layout condition by integrating adapter modules into the base generative model. However, the generated images often exhibit low visual quality and stylistic inconsistency with the base model, indicating a loss of pretrained knowledge. To alleviate this issue, we construct the Layout Synthesis (LaySyn) dataset, which leverages images synthesized by the base model itself to mitigate the distribution shift from the pretraining data. Moreover, we propose the Layout Control (Laytrol) Network, in which parameters are inherited from MM-DiT to preserve the pretrained knowledge of the base model. To effectively activate the copied parameters and avoid disturbance from unstable control conditions, we adopt a dedicated initialization scheme for Laytrol. In this scheme, the layout encoder is initialized as a pure text encoder to ensure that its output tokens remain within the data domain of MM-DiT. Meanwhile, the outputs of the layout control network are initialized to zero. In addition, we apply Object-level Rotary Position Embedding to the layout tokens to provide coarse positional information. Qualitative and quantitative experiments demonstrate the effectiveness of our method. Sida Huang, Ping Luo 0002, Hongyuan Zhang 0001 |
AAAI | 4 |
| 2026 | CoLM: Collaborative Large Models via a Client-Server ParadigmabstractLarge models have achieved remarkable performance across a range of reasoning and understanding tasks. Prior work often utilizes model ensembles or multi-agent systems to collaboratively generate responses, effectively operating in a server-to-server paradigm. However, such approaches do not align well with practical deployment settings, where a limited number of server-side models are shared by many clients under modern internet architectures. In this paper, we introduce CoLM (Collaboration in Large-Models), a novel framework for collaborative reasoning that redefines cooperation among large models from a client-server perspective. Unlike traditional ensemble methods that rely on simultaneous inference from multiple models to produce a single output, CoLM allows the outputs of multiple models to be aggregated or shared, enabling each client model to independently refine and update its own generation based on these high-quality outputs. This design enables collaborative benefits by fully leveraging both client-side and shared server-side models. We further extend CoLM to vision-language models (VLMs), demonstrating its applicability beyond language tasks. Experimental results across multiple benchmarks show that CoLM consistently improves model performance on previously failed queries, highlighting the effectiveness of collaborative guidance in enhancing single-model capabilities. Sida Huang, Hongyuan Zhang 0001 |
AAAI | 3 |
| 2026 | Explore How to Inject Beneficial Noise in MLLMsabstractMultimodal Large Language Models (MLLMs) have played an increasingly important role in multimodal intelligence. However, the existing fine-tuning methods often ignore cross-modal heterogeneity, limiting their full potential. In this work, we propose a novel fine-tuning strategy by injecting beneficial random noise, which outperforms previous methods and even surpasses full fine-tuning, with minimal additional parameters. The proposed Multimodal Noise Generator (MuNG) enables efficient modality fine-tuning by injecting customized noise into the frozen MLLMs. Specifically, we reformulate the reasoning process of MLLMs from a variational inference perspective, upon which we design a multimodal noise generator that dynamically analyzes cross-modal relationships in image-text pairs to generate task adaptive beneficial noise. Injecting this type of noise into the MLLMs effectively suppresses irrelevant semantic components, leading to significantly improved cross-modal representation alignment and enhanced performance on downstream tasks. Experiments on two mainstream MLLMs, QwenVL and LLaVA, demonstrate that our method surpasses full parameter fine-tuning and other existing fine-tuning approaches, while requiring adjustments to only about 1~2% additional parameters. Ruishu Zhu, Sida Huang, Ziheng Jiao, Hongyuan Zhang 0001 |
AAAI | 4 |
| 2026 | Predict overlooked nodes: A fast graph representation learning paradigm
Yanchen Xu, Hongyuan Zhang 0001, Xuelong Li 0001 |
Pattern Recognit. | 3 |
| 2026 | Miles: Metric Learning With Expandable Subspace for Pre-Trained Model-Based Class-Incremental LearningabstractClass Incremental Learning (CIL) aims to learn new concepts consistently from a data stream without forgetting. Unlike typical CIL methods which need to learn a model from scratch, pre-trained model (PTM) can easily adapt to a new task with fine-tuning. However, existing PTM-based CIL methods fail to achieve a trade-off between performance and computational expenditure, i.e., they either adopt the same parameter space so that leading catastrophic forgetting, or expand a new branch for each task but adding more computational cost. To this end, we propose MetrIc Learning with Expandable Subspace (Miles) to harness the prior information within pre-trained knowledge, thereby orchestrating an efficient expansion of the parameter space through guided optimization. Specifically, it decouples the learnable modules with the pre-trained model, exploiting prior information from intermediate features of the backbone network to enable more flexible parameter expansion. Then, a central loss is adopted to guide the new category to cluster towards the corresponding prototype in the new task subspace while incorporating an auxiliary distance regularization term to maintain metric equilibrium across tasks. Extensive experiments on six benchmark datasets demonstrate that Miles achieves state-of-the-art performance in various CIL settings. Zisong Lin, Hongyuan Zhang 0001, Xueru Bai, Xuelong Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Dust to Tower: Prior-Driven Coarse-to-Fine Photo-Realistic Scene Reconstruction From Sparse Uncalibrated ImagesabstractPhoto-realistic scene reconstruction from sparse-view, uncalibrated images is highly required in practice. Although some successes have been made, existing methods are either Sparse-View but require accurate camera parameters (i.e., intrinsic and extrinsic), or SfM-free but need densely captured images. This paper proposes Dust to Tower (D2T), a novel coarse-to-fine framework to address the coupled difficulty. The key idea is to explicitly narrow down the solution space and then introduce reliable supervision at novel viewpoints without resorting to expensive diffusion-based view synthesis. To do this, we first introduce a Coarse Construction Module (CCM) which exploits a fast Multi-View Stereo model to initialize a 3D Gaussian Splatting (3DGS) and recover initial camera poses. To refine the 3D model at novel viewpoints, we introduce Confidence-Aware Depth Alignment (CADA), which aligns a monocular inverse-depth prior to the reliable regions of the coarse depth using DUSt3R confidence, producing sharp and scale-consistent depth maps for accurate warping. We further propose Warped Image-Guided Inpainting (WIGI), which converts the accurate warped views into multi-view-consistent pseudo supervision via elaborate warping and inpainting process. Experiments on three benchmark datasets show that D2T achieves superior novel view synthesis quality and pose accuracy over ten representative baselines, while keeping high efficiency. Yongcai Wang, Zhaoxin Fan, Shuo Wang 0015, Deying Li 0001, Lun Luo, Minhang Wang, Hongyuan Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 10 |
| 2026 | ExGes: Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture SynthesisabstractAudio-driven human gesture synthesis is a crucial task with broad applications in virtual avatars, human-computer interaction, and creative content generation. Despite notable progress, existing methods often produce coarse gestures, lack expressiveness, and fail to fully align with audio semantics. To address these challenges, we propose ExGes, a novel retrieval-enhanced diffusion framework with three key designs: (1) a Motion Base Construction, which builds a gesture library from the training dataset; (2) a Motion Retrieval Module, employing contrastive learning and momentum distillation for retrieving fine-grained reference poses; and (3) a Precise Control Module, integrating partial masking and stochastic masking to enable flexible and fine-grained control. Experimental evaluations on BEAT2 demonstrate that ExGes reduces Fréchet Gesture Distance by 4.55%and improves motion diversity by 5.3% over EMAGE, with user studies revealing a 71.3% preference for its naturalness and semantic relevance. Xukun Zhou, Fengxin Li, Yan Zhou 0003, Pengfei Wan 0001, Yeying Jin, Hongyuan Zhang 0001, Hongyan Liu 0002, Zhaoxin Fan, Jun He 0008, Xuelong Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | Enhance Vision-Language Alignment with NoiseabstractWith the advancement of pre-trained vision-language (VL) models, enhancing the alignment between visual and linguistic modalities in downstream tasks has emerged as a critical challenge. Different from existing fine-tuning methods that add extra modules to these two modalities, we investigate whether the frozen model can be fine-tuned by customized noise. Our approach is motivated by the scientific study of beneficial noise, namely Positive-incentive Noise (Pi-noise) , which quantitatively analyzes the impact of noise. It therefore implies a new scheme to learn beneficial noise distribution that can be employed to fine-tune VL models. Focusing on few-shot classification tasks based on CLIP, we reformulate the inference process of CLIP and apply variational inference, demonstrating how to generate Pi-noise towards visual and linguistic modalities. Then, we propose Positive-incentive Noise Injector (PiNI), which can fine-tune CLIP via injecting noise into both visual and text encoders. Since the proposed method can learn the distribution of beneficial noise, we can obtain more diverse embeddings of vision and language to better align these two modalities for specific downstream tasks within limited computational resources. We evaluate different noise incorporation approaches and network architectures of PiNI. The evaluation across 11 datasets demonstrates its effectiveness. Sida Huang, Hongyuan Zhang 0001, Xuelong Li 0001 |
AAAI | 2 |
| 2025 | Why Does Dropping Edges Usually Outperform Adding Edges in Graph Contrastive Learning?abstractGraph contrastive learning (GCL) has been widely used as an effective self-supervised learning method for graph representation learning. However, how to apply adequate and stable graph augmentation to generating proper views for contrastive learning remains an essential problem. Dropping edges is a primary augmentation in GCL while adding edges is not a common method due to its unstable performance. To our best knowledge, there is no theoretical analysis to study why dropping edges usually outperforms adding edges. To answer this question, we introduce a new metric, namely Error Passing Rate (EPR), to quantify how a graph fits the network. Inspired by the theoretical conclusions and the idea of positive-incentive noise, we propose a novel GCL algorithm, Error-PAssing-based Graph Contrastive Learning (EPAGCL), which uses both edge adding and edge dropping as its augmentations. To be specific, we generate views by adding and dropping edges based on the weights derived from EPR. Extensive experiments on various real-world datasets are conducted to validate the correctness of our theoretical analysis and the effectiveness of our proposed algorithm. Yanchen Xu, Hongyuan Zhang 0001, Xuelong Li 0001 |
AAAI | 3 |
| 2025 | G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object ManipulationabstractRecent advances in imitation learning for 3D robotic manipulation have shown promising results with diffusion-based policies. However, achieving human-level dexterity requires seamless integration of geometric precision and semantic understanding. We present G3Flow, a novel framework that constructs real-time semantic flow, a dynamic, object-centric 3D semantic representation by leveraging foundation models. Our approach uniquely combines 3D generative models for digital twin creation, vision foundation models for semantic feature extraction, and robust pose tracking for continuous semantic flow updates. This integration enables complete semantic understanding even under occlusions while eliminating manual annotation requirements. By incorporating semantic flow into diffusion policies, extensive experiments across five simulation tasks show that G3Flow consistently outperforms existing approaches, achieving up to 68.3% and 50.1% success rates on terminal-constrained manipulation and cross-object generalization respectively. Our results demonstrate the effectiveness of G3Flow in enhancing real-time dynamic semantic feature understanding for robotic policies. Tianxing Chen, Yao Mu 0001, Zhixuan Liang, Zanxin Chen, Shijia Peng, Qiangyu Chen, Mingkun Xu, Ruizhen Hu, Hongyuan Zhang 0001, Xuelong Li 0001, Ping Luo 0002 |
CVPR | 9 |
| 2025 | Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial AttacksabstractRecent Customized Portrait Generation (CPG) methods, taking a facial image and a textual prompt as inputs, have attracted substantial attention. Although these methods generate high-fidelity portraits, they fail to prevent the generated portraits from being tracked and misused by malicious face recognition systems. To address this, this paper proposes a Customized Portrait Generation framework with facial Adversarial attacks (Adv-Cpg). Specifically, to achieve facial privacy protection, we devise a lightweight local ID encryptor and an encryption enhancer. They implement progressive double-layer encryption protection by directly injecting the target identity and adding additional identity guidance, respectively. Furthermore, to accomplish fine-grained and personalized portrait generation, we develop a multi-modal image customizer capable of generating controlled fine-grained facial features. To the best of our knowledge, Adv-Cpg is the first study that introduces facial adversarial attacks into CPG. Extensive experiments demonstrate the superiority of Adv-Cpg, e.g., the average attack success rate of the proposed Adv-Cpg is 28.1% and 2.86% higher compared to the SOTA noise-based attack methods and unconstrained attack methods, respectively. Hongyuan Zhang 0001, Yuan Yuan 0001 |
CVPR | 2 |
| 2025 | Growing a Twig to Accelerate Large Vision-Language Models
Zhenwei Shao, Zhou Yu 0001, Wenwen Pan 0003, Hongyuan Zhang 0001, Wei Chen 0001, Jun Yu 0002 |
ICCV | 7 |
| 2025 | Learn Beneficial Noise as Graph AugmentationabstractAlthough graph contrastive learning (GCL) has been widely investigated, it is still a challenge to generate effective and stable graph augmentations. Existing methods often apply heuristic augmentation like random edge dropping, which may disrupt important graph structures and result in unstable GCL performance. In this paper, we propose **P**ositive-**i**ncentive **N**oise driven **G**raph **D**ata **A**ugmentation (PiNGDA), where positive-incentive noise (pi-noise) scientifically analyzes the beneficial effect of noise under the information theory. To bridge the standard GCL and pi-noise framework, we design a Gaussian auxiliary variable to convert the loss function to information entropy. We prove that the standard GCL with pre-defined augmentations is equivalent to estimate the beneficial noise via the point estimation. Following our analysis, PiNGDA is derived from learning the beneficial noise on both topology and attributes through a trainable noise generator for graph augmentations, instead of the simple estimation. Since the generator learns how to produce beneficial perturbations on graph topology and node attributes, PiNGDA is more reliable compared with the existing methods. Extensive experimental results validate the effectiveness and stability of PiNGDA. Yanchen Xu, Hongyuan Zhang 0001, Xuelong Li 0001 |
ICML | 3 |
| 2025 | NFIG: Multi-Scale Autoregressive Image Generation via Frequency OrderingabstractAutoregressive models have achieved significant success in image generation. However, unlike the inherent hierarchical structure of image information in the spectral domain, standard autoregressive methods typically generate pixels sequentially in a fixed spatial order. To better leverage this spectral hierarchy, we introduce Next-Frequency Image Generation (NFIG). NFIG is a novel framework that decomposes the image generation process into multiple frequency-guided stages. NFIG aligns the generation process with the natural image structure. It does this by first generating low-frequency components, which efficiently capture global structure with significantly fewer tokens, and then progressively adding higher-frequency details. This frequency-aware paradigm offers substantial advantages: it not only improves the quality of generated images but crucially reduces inference cost by efficiently establishing global structure early on. Extensive experiments on the ImageNet-256 benchmark validate NFIG's effectiveness, demonstrating superior performance (FID: 2.81) and a notable 1.25x speedup compared to the strong baseline VAR-d20. Xi Qiu, Yukuo Ma, Yifu Zhou, Hongyuan Zhang 0001, Chi Zhang 0012, Xuelong Li 0001 |
NeurIPS | 6 |
| 2025 | Mixture of Noise for Pre-Trained Model-Based Class-Incremental LearningabstractClass Incremental Learning (CIL) aims to continuously learn new categories while retaining the knowledge of old ones. Pre-trained models (PTMs) show promising capabilities in CIL. However, existing approaches that apply lightweight fine-tuning to backbones still induce parameter drift, thereby compromising the generalization capability of pre-trained models. Parameter drift can be conceptualized as a form of noise that obscures critical patterns learned for previous tasks. However, recent researches have shown that noise is not always harmful. For example, the large number of visual patterns learned from pre-training can be easily abused by a single task, and introducing appropriate noise can suppress some low-correlation features, thus leaving a margin for future tasks. To this end, we propose learning beneficial noise for CIL guided by information theory and propose Mixture of Noise (MiN), aiming to mitigate the degradation of backbone generalization from adapting new tasks. Specifically, task-specific noise is learned from high-dimension features of new tasks. Then, a set of weights is adjusted dynamically for optimal mixture of different task noise. Finally, MiN embeds the beneficial noise into the intermediate features to mask the response of inefficient patterns. Extensive experiments on six benchmark datasets demonstrate that MiN achieves state-of-the-art performance in most incremental settings, with particularly outstanding results in 50-steps incremental settings. This shows the significant potential for beneficial noise in continual learning. Code is available at https://github.com/ASCIIJK/MiN-NeurIPS2025. Zhengyan Shi, Dell Zhang, Hongyuan Zhang 0001, Xuelong Li 0001 |
NeurIPS | 4 |
| 2025 | CNN2GNN: How to Bridge CNN With GNNabstractThanks to extracting the intra-sample representation, the convolution neural network (CNN) has achieved excellent performance in vision tasks. However, its numerous convolutional layers take a higher training expense. Recently, graph neural networks (GNN), a bilinear model, have succeeded in exploring the underlying topological relationship among the graph data with a few graph neural layers. Unfortunately, due to the lack of graph structure and high-cost inference on large-scale scenarios, it cannot be directly utilized on non-graph data. Inspired by these complementary strengths and weaknesses, we discuss a natural question, how to bridge these two heterogeneous networks? In this paper, we propose a novel CNN2GNN framework to unify CNN and GNN together via distillation. First, to break the limitations of GNN, we design a differentiable sparse graph learning module as the head of the networks. It can dynamically learn the graph for inductive learning. Then, a response-based distillation is introduced to transfer the knowledge and bridge these two heterogeneous networks. Notably, due to extracting the intra-sample representation of a single instance and the topological relationship among the datasets simultaneously, the performance of the distilled "boosted" two-layer GNN on Mini-ImageNet is much higher than CNN containing dozens of layers, such as ResNet152. Ziheng Jiao, Hongyuan Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Variational Positive-Incentive Noise: How Noise Benefits ModelsabstractA large number of works aim to alleviate the impact of noise due to an underlying conventional assumption of the negative role of noise. However, some existing works show that the assumption does not always hold. In this paper, we investigate how to benefit the classical models by random noise under the framework of Positive-incentive Noise (Pi-Noise) (Li, 2024). Since the ideal objective of Pi-Noise is intractable, we propose to optimize its variational bound instead, namely variational Pi-Noise (VPN). With the variational inference, a VPN generator implemented by neural networks is designed for enhancing base models and simplifying the inference of base models, without changing the architecture of base models. Benefiting from the independent design of base models and VPN generators, the VPN generator can work with most existing models. From the extensive experiments on different base models (including linear models, ResNet, ViT, etc.) it is shown that the proposed VPN generator can improve the base models. It is appealing that the trained VPN generator prefers to blur the irrelevant ingredients in complicated images, which meets our expectations. Hongyuan Zhang 0001, Sida Huang, Yubin Guo, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | AnchorFormer: Differentiable anchor attention for efficient vision transformer
Jiquan Shan, Lifeng Zhao, Hongyuan Zhang 0001, Ioannis Liritzis |
Pattern Recognit. Lett. | 5 |
| 2025 | SignEye: Traffic Sign Interpretation From Vehicle First-Person ViewabstractTraffic signs play a key role in assisting autonomous driving systems (ADS) by enabling the assessment of vehicle behavior in compliance with traffic regulations and providing navigation instructions. However, current works are limited to basic sign understanding without considering the egocentric vehicle’s spatial position, which fails to support further regulation assessment and direction navigation. Following the above issues, we introduce a new task: traffic sign interpretation from the vehicle’s first-person view, referred to asTSI-FPV. Meanwhile, we develop a traffic guidance assistant (TGA) scenario application to re-explore the role of traffic signs in ADS as a complement to popular autonomous technologies (such as obstacle perception). Notably, TGA is not a replacement for electronic map navigation; rather, TGA can be an automatic tool for updating it and complementing it in situations such as offline conditions or temporary sign adjustments. Lastly, a spatial and semantic logic-aware stepwise reasoning pipeline (SignEye) is constructed to achieve the TSI-FPV and TGA, and an application-specific dataset (Traffic-CN) is built. Experiments show that TSI-FPV and TGA are achievable via our SignEye trained on Traffic-CN. The results also demonstrate that the TGA can provide complementary information to ADS beyond existing popular autonomous technologies. Chuang Yang 0003, Xu Han 0019, Tao Han 0002, Yuejiao Su, Junyu Gao 0001, Hongyuan Zhang 0001, Yi Wang 0068, Lap-Pui Chau |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Deep Graph Multi-View Representation Learning With Self-Augmented View FusionabstractSome current researchers attempt to extend the graph neural network (GNN) on multi-view representation learning and learn the latent structure information among the data. Generally, they concatenate the features of each view and employ a single GNN to extract the representations of this concatenated feature. It causes that the within-view information may not be learned and the pivotal view will not be strengthened during the concatenation. Although some GNN models introduce the Siamese structure to extract the within-view information, the learned representation may not be informative since the Siamese GNNs share the same parameters. To overcome these issues, we propose a novel deep graph auto-encoder for multi-view representation learning. Among them, a self-augmented view-weight technique is theoretically devised for cross-view fusion, which can highlight the pivotal views and maintain the rest views. Then, GNNs of different views can learn the informative representation without sharing parameters. Furthermore, by fitting the fusion distribution with a neural layer, the model unifies these two individual procedures and achieve to extract the fusion representation end-to-end. Compared with numerous recently proposed methods, extensive experiments on clustering and recognition tasks demonstrate our superior performance. Ziheng Jiao, Hongyuan Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Graph Convolutional Network With Self-Augmented Weights for Semi-Supervised Multi-View LearningabstractRecently, owing to the effectiveness in exploiting inherent connections between data in different views, graph-based deep learning approaches have gained widespread popularity in semi-supervised multi-view tasks. Generally, the existing approaches fuse the information from different views via the linear or nonlinear weight strategies, which distinguish the importance of different views by attributing their weights between $[{0, 1}]$ , i.e., some less important views are discarded since assigned with 0 and the pivotal views are not enhanced. However, these view-weighting strategies ignore the complementary information from the less important views. To address this issue, a superior-performing graph convolutional network (GCN) with self-augmented weights is proposed. The proposed self-augmented weight strategy is based on exponential series integration, which preserves the less important views and simultaneously strengthens the key views for multi-view fusion. Specifically, the designed weight strategy can adaptively preserve the complementary information from the less important views by assigning nonzero weights and strengthen the pivotal views by assigning higher weights based on exponential series integration. Besides, to further improve the model performance, an orthogonal constraint layer with a forced orthogonal weight is introduced, which is capable of making the representation more discriminative. Extensive experiments demonstrate the superiority of the proposed method. Hongyuan Zhang 0001, Yuan Yuan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | CNN to GNN: Unsupervised Multi-level Knowledge LearningabstractAlthough graph neural networks (GNNs) can extract the latent relationship-level knowledge among the graph nodes and have achieved excellent performance in unsupervised scenarios, it is weak in learning the instance-level knowledge in contrast to the convolution neural networks (CNNs). Besides, lacking of the graph structure limits the extension of GNNs on non-graph datasets. To solve these problems, we propose a novel unsupervised multi-level knowledge fusion network. It successfully unifies the instance-level and relationship-level knowledge on the non-graph data by distillation from a pre-trained CNN teacher to a GNN student. Meanwhile, a sparse weighted strategy is designed to adaptively extract the sparse graph topology and extend the GNN on non-graph datasets. By optimization of distillation loss, the "boosted'' GNN student can learn the multi-level knowledge and extract more discriminative deep embeddings for clustering. Finally, extensive experiments show it has achieved excellent performance compared with the current methods. Ziheng Jiao, Hongyuan Zhang 0001, Xuelong Li 0001 |
CIKM | 2 |
| 2024 | Graph manifold learning with non-gradient decision layer
Ziheng Jiao, Hongyuan Zhang 0001, Rui Zhang 0017, Xuelong Li 0001 |
Neurocomputing | 2 |
| 2024 | Decouple Graph Neural Networks: Train Multiple Simple GNNs Simultaneously Instead of OneabstractGraph neural networks (GNN) suffer from severe inefficiency due to the exponential growth of node dependency with the increase of layers. It extremely limits the application of stochastic optimization algorithms so that the training of GNN is usually time-consuming. To address this problem, we propose to decouple a multi-layer GNN as multiple simple modules for more efficient training, which is comprised of classical forward training (FT) and designed backward training (BT). Under the proposed framework, each module can be trained efficiently in FT by stochastic algorithms without distortion of graph information owing to its simplicity. To avoid the only unidirectional information delivery of FT and sufficiently train shallow modules with the deeper ones, we develop a backward training mechanism that makes the former modules perceive the latter modules, inspired by the classical backward propagation algorithm. The backward training introduces the reversed information delivery into the decoupled modules as well as the forward information delivery. To investigate how the decoupling and greedy training affect the representational capacity, we theoretically prove that the error produced by linear modules will not accumulate on unsupervised tasks in most cases. The theoretical and experimental results show that the proposed framework is highly efficient with reasonable performance, which may deserve more investigation. Hongyuan Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Orthogonal subspace exploration for matrix completion
Hongyuan Zhang 0001, Ziheng Jiao, Xuelong Li 0001 |
Pattern Recognit. | 1 |
| 2024 | Pivotal-Aware Principal Component AnalysisabstractA conventional principal component analysis (PCA) frequently suffers from the disturbance of outliers, and thus, spectra of extensions and variations of PCA have been developed. However, all the existing extensions of PCA derive from the same motivation, which aims to alleviate the negative effect of the occlusion. In this article, we design a novel collaborative-enhanced learning framework that aims to highlight the pivotal data points in contrast. As for the proposed framework, only a part of well-fitting samples are adaptively highlighted, which indicates more significance during training. Meanwhile, the framework can collaboratively reduce the disturbance of the polluted samples as well. In other words, two contrary mechanisms could work cooperatively under the proposed framework. Based on the proposed framework, we further develop a pivotal-aware PCA (PAPCA), which utilizes the framework to simultaneously augment positive samples and constrain negative ones by retaining the rotational invariance property. Accordingly, extensive experiments demonstrate that our model has superior performance compared with the existing methods that only focus on the negative samples. Xuelong Li 0001, Hongyuan Zhang 0001, Kangjia Zhu, Rui Zhang 0017 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Discretize Relaxed Solution of Spectral Clustering via a Nonheuristic AlgorithmabstractSpectral clustering and its extensions usually consist of two steps: 1) constructing a graph and computing the relaxed solution and 2) discretizing relaxed solutions. Although the former has been extensively investigated, the discretization techniques are mainly heuristic methods, e.g., -means (KM), spectral rotation (SR). Unfortunately, the goal of the existing methods is not to find a discrete solution that minimizes the original objective. In other words, the primary drawback is the neglect of the original objective when computing the discrete solution. Inspired by the first-order optimization algorithms, we propose to develop a first-order term to bridge the original problem and discretization algorithm, which is the first nonheuristic to the best of our knowledge. Since the nonheuristic method is aware of the original graph cut problem, the final discrete solution is more reliable and achieves the preferable loss value. We also theoretically show that the continuous optimum is beneficial to discretization algorithms though simply finding its closest discrete solution is an existing heuristic algorithm which is also unreliable. Sufficient experiments significantly show the superiority of our method. Hongyuan Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Learn Topological Representation with Flexible Manifold LayerabstractDeep neural networks for classification are one of the most fundamental topics in machine learning. However, the widely used softmax decision layer, originating from the conventional softmax regression for multi-class classification, fails to guide the powerful feature extractor to explore the topological structure hidden in data, which limits the quality of produced representations. Therefore, we propose a flexible manifold layer for better representation learning in this paper, rather than adding some regularized losses to introduce extra mechanisms. The flexible manifold layer is inspired by SVM which maximizes the geometric margin and usually achieves better results than softmax regression. Since the potential structure is captured due to the explicit guidance of the proposed flexible manifold layer, the noise can be easily detected and filtered out via a designed automatic mechanism. In addition to passing the topological knowledge by latent representations, a knowledge distillation from the manifold layer and decision layer is developed. Extensive experiments show that the proposed model has achieved excellent performance. The code is available at https://github.com/jzh9830/DeepSVM. Ziheng Jiao, Hongyuan Zhang 0001, Xuelong Li 0001 |
ICASSP | 2 |
| 2023 | Matrix Completion via Non-Convex Relaxation and Adaptive Correlation LearningabstractThe existing matrix completion methods focus on optimizing the relaxation of rank function such as nuclear norm, Schatten- p norm, etc. They usually need many iterations to converge. Moreover, only the low-rank property of matrices is utilized in most existing models and several methods that incorporate other knowledge are quite time-consuming in practice. To address these issues, we propose a novel non-convex surrogate that can be optimized by closed-form solutions, such that it empirically converges within dozens of iterations. Besides, the optimization is parameter-free and the convergence is proved. Compared with the relaxation of rank, the surrogate is motivated by optimizing an upper-bound of rank. We theoretically validate that it is equivalent to the existing matrix completion models. Besides the low-rank assumption, we intend to exploit the column-wise correlation for matrix completion, and thus an adaptive correlation learning, which is scaling-invariant, is developed. More importantly, after incorporating the correlation learning, the model can be still solved by closed-form solutions such that it still converges fast. Experiments show the effectiveness of the non-convex surrogate and adaptive correlation learning. Xuelong Li 0001, Hongyuan Zhang 0001, Rui Zhang 0017 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Manifold Neural Network With Non-Gradient OptimizationabstractDeep neural network (DNN) generally takes thousands of iterations to optimize via gradient descent and thus has a slow convergence. In addition, softmax, as a decision layer, may ignore the distribution information of the data during classification. Aiming to tackle the referred problems, we propose a novel manifold neural network based on non-gradient optimization, i.e., the analytical-form solutions. Considering that the activation function is generally invertible, we reconstruct the network via forward ridge regression and low-rank backward approximation, which achieve rapid convergence. Moreover, by unifying the flexible Stiefel manifold and adaptive support vector machine, we devise the novel decision layer which efficiently fits the manifold structure of the data and label information. Consequently, a jointly non-gradient optimization method is designed to generate the network with analytical-form results. Furthermore, an acceleration strategy is utilize to reduce the time complexity for handling high dimensional datasets. Eventually, extensive experiments validate the superior performance of the model. Rui Zhang 0017, Ziheng Jiao, Hongyuan Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Non-Graph Data Clustering via $\mathcal {O}(n)$O(n) Bipartite Graph ConvolutionabstractSince the representative capacity of graph-based clustering methods is usually limited by the graph constructed on the original features, it is attractive to find whether graph neural networks (GNNs), a strong extension of neural networks to graphs, can be applied to augment the capacity of graph-based clustering methods. The core problems mainly come from two aspects. On the one hand, the graph is unavailable in the most general clustering scenes so that how to construct graph on the non-graph data and the quality of graph is usually the most important part. On the other hand, given$n$samples, the graph-based clustering methods usually consume at least$\mathcal {O}(n^{2})$time to build graphs and the graph convolution requires nearly$\mathcal {O}(n^{2})$for a dense graph and$\mathcal {O}(|\mathcal {E}|)$for a sparse one with$|\mathcal {E}|$edges. Accordingly, both graph-based clustering and GNNs suffer from the severe inefficiency problem. To tackle these problems, we propose a novel clustering method,AnchorGAE, with the self-supervised estimation of graph and efficient graph convolution. We first show how to convert a non-graph dataset into a graph dataset, by introducing the generative graph model and anchors. A bipartite graph is built via generating anchors and estimating the connectivity distributions of original points and anchors. We then show that the constructed bipartite graph can reduce the computational complexity of graph convolution from$\mathcal {O}(n^{2})$and$\mathcal {O}(|\mathcal {E}|)$to$\mathcal {O}(n)$. The succeeding steps for clustering can be easily designed as$\mathcal {O}(n)$operations. Interestingly, the anchors naturally lead to siamese architecture with the help of the Markov process. Furthermore, the estimated bipartite graph is updated dynamically according to the features extracted by GNN modules, to promote the quality of the graph by exploiting the high-level information by GNNs. However, we theoretically prove that the self-supervised paradigm frequently results in a collapse that often occurs after 2-3 update iterations in experiments, especially when the model is well-trained. A specific strategy is accordingly designed to prevent the collapse. The experiments support the theoretical analysis and show the superiority of AnchorGAE. Hongyuan Zhang 0001, Jiankun Shi, Rui Zhang 0017, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Toward Projected Clustering With Aggregated MappingabstractProjected clustering is the foundation of deep clustering models. Aiming at catching the essence of deep clustering, we propose a novel projected clustering framework by summarizing the core properties of prevalent powerful models, especially deep models. At first, we introduce the aggregated mapping, consisting of projection learning and neighbor estimation, to obtain clustering-friendly representation. Importantly, we theoretically prove that the simple clustering-friendly representation learning may suffer from severe degeneration, which can be regarded as over-fitting. Roughly speaking, the well-trained model would group neighboring points into plenty of sub-clusters. These small sub-clusters may scatter randomly due to no connection between them. The degeneration may occur more frequently with the increasing of model capacity. We accordingly develop a self-evolution mechanism that implicitly aggregates the sub-clusters and the proposed method can alleviate the potential risk of over-fitting and obtain prominent improvement. The ablation experiments support the theoretical analysis and verify the effectiveness of the neighbor-aggregation mechanism. Finally, we show how to choose the unsupervised projection function through two specific examples, including a linear method (namely locality analysis) and a non-linear model. Hongyuan Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Embedding Graph Auto-Encoder for Graph ClusteringabstractGraph clustering, aiming to partition nodes of a graph into various groups via an unsupervised approach, is an attractive topic in recent years. To improve the representative ability, several graph auto-encoder (GAE) models, which are based on semisupervised graph convolution networks (GCN), have been developed and they have achieved impressive results compared with traditional clustering methods. However, all existing methods either fail to utilize the orthogonal property of the representations generated by GAE or separate the clustering and the training of neural networks. We first prove that the relaxed k -means will obtain an optimal partition in the inner-product distance used space. Driven by theoretical analysis about relaxed k -means, we design a specific GAE-based model for graph clustering to be consistent with the theory, namely Embedding GAE (EGAE). The learned representations are well explainable so that the representations can be also used for other tasks. To induce the neural network to produce deep features that are appropriate for the specific clustering model, the relaxed k -means and GAE are learned simultaneously. Meanwhile, the relaxed k -means can be equivalently regarded as a decoder that attempts to learn representations that can be linearly constructed by some centroid vectors. Accordingly, EGAE consists of one encoder and dual decoders. Extensive experiments are conducted to prove the superiority of EGAE and the corresponding theoretical analyses. Hongyuan Zhang 0001, Rui Zhang 0017, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Adaptive Graph Auto-Encoder for General Data ClusteringabstractGraph-based clustering plays an important role in the clustering area. Recent studies about graph neural networks (GNN) have achieved impressive success on graph-type data. However, in general clustering tasks, the graph structure of data does not exist such that GNN can not be applied to clustering directly and the strategy to construct a graph is crucial for performance. Therefore, how to extend GNN into general clustering tasks is an attractive problem. In this paper, we propose a graph auto-encoder for general data clustering, AdaGAE, which constructs the graph adaptively according to the generative perspective of graphs. The adaptive process is designed to induce the model to exploit the high-level information behind data and utilize the non-euclidean structure sufficiently. Importantly, we find that the simple update of the graph will result in severe degeneration, which can be concluded as better reconstruction means worse update. We provide rigorous analysis theoretically and empirically. Then we further design a novel mechanism to avoid the collapse. Via extending the generative graph models to general type data, a graph auto-encoder with a novel decoder is devised and the weighted graphs can be also applied to GNN. AdaGAE performs well and stably in different scale and type datasets. Besides, it is insensitive to the initialization of parameters and requires no pretraining. Xuelong Li 0001, Hongyuan Zhang 0001, Rui Zhang 0017 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Hierarchical Deep Click Feature Prediction for Fine-Grained Image RecognitionabstractThe click feature of an image, defined as the user click frequency vector of the image on a predefined word vocabulary, is known to effectively reduce the semantic gap for fine-grained image recognition. Unfortunately, user click frequency data are usually absent in practice. It remains challenging to predict the click feature from the visual feature, because the user click frequency vector of an image is always noisy and sparse. In this paper, we devise a Hierarchical Deep Word Embedding (HDWE) model by integrating sparse constraints and an improved RELU operator to address click feature prediction from visual features. HDWE is a coarse-to-fine click feature predictor that is learned with the help of an auxiliary image dataset containing click information. It can therefore discover the hierarchy of word semantics. We evaluate HDWE on three dog and one bird image datasets, in which Clickture-Dog and Clickture-Bird are utilized as auxiliary datasets to provide click data, respectively. Our empirical studies show that HDWE has 1) higher recognition accuracy, 2) a larger compression ratio, and 3) good one-shot learning ability and scalability to unseen categories. Jun Yu 0002, Min Tan 0005, Hongyuan Zhang 0001, Yong Rui, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Maximum Joint Probability With Multiple Representations for ClusteringabstractClassical generative models in unsupervised learning intend to maximize p(X) . In practice, samples may have multiple representations caused by various transformations, measurements, and so on. Therefore, it is crucial to integrate information from different representations, and lots of models have been developed. However, most of them fail to incorporate the prior information about data distribution p(X) to distinguish representations. In this article, we propose a novel clustering framework that attempts to maximize the joint probability of data and parameters. Under this framework, the prior distribution can be employed to measure the rationality of diverse representations. K -means is a special case of the proposed framework. Meanwhile, a specific clustering model considering both multiple kernels and multiple views is derived to verify the validity of the designed framework and model. Rui Zhang 0017, Hongyuan Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Unsupervised Feature Selection With Extended OLSDA via Embedding Nonnegative Manifold StructureabstractAs to unsupervised learning, most discriminative information is encoded in the cluster labels. To obtain the pseudo labels, unsupervised feature selection methods usually utilize spectral clustering to generate them. Nonetheless, two related disadvantages exist accordingly: 1) the performance of feature selection highly depends on the constructed Laplacian matrix and 2) the pseudo labels are obtained with mixed signs, while the real ones should be nonnegative. To address this problem, a novel approach for unsupervised feature selection is proposed by extending orthogonal least square discriminant analysis (OLSDA) to the unsupervised case, such that nonnegative pseudo labels can be achieved. Additionally, an orthogonal constraint is imposed on the class indicator to hold the manifold structure. Furthermore,$\ell _{2,1}$regularization is imposed to ensure that the projection matrix is row sparse for efficient feature selection and proved to be equivalent to$\ell _{2,0}$regularization. Finally, extensive experiments on nine benchmark data sets are conducted to demonstrate the effectiveness of the proposed approach. Rui Zhang 0017, Hongyuan Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Robust multi-view fuzzy clustering via softmin
Hongyuan Zhang 0001, Rui Zhang 0017, Xuelong Li 0001, Yueshen Xu |
Neurocomputing | 1 |
| 2021 | Robust Multi-Task Learning With Flexible Manifold ConstraintabstractMulti-Task Learning attempts to explore and mine the sufficient information within multiple related tasks for the better solutions. However, the performance of the existing multi-task approaches would largely degenerate when dealing with the polluted data, i.e., outliers. In this paper, we propose a novel robust multi-task model by incorporating a flexible manifold constraint (FMC-MTL) and a robust loss. Specifically speaking, multi-task subspace is embedded with a relaxed and generalized Stiefel Manifold for considering point-wise correlation and preserving the data structure simultaneously. In addition, a robust loss function is developed to ensure the robustness to outliers by smoothly interpolating betweenl2,1ℓ2,1-norm and squared Frobenius norm. Equipped with an efficient algorithm, FMC-MTL serves as a robust solution to tackling the severely polluted data. Moreover, extensive experiments are conducted to verify the superiority of our model. Compared to the state-of-the-art multi-task models, the proposed FMC-MTL model demonstrates remarkable robustness to the contaminated data. Rui Zhang 0017, Hongyuan Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Autoencoder Constrained Clustering With Adaptive NeighborsabstractThe conventional subspace clustering method obtains explicit data representation that captures the global structure of data and clusters via the associated subspace. However, due to the limitation of intrinsic linearity and fixed structure, the advantages of prior structure are limited. To address this problem, in this brief, we embed the structured graph learning with adaptive neighbors into the deep autoencoder networks such that an adaptive deep clustering approach, namely, autoencoder constrained clustering with adaptive neighbors (ACC_AN), is developed. The proposed method not only can adaptively investigate the nonlinear structure of data via a parameter-free graph built upon deep features but also can iteratively strengthen the correlations among the deep representations in the learning process. In addition, the local structure of raw data is preserved by minimizing the reconstruction error. Compared to the state-of-the-art works, ACC_AN is the first deep clustering method embedded with the adaptive structured graph learning to update the latent representation of data and structured deep graph simultaneously. Xuelong Li 0001, Rui Zhang 0017, Qi Wang 0009, Hongyuan Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Deep Fuzzy K-Means With Adaptive Loss and Entropy RegularizationabstractNeural network based clustering methods usually have better performance compared to the conventional approaches due to more efficient feature extraction. Most of existing deep clustering techniques either exploit graph information as prior to extract pivotal deep structure from the raw data and simply utilizes stochastic gradient descent (SGD). However, they often suffer from separating the learning steps regarding dimensionality reduction and clustering. To address these issues, a novel deep model named as deep fuzzy k-means (DFKM) with adaptive loss function and entropy regularization is proposed. DFKM performs deep feature extraction and fuzzy clustering simultaneously to generate a more appropriate nonlinear feature map. Additionally, DFKM incorporates FKM so that fuzzy information is utilized to represent a clear structure of deep clusters. To further promote the robustness of the model, a robust loss function is applied to the objective with adaptive weights. Moreover, an entropy regularization is employed for affinity to provide confidence of each assignment and the corresponding membership and centroid matrices are updated by close-form solutions rather than SGD. Extensive experiments show that DFKM has better performance compared to the state-of-the-art fuzzy clustering techniques under three clustering metrics. Rui Zhang 0017, Xuelong Li 0001, Hongyuan Zhang 0001, Feiping Nie 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2019 | Image Recognition by Predicted User Click Feature With Multidomain Multitask Transfer Deep NetworkabstractThe click feature of an image, defined as a user click count vector based on click data, has been demonstrated to be effective for reducing the semantic gap for image recognition. Unfortunately, most of the traditional image recognition datasets do not contain click data. To address this problem, researchers have begun to develop a click prediction model using assistant datasets containing click information and have adapted this predictor to a common click-free dataset for different tasks. This method can be customized to our problem, but it has two main limitations: 1) the predicted click feature often performs badly in the recognition task since the prediction model is constructed independently of the subsequent recognition problem and 2) transferring the predictor from one dataset to another is challenging due to the large cross-domain diversity. In this paper, we devise a multitask and multidomain deep network with varied modals (MTMDD-VM) to formulate image recognition and click prediction tasks in a unified framework. Datasets with and without click information are integrated in the training. Furthermore, a nonlinear word embedding with a position-sensitive loss function is designed to discover the visual click correlation. We evaluate the proposed method on three public dog breed image datasets, and we utilize the Clickture-Dog dataset as the auxiliary dataset that provides click data. The experimental results show that: 1) the nonlinear word embedding and position-sensitive loss function largely enhance the predicted click feature in the recognition task, realizing a 32% improvement in accuracy; 2) the multitask learning framework improves accuracies in both image recognition and click prediction; and 3) the unified training using the combined dataset with and without click data further improves the performance. Compared with the state-of-the-art methods, the proposed approach not only performs much better in accuracy but also achieves good scalability and one-shot learning ability. Min Tan 0005, Jun Yu 0002, Hongyuan Zhang 0001, Yong Rui, Dacheng Tao |
IEEE Trans. Image Process. | 3 |