VLDB 2026 Research / reviewers in the wild / expert
Jiaqi Jin
dblp:198/6163
· DBLP profile ↗
30ranked-venue papers
6as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DMCAR: Disentangled Mixture-of-Experts with Context-Aware Routing for Multi-View ClusteringabstractMulti-View Clustering (MVC) aims to enhance clustering performance by integrating multi-source complementary information. However, existing deep MVC methods face inherent challenges in balancing the learning of shared consensus representations with the preservation of view-specific information: independent encoders hinder effective cross-view collaboration, while a single shared encoder tends to sacrifice representation diversity. Although the recently introduced Mixture-of-Experts (MoE) model offers a novel approach to facilitating view collaboration, its flattened expert pool design often leads to entanglement between shared and specific information, and its routing mechanism limits collaboration potential by neglecting cross-view context. To address these challenges, this paper proposes a novel deep multi-view clustering framework—Decoupled Mixture-of-Experts with Context-Aware Routing for Multi-View Clustering (DMCAR-MVC). At its core is an innovative Decoupled MoE (D-MoE) architecture. We establish a public expert pool to learn cross-view shared representations while equipping each view with an independent private expert pool to capture its unique information, thereby structurally enforcing the decoupling of shared and specific representations. Building on this, we further design a Context-Aware Hierarchical Routing (CAHR) mechanism. When routing for the public expert pool, this mechanism introduces a global context vector to guide expert selection, enabling more efficient and globally informed cross-view collaboration. Finally, to optimize the model, we adopt a multi-level contrastive learning paradigm: on one hand, a cross-view alignment loss ensures semantic consistency in shared representations; on the other, an orthogonality constraint is imposed to further enhance separability between shared and specific representations. Extensive experiments on multiple benchmark datasets demonstrate that DMCAR-MVC significantly outperforms state-of-the-art methods across key clustering metrics. Additionally, comprehensive ablation studies thoroughly validate the effectiveness and necessity of each proposed component. Baili Xiao, Ke Liang 0006, Jiaqi Jin, Jun Wang 0118, Yinbo Xu, Siwei Wang 0001, En Zhu |
AAAI | 3 |
| 2026 | Hierarchical Cross-View Alignment for Multi-View Clustering via Decoupled Information DistillationabstractMulti-view clustering aims to uncover shared semantics and complementary information across different views. However, the inherent heterogeneity among views poses significant challenges to effective collaborative modeling and information integration. While recent studies have introduced distillation-based mechanisms to enhance cross-view consistency and alleviate heterogeneity, these approaches often rely on manually defined knowledge transfer paths or fixed fusion weights, which are inflexible in handling complex and dynamic view relationships in practice. To address this issue, we propose HOARD: a novel framework for Hierarchical crOss-view Alignment for multi-view clusteRing via Decoupled information distillation. HOARD structurally decouples multi-view representations into shared and specific components, and performs hierarchical alignment. Specifically, we introduce a granular-ball contrastive alignment to enhance the semantic consistency of shared features, and a prototype collaborative transmission alignment strategy to align specific features while preserving view-specific structural characteristics. Moreover, we design an information distillation unit to adaptively model cross-view knowledge transfer in both feature spaces. An attention mechanism is further employed to integrate shared and specific information. Extensive experiments on benchmark datasets demonstrate that HOARD significantly improves alignment quality and clustering performance, achieving state-of-the-art results. Taichun Zhou, Siwei Wang 0001, Zhibin Dong, Jiaqi Jin, Ke Liang 0006, Baili Xiao, Miaomiao Li 0001, Xinwang Liu 0002, En Zhu |
AAAI | 4 |
| 2026 | LRGFormer: A Multiscale Feature Fusion Transformer for Image RestorationabstractAdverse weather conditions can significantly degrade image quality and impair the capture of critical information. Existing restoration networks struggle to effectively combine local, regional, and global features, thereby limiting their ability to handle diverse impacts of such weather. This study proposes the local-region-global transformer (LRGFormer), a transformer-based image restoration model for multiscale feature perception. The model comprises a basic module composed of multi-scale fusion attention (MSFSA) and a channel-spatial dual-attention feed-forward network (CSDF). Specifically, this study designs an MSFSA module. For the first time, it combines rotation-equivariant convolution with local attention for local information extraction and introduces a frequency-domain adaptive attention mechanism. By incorporating a query-aware global adaptive sparse attention mechanism for global information extraction, the network gradually fuses along the channel dimension, enabling progressive capture of spatial and frequency-domain information from the local and regional to global scale. Secondly, a CSDF network structure was designed to enhance channel-spatial interaction and improve the representational capacity of the model. By constructing a basic U-Net framework, the excellent basic modules for image restoration proposed in recent years are compared on a unified framework. Experimental results demonstrated that the proposed basic module can not only better extracts multi-scale features of images and restores image distortion caused by various degradation factors, and also exhibits good universality and generalization. Jiafeng Li 0001, Wanying Hu, Tongyao Jia, Jiaqi Jin, Jing Zhang 0023, Li Zhuo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Imputation-free and Alignment-free: Incomplete Multi-view Clustering Driven by Consensus Semantic LearningabstractIn incomplete multi-view clustering (IMVC), missing data induce prototype shifts within views and semantic inconsistencies across views. A feasible solution is to explore cross-view consistency in paired complete observations, further imputing and aligning the similarity relationships inherently shared across views. Nevertheless, existing methods are constrained by two-tiered limitations: (1) Neither instance- nor cluster-level consistency learning construct a semantic space shared across views to learn consensus semantics. The former enforces cross-view instances alignment, and wrongly regards unpaired observations with semantic consistency as negative pairs; the latter focuses on cross-view cluster counterparts while coarsely handling fine-grained intra-cluster relationships within views. (2) Excessive reliance on consistency results in unreliable imputation and alignment without incorporating view-specific cluster information. Thus, we propose an IMVC framework, imputation- and alignment-free for consensus semantics learning (FreeCSL). To bridge semantic gaps across all observations, we learn consensus prototypes from available data to discover a shared space, where semantically similar observations are pulled closer for consensus semantics learning. To capture semantic relationships within specific views, we design a heuristic graph clustering based on modularity to recover cluster structure with intra-cluster compactness and inter-cluster separation for cluster semantics enhancement. Extensive experiments demonstrate, compared to state-of-the-art competitors, FreeCSL achieves more confident and robust assignments on IMVC task. Yuzhuo Dai, Jiaqi Jin, Zhibin Dong, Siwei Wang 0001, Xinwang Liu 0002, En Zhu, Xihong Yang, Xinbiao Gan |
CVPR | 2 |
| 2025 | Enhanced then Progressive Fusion with View Graph for Multi-View ClusteringabstractMulti-view clustering aims to improve clustering accuracy by effectively integrating complementary information from multiple perspectives. However, existing methods often encounter challenges such as feature conflicts between views and insufficient enhancement of individual view features, which hinder clustering performance. To address these challenges, we propose a novel framework, EPFMVC, which integrates feature enhancement with progressive fusion to more effectively align multi-view data. Specifically, we introduce two key innovations: (1) a Feature Channel Attention Encoder (FCAencoder), which adaptively enhances the most discriminative features in each view, and (2) a View Graph-based Progressive Fusion Mechanism, which constructs a view graph using optimal transport (OT) distance to progressively fuse similar views while minimizing inter-view conflicts. By leveraging multi-head attention, the fusion process gradually integrates complementary information, ensuring more consistent and robust shared representations. These innovations enable superior representation learning and effective fusion across views. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art techniques, achieving notable improvements in multi-view clustering tasks across various datasets and evaluation metrics. Zhibin Dong, Meng Liu 0014, Siwei Wang 0001, Ke Liang 0006, Yi Zhang 0104, Suyuan Liu, Jiaqi Jin, Xinwang Liu 0002, En Zhu |
CVPR | 7 |
| 2025 | Diffusion Model Is a Good Steganalyzer: Magnifying Subtle Perturbations in Image DataabstractDigital steganography embeds secret messages into images via invisible modifications, posing challenges for steganalysis, which seeks to detect these alterations by analyzing subtle shifts in image distributions. Previous steganalysis efforts primarily focus on enhancing the steganographic signal while suppressing image semantic content, such as through high-pass filtering. However, these empirically designed methods often lack theoretical underpinnings, exhibiting reduced detection accuracy, particularly at low embedding capacities. To address these limitations, we propose an innovative steganalysis approach that transforms images into pure Gaussian noise representations, actively amplifying the subtle distribution shifts introduced by the steganographic processes in spatial images. This paper pioneers the application of diffusion models to magnify steganographic signals, proposing a new paradigm for further research. Specifically, we iteratively perform forward steps of the probability flow in diffusion models to diminish semantic information. By utilizing the natural spreading properties of the diffusion process, we have theoretically validated the efficacy of each forward step in amplifying differences in noise patterns between cover and stego samples. These magnified differences can be easily captured by a simple classifier—a two-layer MLP. Extensive experiments demonstrate the effectiveness of our method, highlighting detection accuracy gains of 10% to 20% under standard conditions and an average 6.7% increase at low embedding rates compared to existing schemes. Xiaoxiao Hu, Jiaqi Jin, Shengjiu Dai, Sheng Li 0006, Xinpeng Zhang 0001, Zhenxing Qian |
ECAI | 2 |
| 2025 | Deep Incomplete Multi-View Clustering with Distribution Dual-Consistency Recovery GuidanceabstractMulti-view clustering leverages complementary representations from diverse sources to enhance performance. However, real-world data often suffer incomplete cases due to factors like privacy concerns and device malfunctions. A key challenge is effectively utilizing available instances to recover missing views. Existing methods frequently overlook the heterogeneity among views during recovery, leading to significant distribution discrepancies between recovered and true data. Additionally, many approaches focus on cross-view correlations, neglecting insights from intra-view reliable structure and cross-view clustering structure. To address these issues, we propose BURG, a novel method for incomplete multi-view clustering with distriBution dUal-consistency Recovery Guidance. We treat each sample as a distinct category and perform cross-view distribution transfer to predict the distribution space of missing views. To compensate for the lack of reliable category information, we design a dual-consistency guided recovery strategy that includes intra-view alignment guided by neighbor-aware consistency and cross-view alignment guided by prototypical consistency. Extensive experiments on benchmarks demonstrate the superiority of BURG in the incomplete multi-view scenario. Jiaqi Jin, Siwei Wang 0001, Zhibin Dong, Xihong Yang, Xinwang Liu 0002, En Zhu, Kunlun He |
ICCV | 1 |
| 2025 | Generalized Deep Multi-View Clustering Via Causal Learning With Partially Aligned Cross-View CorrespondenceabstractMulti-view clustering (MVC) aims to explore the common clustering structure across multiple views. Many existing MVC methods heavily rely on the assumption of view consistency, where alignments for corresponding samples across different views are ordered in advance. However, real-world scenarios often present a challenge as only partial data is consistently aligned across different views, restricting the overall clustering performance. In this work, we consider the model performance decreasing phenomenon caused by data order shift (i.e., from fully to partially aligned) as a generalized multi-view clustering problem. To tackle this problem, we design a causal multi-view clustering network, termed CauMVC. We adopt a causal modeling approach to understand multi-view clustering procedure. To be specific, we formulate the partially aligned data as an intervention and multi-view clustering with partially aligned data as an post-intervention inference. However, obtaining invariant features directly can be challenging. Thus, we design a Variational Auto-Encoder for causal learning by incorporating an encoder from existing information to estimate the invariant features. Moreover, a decoder is designed to perform the post-intervention inference. Lastly, we design a contrastive regularizer to capture sample correlations. To the best of our knowledge, this paper is the first work to deal generalized multi-view clustering via causal learning. Empirical experiments on both fully and partially aligned data illustrate the strong generalization and effectiveness of CauMVC. Xihong Yang, Siwei Wang 0001, Jiaqi Jin, Fangdi Wang, Tianrui Liu 0001, Yueming Jin, Xinwang Liu 0002, En Zhu, Kunlun He |
ICCV | 3 |
| 2025 | Automatically Identify and Rectify: Robust Deep Contrastive Multi-view Clustering in Noisy ScenariosabstractLeveraging the powerful representation learning capabilities, deep multi-view clustering methods have demonstrated reliable performance by effectively integrating multi-source information from diverse views in recent years. Most existing methods rely on the assumption of clean views. However, noise is pervasive in real-world scenarios, leading to a significant degradation in performance. To tackle this problem, we propose a novel multi-view clustering framework for the automatic identification and rectification of noisy data, termed AIRMVC. Specifically, we reformulate noisy identification as an anomaly identification problem using GMM. We then design a hybrid rectification strategy to mitigate the adverse effects of noisy data based on the identification results. Furthermore, we introduce a noise-robust contrastive mechanism to generate reliable representations. Additionally, we provide a theoretical proof demonstrating that these representations can discard noisy information, thereby improving the performance of downstream tasks. Extensive experiments on six benchmark datasets demonstrate that AIRMVC outperforms state-of-the-art algorithms in terms of robustness in noisy scenarios. The code of AIRMVC are available at https://github.com/xihongyang1999/AIRMVC on Github. Xihong Yang, Siwei Wang 0001, Fangdi Wang, Jiaqi Jin, Suyuan Liu, Yue Liu 0008, En Zhu, Xinwang Liu 0002, Yueming Jin |
ICML | 4 |
| 2025 | SMA-TENG Actuator with Tactile Sensing CapabilityabstractShape memory alloy (SMA) is widely employed in developing actuators. However, the lack of sensing capabilities limits its application. This study presents a sensing-actuation integrated device based on SMA and triboelectric nanogenerator (TENG), achieving tactile sensing while maintaining the actuation performance. The proposed core-shell structure not only repurposes the SMA spring as a key component of actuation and sensing, but also effectively isolates the actuation current to prevent interference with the sensing signal. The aerogel-modified silicone composite layer is applied to the SMA to reduce temperature rise by 30.56%, ensuring the sensing performance. With a rapid response time of less than 31 ms and stable sensing performance exceeding 2000 cycles, the SMA-TENG actuator reliably detects dynamically varying forces and bending. Additionally, it generates a maximum actuation force of 3.21 N, which represents a 12.2% increase compared to a standard SMA spring, due to the pre-stress introduced by the composite layer. Moreover, it can actuate a displacement of 7.7 cm and exhibiting a power density of 7.15 × 103W/m3(at 0.84 V, 6 A). Finally, we validate its haptic sensing capability during actuation, demonstrating its potential towards interactive robotic systems. Jiaqi Jin, Boan Yang, Ziyu Ren |
IROS | 4 |
| 2025 | PostMan: A Productive System for Spatio-temporal Data Management and AnalysisabstractAbstract In daily life, there is an increasing demand for efficient management and analysis of spatio-temporal data. However, current systems struggle to balance multi-functionality, scalability, and computational efficiency in this domain. To address this challenge, we introduce PostMan: a productive spatio-temporal data management system. PostMan is based on Apache Spark and Apache Hadoop HDFS. It extensively, efficiently, and scalably supports spatio-temporal data types and operators across multiple API levels. To realize effective data management and analysis, PostMan designs the unified partition management and hybrid index. Based on this, PostMan has designed and implemented a variety of optimization strategies for vector and raster operators. PostMan also introduces a two-phase static partitioning (TPSP) method to maintain load balance before and after partition filtering during the query process. In the first phase, partitions are generated using an enhanced R*-Tree algorithm, while the second phase allocates partitions by modeling the task as an optimization problem solved through greedy algorithms. For faster computation, PostMan introduces processes and program interfaces for GPU accelerated spatio-temporal operators in Spark. Moreover, extensive evaluations using real-world datasets show PostMan’s notable efficiency and scalability advantages (e.g., 13%-36% improvement) over baseline systems, as well as their constituent techniques. Finally, PostMan has been deployed on the public cloud in a Software as a Service (SaaS) model, garnering substantial attention from customers. Jiaqi Jin, Ziquan Fang, Lu Chen 0001, Yunjun Gao |
Data Sci. Eng. | 1 |
| 2025 | Selective Cross-View Topology for Deep Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering has gained significant attention due to the prevalence of incomplete multi-view data in real-world scenarios. However, existing methods often overlook the critical role of inter-view relationships. In unsupervised settings, selectively leveraging cross-view topological relationships can effectively guide view completion and representation learning. To address this challenge, we propose a novel framework called Selective Cross-View Topology Incomplete Multi-View Clustering (SCVT). Our approach constructs a view topology graph using the Optimal Transport (OT) distance between view. This graph helps identify neighboring views for those with missing data, enabling the inference of topological relationships and accurate completion of missing samples. Additionally, we introduce the Max View Graph Contrastive Alignment module to facilitate information transfer and alignment across neighboring views. Furthermore, we propose the View Graph Weighted Intra-View Contrastive Learning module, which enhances representation learning by pulling representations of samples within the same cluster closer, while applying varying degrees of enhancement across different views based on the view graph. Our method achieves state-of-the-art performance on seven benchmark datasets, significantly outperforming existing methods for incomplete multi-view clustering and demonstrating its effectiveness. Zhibin Dong, Dayu Hu, Jiaqi Jin, Siwei Wang 0001, Xinwang Liu 0002, En Zhu |
IEEE Trans. Image Process. | 3 |
| 2025 | Subgraph Propagation and Contrastive Calibration for Incomplete Multiview Data ClusteringabstractThe success of multiview raw data mining relies on the integrity of attributes. However, each view faces various noises and collection failures, which leads to a condition that attributes are only partially available. To make matters worse, the attributes in multiview raw data are composed of multiple forms, which makes it more difficult to explore the structure of the data especially in multiview clustering task. Due to the missing data in some views, the clustering task on incomplete multiview data confronts the following challenges, namely: 1) mining the topology of missing data in multiview is an urgent problem to be solved; 2) most approaches do not calibrate the complemented representations with common information of multiple views; and 3) we discover that the cluster distributions obtained from incomplete views have a cluster distribution unaligned problem (CDUP) in the latent space. To solve the above issues, we propose a deep clustering framework based on subgraph propagation and contrastive calibration (SPCC) for incomplete multiview raw data. First, the global structural graph is reconstructed by propagating the subgraphs generated by the complete data of each view. Then, the missing views are completed and calibrated under the guidance of the global structural graph and contrast learning between views. In the latent space, we assume that different views have a common cluster representation in the same dimension. However, in the unsupervised condition, the fact that the cluster distributions of different views do not correspond affects the information completion process to use information from other views. Finally, the complemented cluster distributions for different views are aligned by contrastive learning (CL), thus solving the CDUP in the latent space. Our method achieves advanced performance on six benchmarks, which validates the effectiveness and superiority of our SPCC. Zhibin Dong, Jiaqi Jin, Yuyang Xiao, Bin Xiao 0002, Siwei Wang 0001, Xinwang Liu 0002, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | View Gap Matters: Cross-view Topology and Information Decoupling for Multi-view ClusteringabstractMulti-view clustering, a pivotal technology in multimedia research, aims to leverage complementary information from diverse perspectives to enhance clustering performance. The current multi-view clustering methods normally enforce the reduction of distances between any pair of views, overlooking the heterogeneity between views, thereby sacrificing the diverse and valuable insights inherent in multi-view data. In this paper, we propose a Tree-Based View-Gap Maintaining Multi-View Clustering (TGM-MVC) method. Our approach introduces a novel conceptualization of multiple views as a graph structure. In this structure, each view corresponds to a node, with the view gap, calculated by the cosine distance between views, acting as the edge. Through graph pruning, we derive the minimum spanning tree of the views, reflecting the neighbouring relationships among them. Specifically, we applied a share-specific learning framework, and generate view trees for both view-shared and view-specific information. Concerning shared information, we only narrow the distance between adjacent views, while for specific information, we maintain the view gap between neighboring views. Theoretical analysis highlights the risks of eliminating the view gap, and comprehensive experiments validate the efficacy of our proposed TGM-MVC method. Fangdi Wang, Jiaqi Jin, Zhibin Dong, Xihong Yang, Xinwang Liu 0002, Xinzhong Zhu, Siwei Wang 0001, Tianrui Liu 0001, En Zhu |
ACM Multimedia | 2 |
| 2024 | Evaluate then Cooperate: Shapley-based View Cooperation Enhancement for Multi-view ClusteringabstractThe fundamental goal of deep multi-view clustering is to achieve preferable task performance through inter-view cooperation. Although numerous DMVC approaches have been proposed, the collaboration role of individual views have not been well investigated in existing literature. Moreover, how to further enhance view cooperation for better fusion still needs to be explored. In this paper, we firstly consider DMVC as an unsupervised cooperative game where each view can be regarded as a participant. Then, we introduce the Shapley value and propose a novel MVC framework termed Shapley-based Cooperation Enhancing Multi-view Clustering (SCE-MVC), which evaluates view cooperation with game theory. Specially, we employ the optimal transport distance between fused cluster distributions and single view component as the utility function for computing shapley values. Afterwards, we apply shapley values to assess the contribution of each view and utilize these contributions to promote view cooperation. Comprehensive experimental results well support the effectiveness of our framework adopting to existing DMVC frameworks, demonstrating the importance and necessity of enhancing the cooperation among views. Fangdi Wang, Jiaqi Jin, Jingtao Hu, Suyuan Liu, Xihong Yang, Siwei Wang 0001, Xinwang Liu 0002, En Zhu |
NeurIPS | 2 |
| 2024 | HDUD-Net: heterogeneous decoupling unsupervised dehaze network
Jiafeng Li 0001, Lingyan Kuang, Jiaqi Jin, Li Zhuo 0001, Jing Zhang 0023 |
Neural Comput. Appl. | 3 |
| 2024 | Iterative Deep Structural Graph Contrast Clustering for Multiview Raw DataabstractMultiview clustering has attracted increasing attention to automatically divide instances into various groups without manual annotations. Traditional shadow methods discover the internal structure of data, while deep multiview clustering (DMVC) utilizes neural networks with clustering-friendly data embeddings. Although both of them achieve impressive performance in practical applications, we find that the former heavily relies on the quality of raw features, while the latter ignores the structure information of data. To address the above issue, we propose a novel method termed iterative deep structural graph contrast clustering (IDSGCC) for multiview raw data consisting of topology learning (TL), representation learning (RL), and graph structure contrastive learning to achieve better performance. The TL module aims to obtain a structured global graph with constraint structural information and then guides the RL to preserve the structural information. In the RL module, graph convolutional network (GCN) takes the global structural graph and raw features as inputs to aggregate the samples of the same cluster and keep the samples of different clusters away. Unlike previous methods performing contrastive learning at the representation level of the samples, in the graph contrastive learning module, we conduct contrastive learning at the graph structure level by imposing a regularization term on the similarity matrix. The credible neighbors of the samples are constructed as positive pairs through the credible graph, and other samples are constructed as negative pairs. The three modules promote each other and finally obtain clustering-friendly embedding. Also, we set up an iterative update mechanism to update the topology to obtain a more credible topology. Impressive clustering results are obtained through the iterative mechanism. Comparative experiments on eight multiview datasets show that our model outperforms the state-of-the-art traditional and deep clustering competitors. Zhibin Dong, Jiaqi Jin, Yuyang Xiao, Siwei Wang 0001, Xinzhong Zhu, Xinwang Liu 0002, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Deep Incomplete Multi-View Clustering with Cross-View Partial Sample and Prototype AlignmentabstractThe success of existing multi-view clustering relies on the assumption of sample integrity across multiple views. However, in real-world scenarios, samples of multi-view are partially available due to data corruption or sensor failure, which leads to incomplete multi-view clustering study (IMVC). Although several attempts have been proposed to address IMVC, they suffer from the following draw-backs: i) Existing methods mainly adopt cross-view contrastive learning forcing the representations of each sample across views to be exactly the same, which might ignore view discrepancy and flexibility in representations; ii) Due to the absence of non-observed samples across multiple views, the obtained prototypes of clusters might be unaligned and biased, leading to incorrect fusion. To address the above issues, we propose a Cross-view Partial Sample and Prototype Alignment Network (CPSPAN) for Deep Incomplete Multi-view Clustering. Firstly, unlike existing contrastive-based methods, we adopt pair-observed data alignment as 'proxy supervised signals' to guide instance-to-instance correspondence construction among views. Then, regarding of the shifted prototypes in IMVC, we further propose a prototype alignment module to achieve incomplete distribution calibration across views. Extensive experimental results showcase the effectiveness of our proposed modules, attaining noteworthy performance improvements when compared to existing IMVC competitors on benchmark datasets. Jiaqi Jin, Siwei Wang 0001, Zhibin Dong, Xinwang Liu 0002, En Zhu |
CVPR | 1 |
| 2023 | Cross-view Topology Based Consistent and Complementary Information for Deep Multi-view ClusteringabstractMulti-view clustering aims to extract valuable information from different sources or perspectives. Over the years, the deep neural network has demonstrated its superior representation learning capability in multi-view clustering and achieved impressive performance. However, most existing deep clustering approaches are dedicated to merging and exploring the consistent latent representation across multiple views while overlooking the abundant complementary information in each view. Furthermore, finding correlations between multiple views in an unsupervised setting is a significant challenge. To tackle these issues, we present a novel Cross-view Topology based Consistent and Complementary information extraction framework, termed CTCC. In detail, deep embedding can be obtained from the bipartite graph learning module for each view individually. CTCC then constructs the cross-view topological graph based on the OT distance between the bipartite graph of each view. Utilizing the above graph, we maximize the mutual information across views to learn consistent information and enhance the complementarity of each view by selectively isolating distributions from each other. Extensive experiments on five challenging datasets verify that CTCC outperforms existing methods significantly. Zhibin Dong, Siwei Wang 0001, Jiaqi Jin, Xinwang Liu 0002, En Zhu |
ICCV | 3 |
| 2023 | DealMVC: Dual Contrastive Calibration for Multi-view ClusteringabstractBenefiting from the strong view-consistent information mining capacity, multi-view contrastive clustering has attracted plenty of attention in recent years. However, we observe the following drawback, which limits the clustering performance from further improvement. The existing multi-view models mainly focus on the consistency of the same samples in different views while ignoring the circumstance of similar but different samples in cross-view scenarios. To solve this problem, we propose a novel Dual contrastive calibration network for Multi-View Clustering (DealMVC). Specifically, we first design a fusion mechanism to obtain a global cross-view feature. Then, a global contrastive calibration loss is proposed by aligning the view feature similarity graph and the high-confidence pseudo-label graph. Moreover, to utilize the diversity of multi-view information, we propose a local contrastive calibration loss to constrain the consistency of pair-wise view features. The feature structure is regularized by reliable class information, thus guaranteeing similar samples have similar features in different views. During the training procedure, the interacted cross-view feature is jointly optimized at both local and global levels. In comparison with other state-of-the-art approaches, the comprehensive experimental results obtained from eight benchmark datasets provide substantial validation of the effectiveness and superiority of our algorithm. We release the code of DealMVC at https://github.com/xihongyang1999/DealMVC on GitHub. Xihong Yang, Jiaqi Jin, Siwei Wang 0001, Ke Liang 0006, Yue Liu 0008, Yi Wen 0001, Suyuan Liu, Sihang Zhou 0001, Xinwang Liu 0002, En Zhu |
ACM Multimedia | 2 |
| 2023 | A Hybrid Algorithm for Dust Aerosol Detection: Integrating Forward Radiative Transfer Simulations and Machine LearningabstractA hybrid algorithm based on radiative transfer simulations and machine learning for dust aerosol detection, is developed for the Advanced Himawari Imager (AHI) carried by the geostationary satellite Himawari-8. The sensitivities of the AHI thermal infrared (TIR) channels for dust aerosols are analyzed through radiative transfer simulations. The sensitivity study demonstrates that the simulated clear-sky brightness temperatures (BTs) show an obvious improvement in identifying dust aerosols compared to brightness temperature difference techniques, especially optically thin dust. Therefore, the simulated clear-sky BTs and AHI TIR observed BTs, in addition to ground information, are used as inputs to add physical knowledge in the machine learning model. The performance of an artificial neural network constructed for dust aerosol detection is evaluated by comparing its results with those of active Cloud-Aerosol Lidar with Orthogonal Polarization measurements. The proposed algorithm effectively achieves dust aerosol detection during both daytime and nighttime, with a precision of over 86% and a recall of over 85% on an independent testing dataset. The proposed algorithm is applied to three typical dust events to further illustrate its applicability. Although some thin dust aerosols near the ground are misclassified due to weak signals, most dust aerosols are successfully detected, and the identification is generally not affected by other types of aerosols. The results of the regional classification demonstrate that our algorithm is superior in detecting tenuous dust aerosols compared to Dust RGB images using the AHI TIR channels and physical-based algorithm. Jiaqi Jin, Feng Zhang 0041, Linlu Mei, Lin Chen 0017 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Detecting Insulator Strings as Linked Chain Structure in Smart Grid InspectionabstractIn high-voltage power systems, insulators are essential components in transmission lines for increasing shooting distance and securing wires. Unmanned aerial vehicle imaging becomes a common way of inspecting the state of the insulators. However, the automatic detection of insulators with complex backgrounds is still a challenging task. Most of the existing object detection methods are based on anchors, which do not have sufficient ability to describe objects that have a string-like structure. To tackle it, inspired by the keypoints-based object detection method, we propose a novel chain structure framework to detect insulators. First, we model the string-like object as a chain structure consisting of keypoints and their linkage, by which the insulator strings features are efficiently encoded and trained in the proposed ChainNet. Then, an assembling algorithm is proposed to assemble the estimated keypoints and linkages into chains, by which the insulator strings can be tightly enclosed in rotational bounding boxes. We evaluate the proposed approach on our collected large-scale rigorous dataset under the Directed Intersection over Union metric. The extensive experimental results show that the proposed method achieves 75.7% mAP, which yields 3.4, 15, and 10.6 improvements to the state-of-the-art rotational anchor, axial-aligned anchor, and anchor-free detection methods, respectively. Moreover, the proposed framework can easily be extended to detect other string-like manmade objects in the industrial area. Jiaqi Jin, Shuifa Sun |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | DynGCF: Augmenting Inactive Users and Items in Dynamic Graph-based Collaborative FilteringabstractModeling user-item interactions in a dynamic manner bring new insight to the representation learning for recom-mender systems. Distinct from static graph-based approaches that model the whole user-item interaction graph, dynamic graph-based approaches model both the structural and temporal information from a sequence of snapshot graphs. Despite effectiveness, we argue that existing approaches do not explicitly address the temporal sparsity issue, which degrades the representation learning performance for inactive users and items. Therefore, we propose a new Dynamic Graph-based Collaborative Filtering(DynGCF) framework. In particular, it utilizes the vanilla interaction graph with the co-occurrence graph(co-graph) to jointly explores 1-hop collaborative and 2-hop implicit similarity for dynamic representation learning. Moreover, to further alleviate temporal sparsity, we explore representative(active) users and items via graph pooling and design an activity-guided gating(AGate) layer to augment inactive users and items. At last, we further stack a temporal aggregator layer to obtain the final representation. We conduct extensive experiments on four real-world benchmark datasets to demonstrate the significant performance gains for DynGCF over several state-of-the-art methods. Further analyses also show the necessity of alleviating temporal sparsity for improving recommendation performance. Jiaqi Jin, Mengfei Zhang, Mao Pan, Jinyun Fang |
IJCNN | 1 |
| 2022 | Align then Fusion: Generalized Large-scale Multi-view Clustering with Anchor Matching CorrespondencesabstractMulti-view anchor graph clustering selects representative anchors to avoid full pair-wise similarities and therefore reduce the complexity of graph methods. Although widely applied in large-scale applications, existing approaches do not pay sufficient attention to establishing correct correspondences between the anchor sets across views. To be specific, anchor graphs obtained from different views are not aligned column-wisely. Such an Anchor-Unaligned Problem (AUP) would cause inaccurate graph fusion and degrade the clustering performance. Under multi-view scenarios, generating correct correspondences could be extremely difficult since anchors are not consistent in feature dimensions. To solve this challenging issue, we propose the first study of the generalized and flexible anchor graph fusion framework termed Fast Multi-View Anchor-Correspondence Clustering (FMVACC). Specifically, we show how to find anchor correspondence with both feature and structure information, after which anchor graph fusion is performed column-wisely. Moreover, we theoretically show the connection between FMVACC and existing multi-view late fusion and partial view-aligned clustering, which further demonstrates our generality. Extensive experiments on seven benchmark datasets demonstrate the effectiveness and efficiency of our proposed method. Moreover, the proposed alignment module also shows significant performance improvement applying to existing multi-view anchor graph competitors indicating the importance of anchor alignment. Our code is available at \url{https://github.com/wangsiwei2010/NeurIPS22-FMVACC}. Siwei Wang 0001, Xinwang Liu 0002, Suyuan Liu, Jiaqi Jin, Wenxuan Tu, Xinzhong Zhu, En Zhu |
NeurIPS | 4 |
| 2022 | Asymptotically Unbiased Estimation for Delayed Feedback Modeling via Label CorrectionabstractAlleviating the delayed feedback problem is of crucial importance for the conversion rate(CVR) prediction in online advertising. Previous delayed feedback modeling methods using an observation window to balance the trade-off between waiting for accurate labels and consuming fresh feedback. Moreover, to estimate CVR upon the freshly observed but biased distribution with fake negatives, the importance sampling is widely used to reduce the distribution bias. While effective, we argue that previous approaches falsely treat fake negative samples as real negative during the importance weighting and have not fully utilized the observed positive samples, leading to suboptimal performance. Jiaqi Jin, Pengjie Wang 0002, Jian Xu 0015, Bo Zheng 0007 |
WWW | 2 |
| 2021 | Sequential Recommendation with Context-Aware Collaborative Graph Attention NetworksabstractRecently, sequence features have been extensively studied to improve the performance of recommender systems. However, advanced sequential recommendation methods that rely only on item IDs still face the challenge of modeling fine-grained user preference from interactive data. Furthermore, context-aware sequential recommendations have the hardness of modeling the relationship between items and items, items and users. Both of these two methods ignore the effect of categories on users' next click tendency and the interactive learning between categories and items. In this paper, we propose a method named Contextual Collaborative Graph Attention Network (CCGAT) to model the sequence. Methodologically, user behavior sequences are constructed as graph-structured data, and we apply two similar graph self-attention networks to model the item transitions and the category click probability. CCGAT takes advantage of the fact that users tend to click on the same or similar categories under specific purposes, and provides a simple but effective way to train two networks collaboratively. Extensive experiments on five real-world datasets show that our model outperforms state-of-the-art methods, and demonstrate the validity of modeling both contextual information and graph features. Mengfei Zhang, Jiaqi Jin, Mao Pan, Jinyun Fang |
IJCNN | 3 |
| 2021 | Modeling Hierarchical Intents and Selective Current Interest for Session-Based Recommendation
Mengfei Zhang, Jiaqi Jin, Mao Pan, Jinyun Fang |
PAKDD (2) | 3 |
| 2021 | m6A regulator-mediated methylation modification patterns and characteristics of immunity and stemness in low-grade gliomaabstractm6A RNA methylation is an emerging epigenetic modification, and its potential role in immunity and stemness remains unknown. Based on 17 widely recognized m6A regulators, the m6A modification patterns and corresponding characteristics of immune infiltration and stemness of 1152 low-grade glioma samples were comprehensively analyzed. Machine-learning strategies for constructing m6AScores were trained to quantify the m6A modification patterns of individual samples. Here, we reveal a significant correlation between the multi-omics data of regulators and clinicopathological parameters. We identified two distinct m6A modification patterns (an immune-activated differentiation pattern and an immune-desert dedifferentiation pattern) and four regulatory patterns of m6A methylation on immunity and stemness. We show that the m6AScores can predict the molecular subtype of low-grade glioma, the abundance of immune infiltration, the enrichment of signaling pathways, gene variation and prognosis. The concentration of high immunogenicity and clinical benefits in the low-m6AScore group confirmed the sensitive response to radio-chemotherapy and immunotherapy in patients with high-m6AScore. The results of the pan-cancer analyses illustrate the significant correlation between m6AScore and clinical outcome, the burden of neoepitope, immune infiltration and stemness. The assessment of individual tumor m6A modification patterns will guide us in improving treatment strategies and developing objective diagnostic tools. Jianyang Du, Hang Ji, Jiaqi Jin, Shan Mi, Kuiyuan Hou, Chaochao Zhang, Shaoshan Hu |
Briefings Bioinform. | 4 |
| 2020 | Session-based Recommendation with Hierarchical Leaping NetworksabstractSession-based recommendation aims to predict the next item that users will interact based solely on anonymous sessions. In real-life scenarios, the user's preferences are usually various, and distinguishing different preferences in the session is important. However previous studies focus mostly on the transition modeling between items, ignoring the mining of various user preferences. In this paper, we propose a Hierarchical Leaping Network (HLN) to explicitly model the users' multiple preferences by grouping items that share some relationships. We first design a Leap Recurrent Unit (LRU) which is capable of skipping preference-unrelated items and accepting knowledge of previously learned preferences. Then we introduce a Preference Manager (PM) to manage those learned preferences and produce an aggregated preference representation each time LRU reruns. The final output of PM which contains multiple preferences of the user is used to make recommendations. Experiments on two benchmark datasets demonstrate the effectiveness of HLN. Furthermore, the visualization of explicitly learned subsequences also confirms our idea. Mengfei Zhang, Jinyun Fang, Jiaqi Jin, Mao Pan |
SIGIR | 4 |
| 2019 | A Geohash Based Place2vec ModelabstractLearning the vector representing of Point Of Interest(POI) is a key aspect of POI recommender systems. As for shop POI embedding, in addition to the goods selling in shops, the location of shops is also an important factor that must be considered. Word2vec is a commonly used POI embedding model but it cannot be trained directly using location data. In this paper, we present a geohash based Place2vec model, geohash is a geocoding system that can encoding the location of shops in a string form, which can be treated as a spatial context of the Word2Vec model. We investigate the extent to which similar shops occur within the same products contexts and similar spatial contexts, and enrich a dataset of location, type and product lists of shops from YIWUGOU Online Shop Data1. The evaluation results shows that the shop vector trained by combined contexts outperform the vector trained by the products contexts. Jiaqi Jin, Zhuojian Xiao, Qiang Qiu 0003, Jinyun Fang |
IGARSS | 1 |