Junlin He

dblp:304/9915 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Representation and self-supervised learning · 39% Deep learning architectures and training · 38% Language models and text generation · 11%
Databases, data mining, and information retrieval
1 paper
Data mining · 75% Recommender systems · 25%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
regularization
1.522024
Preventing Model Collapse in Deep Canonical Correlation Analysis by Noise Regularization · NeurIPS 2024
Preventing Dimensional Collapse in Self-Supervised Learning via Orthogonality Regularization · NeurIPS 2024
Recommender systems › representation learning for recommendation
contrastive learning for recommendation
0.912025
Deep Multi-View Contrastive Clustering via Graph Structure Awareness · IEEE Trans. Image Process. 2025
Data mining › clustering › multi-view clustering › deep multi-view clustering
contrastive multi-view clustering
0.912025
Deep Multi-View Contrastive Clustering via Graph Structure Awareness · IEEE Trans. Image Process. 2025
Data mining › clustering › multi-view clustering
deep multi-view clustering
0.912025
Deep Multi-View Contrastive Clustering via Graph Structure Awareness · IEEE Trans. Image Process. 2025
Data mining › clustering
multi-view clustering
0.912025
Deep Multi-View Contrastive Clustering via Graph Structure Awareness · IEEE Trans. Image Process. 2025
Machine learning › Representation and self-supervised learning › multi-view learning › canonical correlation analysis
deep canonical correlation analysis
0.812024
Preventing Model Collapse in Deep Canonical Correlation Analysis by Noise Regularization · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation analysis
dimensional collapse
0.812024
Preventing Dimensional Collapse in Self-Supervised Learning via Orthogonality Regularization · NeurIPS 2024
Machine learning › Generative modeling
model collapse
0.812024
Preventing Model Collapse in Deep Canonical Correlation Analysis by Noise Regularization · NeurIPS 2024
Machine learning › Representation and self-supervised learning › multi-view learning
multi-view representation learning
0.812024
Preventing Model Collapse in Deep Canonical Correlation Analysis by Noise Regularization · NeurIPS 2024
Machine learning › Deep learning architectures and training › regularization
noise-based regularization
0.812024
Preventing Model Collapse in Deep Canonical Correlation Analysis by Noise Regularization · NeurIPS 2024
Machine learning › Deep learning architectures and training › regularization
orthogonal constraint
0.812024
Preventing Dimensional Collapse in Self-Supervised Learning via Orthogonality Regularization · NeurIPS 2024
Machine learning › Time series and sequential data
spatiotemporal forecasting
0.312025
Geolocation Representation from Large Language Models Are Generic Enhancers for Spatio-Temporal Learning · AAAI 2025

Methods — techniques the papers use, named apart from their topics

large language model · 0.9graph convolutional network · 0.9feature concatenation · 0.9contrastive learning · 0.9autoencoder · 0.9orthogonal regularization · 0.8noise regularization · 0.8canonical correlation analysis · 0.8
YearPublicationVenuePosition
2025 Geolocation Representation from Large Language Models Are Generic Enhancers for Spatio-Temporal Learning
abstract
In the geospatial domain, universal representation models are significantly less prevalent than their extensive use in natural language processing and computer vision. This discrepancy arises primarily from the high costs associated with the input of existing representation models, which often require street views and mobility data. To address this, we develop a novel, training-free method that leverages large language models (LLMs) and auxiliary map data from OpenStreetMap to derive geolocation representations (LLMGeovec). LLMGeovec can represent the geographic semantics of city, country, and global scales, which acts as a generic enhancer for spatio-temporal learning. Specifically, by direct feature concatenation, we introduce a simple yet effective paradigm for enhancing multiple spatio-temporal tasks including geographic prediction (GP), long-term time series forecasting (LTSF), and graph-based spatio-temporal forecasting (GSTF). LLMGeovec can seamlessly integrate into a wide spectrum of spatio-temporal learning models, providing immediate enhancements. Experimental results demonstrate that LLMGeovec achieves global coverage and significantly boosts the performance of leading GP, LTSF, and GSTF models.
Junlin He, Tong Nie 0001, Wei Ma 0016
AAAI1
2025 Deep Multi-View Contrastive Clustering via Graph Structure Awareness
abstract
Multi-view clustering (MVC) aims to exploit the latent relationships between heterogeneous samples in an unsupervised manner, which has served as a fundamental task in the unsupervised learning community and has drawn widespread attention. In this work, we propose a new deep multi-view contrastive clustering method via graph structure awareness (DMvCGSA) by conducting both instance-level and cluster-level contrastive learning to exploit the collaborative representations of multi-view samples. Unlike most existing deep multi-view clustering methods, which usually extract only the attribute features for multi-view representation, we first exploit the view-specific features while preserving the latent structural information between multi-view data via a GCN-embedded autoencoder, and further develop a similarity-guided instance-level contrastive learning scheme to make the view-specific features discriminative. Moreover, unlike existing methods that separately explore common information, which may not contribute to the clustering task, we employ cluster-level contrastive learning to explore the clustering-beneficial consistency information directly, resulting in improved and reliable performance for the final multi-view clustering task. Extensive experimental results on twelve benchmark datasets clearly demonstrate the encouraging effectiveness of the proposed method compared with the state-of-the-art models.
Lunke Fei, Junlin He, Qi Zhu 0001, Shuping Zhao, Jie Wen 0001, Yong Xu 0001
IEEE Trans. Image Process.2
2025 Activity-Aware Human Mobility Prediction With Hierarchical Graph Attention Recurrent Network
abstract
Human mobility prediction is a fundamental task essential for various applications in urban planning, location-based services and intelligent transportation systems. Existing methods often ignore activity information crucial for reasoning human preferences and routines, or adopt a simplified representation of the dependencies between time, activities and locations. To address these issues, we present Hierarchical Graph Attention Recurrent Network (Hgarn) for human mobility prediction. Specifically, we construct a hierarchical graph based on past mobility records and employ a Hierarchical Graph Attention Module to capture complex time-activity-location dependencies. This way, Hgarn can learn representations with rich human travel semantics to model user preferences at the global level. We also propose a model-agnostic history-enhanced confidence (MaHec) label to incorporate each user’s individual-level preferences. Finally, we introduce a Temporal Module, which employs recurrent structures to jointly predict users’ next activities and their associated locations, with the former used as an auxiliary task to enhance the latter prediction. For model evaluation, we test the performance of Hgarn against existing state-of-the-art methods in both the recurring (i.e., returning to a previously visited location) and explorative (i.e., visiting a new location) settings. Overall, Hgarn outperforms other baselines significantly in all settings based on two real-world human mobility data benchmarks. These findings confirm the important role that human activities play in determining mobility decisions, illustrating the need to develop activity-aware intelligent transportation systems. Source codes of this study are available athttps://github.com/YihongT/HGARN
Yihong Tang, Junlin He, Zhan Zhao
IEEE Trans. Intell. Transp. Syst.2
2024 ARPSSO: An OIDC-Compatible Privacy-Preserving SSO Scheme Based on RP Anonymization
Junlin He, Lingguang Lei, Yuewu Wang, Pingjian Wang, Jiwu Jing
ESORICS (2)1
2024 Preventing Dimensional Collapse in Self-Supervised Learning via Orthogonality Regularization
abstract
Self-supervised learning (SSL) has rapidly advanced in recent years, approaching the performance of its supervised counterparts through the extraction of representations from unlabeled data. However, dimensional collapse, where a few large eigenvalues dominate the eigenspace, poses a significant obstacle for SSL. When dimensional collapse occurs on features (e.g. hidden features and representations), it prevents features from representing the full information of the data; when dimensional collapse occurs on weight matrices, their filters are self-related and redundant, limiting their expressive power. Existing studies have predominantly concentrated on the dimensional collapse of representations, neglecting whether this can sufficiently prevent the dimensional collapse of the weight matrices and hidden features. To this end, we first time propose a mitigation approach employing orthogonal regularization (OR) across the encoder, targeting both convolutional and linear layers during pretraining. OR promotes orthogonality within weight matrices, thus safeguarding against the dimensional collapse of weight matrices, hidden features, and representations. Our empirical investigations demonstrate that OR significantly enhances the performance of SSL methods across diverse benchmarks, yielding consistent gains with both CNNs and Transformer-based architectures.
Junlin He, Jinxiao Du
NeurIPS1
2024 Preventing Model Collapse in Deep Canonical Correlation Analysis by Noise Regularization
abstract
Multi-View Representation Learning (MVRL) aims to learn a unified representation of an object from multi-view data. Deep Canonical Correlation Analysis (DCCA) and its variants share simple formulations and demonstrate state-of-the-art performance. However, with extensive experiments, we observe the issue of model collapse, i.e., the performance of DCCA-based methods will drop drastically when training proceeds. The model collapse issue could significantly hinder the wide adoption of DCCA-based methods because it is challenging to decide when to early stop. To this end, we develop NR-DCCA, which is equipped with a novel noise regularization approach to prevent model collapse. Theoretical analysis shows that the Correlation Invariant Property is the key to preventing model collapse, and our noise regularization forces the neural network to possess such a property. A framework to construct synthetic data with different common and complementary information is also developed to compare MVRL methods comprehensively. The developed NR-DCCA outperforms baselines stably and consistently in both synthetic and real-world datasets, and the proposed noise regularization approach can also be generalized to other DCCA-based methods such as DGCCA.
Junlin He, Jinxiao Du, Susu Xu
NeurIPS1
2023 Query-dominant User Interest Network for Large-Scale Search Ranking
abstract
Historical behaviors have shown great effect and potential in various prediction tasks, including recommendation and information retrieval. The overall historical behaviors are various but noisy while search behaviors are always sparse. Most existing approaches in personalized search ranking adopt the sparse search behaviors to learn representation with bottleneck, which do not sufficiently exploit the crucial long-term interest. In fact, there is no doubt that user long-term interest is various but noisy for instant search, and how to exploit it well still remains an open problem.
Yong Yuan 0004, Jingyou Hou, Bingqing Ke, Junlin He, Shunyu Zhang, Enyun Yu, Wenwu Ou
CIKM9
2022 HQANN: Efficient and Robust Similarity Search for Hybrid Queries with Structured and Unstructured Constraints
abstract
The in-memory approximate nearest neighbor search (ANNS) algorithms have achieved great success for fast high-recall query processing, but are extremely inefficient when handling hybrid queries with unstructured (i.e., feature vectors) and structured (i.e., related attributes) constraints. In this paper, we present HQANN, a simple yet highly efficient hybrid query processing framework which can be easily embedded into existing proximity graph-based ANNS algorithms. We guarantee both low latency and high recall by leveraging navigation sense among attributes and fusing vector similarity search with attribute filtering. Experimental results on both public and in-house datasets demonstrate that HQANN is 10x faster than the state-of-the-art hybrid ANNS solutions to reach the same recall quality and its performance is hardly affected by the complexity of attributes. It can reach 99% [email protected] in just around 50 microseconds On GLOVE-1.2M with thousands of attribute constraints.
Junlin He, Guoheng Fu
CIKM2