Jiahe Li 0008

dblp:265/2448-8 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2024
0009-0003-5092-1326ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Data mining · 72% Recommender systems · 22% Data stream processing · 7%
Artificial intelligence
2 papers
Graph learning · 74% Trustworthy machine learning · 26%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
context-aware learning
0.812024
Con4m: Context-aware Consistency Learning Framework for Segmented Time Series Classification · NeurIPS 2024
Recommender systems › collaborative filtering
matrix factorization
0.812024
Fast Updating Truncated SVD for Representation Learning with Sparse Matrices · ICLR 2024
Data mining
representation learning
0.812024
Fast Updating Truncated SVD for Representation Learning with Sparse Matrices · ICLR 2024
Data mining › time series analysis
time series classification
0.812024
Con4m: Context-aware Consistency Learning Framework for Segmented Time Series Classification · NeurIPS 2024
Algorithms and data structures › numerical linear algebra
matrix factorization
0.812024
Fast Updating Truncated SVD for Representation Learning with Sparse Matrices · ICLR 2024
Machine learning › Graph learning › dynamic graph learning
dynamic network embedding
0.712023
Accelerating Dynamic Network Embedding with Billions of Parameter Updates to Milliseconds · KDD 2023
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.212024
Con4m: Context-aware Consistency Learning Framework for Segmented Time Series Classification · NeurIPS 2024
Data stream processing
evolving data
0.212024
Fast Updating Truncated SVD for Representation Learning with Sparse Matrices · ICLR 2024

Methods — techniques the papers use, named apart from their topics

rayleigh-ritz projection · 1.5orthogonalization · 1.5consistency learning · 1.5con4m · 1.5personalized pagerank · 1.3matrix factorization · 1.3
YearPublicationVenuePosition
2024 Fast Updating Truncated SVD for Representation Learning with Sparse Matrices
abstract
Updating truncated Singular Value Decomposition (SVD) has extensive applications in representation learning. The continuous evolution of massive-scaled data matrices in practical scenarios highlights the importance of aligning SVD-based models with fast-paced updates. Recent methods for updating truncated SVD can be recognized as Rayleigh-Ritz projection procedures where their projection matrices are augmented based on the original singular vectors. However, the updating process in these methods densifies the update matrix and applies the projection to all singular vectors, resulting in inefficiency. This paper presents a novel method for dynamically approximating the truncated SVD of a sparse and temporally evolving matrix. The proposed method takes advantage of sparsity in the orthogonalization process of the augment matrices and employs an extended decomposition to store projections in the column space of singular vectors independently. Numerical experimental results on updating truncated SVD for evolving sparse matrices show an order of magnitude improvement in the efficiency of our proposed method while maintaining precision comparing to previous methods.
Yang Yang 0009, Jiahe Li 0008, Shiliang Pu
ICLR3
2024 Con4m: Context-aware Consistency Learning Framework for Segmented Time Series Classification
abstract
Time Series Classification (TSC) encompasses two settings: classifying entire sequences or classifying segmented subsequences. The raw time series for segmented TSC usually contain Multiple classes with Varying Duration of each class (MVD). Therefore, the characteristics of MVD pose unique challenges for segmented TSC, yet have been largely overlooked by existing works. Specifically, there exists a natural temporal dependency between consecutive instances (segments) to be classified within MVD. However, mainstream TSC models rely on the assumption of independent and identically distributed (i.i.d.), focusing on independently modeling each segment. Additionally, annotators with varying expertise may provide inconsistent boundary labels, leading to unstable performance of noise-free TSC models. To address these challenges, we first formally demonstrate that valuable contextual information enhances the discriminative power of classification instances. Leveraging the contextual priors of MVD at both the data and label levels, we propose a novel consistency learning framework Con4m, which effectively utilizes contextual information more conducive to discriminating consecutive segments in segmented TSC tasks, while harmonizing inconsistent boundary labels for training. Extensive experiments across multiple datasets validate the effectiveness of Con4m in handling segmented TSC tasks on MVD. The source code is available at https://github.com/MrNobodyCali/Con4m.
Tianyu Cao 0006, Jiahe Li 0008, Zhilong Chen, Yang Yang 0009
NeurIPS4
2023 Accelerating Dynamic Network Embedding with Billions of Parameter Updates to Milliseconds
abstract
Network embedding, a graph representation learning method illustrating network topology by mapping nodes into lower-dimension vectors, is challenging to accommodate the ever-changing dynamic graphs in practice. Existing research is mainly based on node-by-node embedding modifications, which falls into the dilemma of efficient calculation and accuracy. Observing that the embedding dimensions are usually much smaller than the number of nodes, we break this dilemma with a novel dynamic network embedding paradigm that rotates and scales the axes of embedding space instead of a node-by-node update. Specifically, we propose the Dynamic Adjacency Matrix Factorization (DAMF) algorithm, which achieves an efficient and accurate dynamic network embedding by rotating and scaling the coordinate system where the network embedding resides with no more than the number of edge modifications changes of node embeddings. Moreover, a dynamic Personalized PageRank is applied to the obtained network embeddings to enhance node embeddings and capture higher-order neighbor information dynamically. Experiments of node classification, link prediction, and graph reconstruction on different-sized dynamic graphs suggest that DAMF advances dynamic network embedding. Further, we unprecedentedly expand dynamic network embedding experiments to billion-edge graphs, where DAMF updates billion-level parameters in less than 10ms.
Yang Yang 0009, Jiahe Li 0008, Haoyang Cai, Shiliang Pu
KDD3