VLDB 2026 Research / reviewers in the wild / expert
Yiheng Jiang
dblp:260/4584
· DBLP profile ↗
27ranked-venue papers
9as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NP-MiSR: Neural Process-based Multi-Interest Learning for Session-Based RecommendationabstractSession-based recommendation (SBR) aims to provide users with satisfactory suggestions via modeling preferences based on short-term, anonymous user-item interaction sequences. Traditional single interest learning methods struggle to align with the diverse nature of preferences. Recent advances resolved this bottleneck by learning multiple interest embeddings for each session. However, due to the pre-defining scheme of interest quantity (e.g. the number of interests), these approaches are deficient in adaptive ability towards distinctive preference patterns across different users. Moreover, these methods rely solely on the current session and ignore useful information from related ones. The short-term property of sessions would magnify the insufficient representation issue. To address these limitations, we propose a Neural Process-based Multi-interest learning framework for Session-based Recommendation, namely NP-MiSR. To be specific, our method enables adaptive multi-interest representation learning through two complementary mechanisms: 1) Neural Process-based Intra-session interest modeling: We employ Neural Processes to model the distribution of interests within a session, where the fixed interest configurations are no longer needed. 2) Cross-session context fusion: We extract interest distributions of similar sessions as contextual priors to refine the current session’s interest representation. Extensive experiments on three datasets demonstrate that our method consistently outperforms state-of-the-art SBR approaches with an average improvement of 38.8%. Moreover, the few-shot learning task reveals that NP-MiSR achieves a surprisingly favorable efficiency v.s. performance trade-off where utilizing only 10% of the training data attains 95% of the recommendation performance. Jun Bao, Yiheng Jiang, Xiangfeng Liu, Yuanbo Xu |
AAAI | 3 |
| 2026 | Practical Image Compression with Energy-Guided Asymmetric Entropy ModelingabstractRecent learned image compression (LIC) methods have demonstrated superior ratedistortion performance compared to classical image compression standards. However, their substantially higher computational complexity and inefficient entropy modeling designs raise challenges to practical deployment. To address this issue, we propose Energy-Guided Asymmetric Entropy Modeling for ultra-low-complexity learned image compression. Building on the observation that low-energy channels can be effectively modeled with less complexity, we propose a novel Asymmetric Hyperprior Transform (AHT). AHT evenly split the latent features into channel groups. Low-energy groups are processed by lightweight subnetworks, achieving reduced complexity. To further ensure proper energy allocation on channel groups, we design an Energy Allocation (EA) Loss that constrains latent features with respect to their estimated mean, thus enabling flexible control over channel energy distribution. Overall, our entropy modeling design is extremely lightweight and can be efficiently deployed on CPUs, making it well-suited for practical deployment. Experiments demonstrate that our proposed model achieves a 5% reduction in BD-rate over BPG with decoding complexity below$10 \text{kMACs} /$pixel, and achieves a BD-rate gain per decoding MACs/pixel of -645.89, striking a favorable balance between rate-distortion performance and computational cost. Yiheng Jiang, Haotian Zhang 0009, Li Li 0040, Dong Liu 0002 |
DCC | 3 |
| 2026 | FNeRV: Frequency-Aware Neural Video Representations with Dynamic Spatial Guidance
Ziwen Xiong, Yiheng Jiang, Li Li 0040, Dong Liu 0002 |
ISCAS | 2 |
| 2026 | TMMSRec: Time-interval-aware Multi-Modal Sequential RecommenderabstractConventional multi-modal sequential recommenders usually employ sequence models (e.g., SASRec and GRU4Rec) as the recommender framework for sequential dependency modeling and multi-modal information as add-on knowledge for fine-grained preference learning. However, the existing methods ignore the time interval information in the interaction sequences, which reflects the user’s preference transition. Consequently, they acquire biased preference representations, leading to suboptimal performance. Along these lines, we concentrate on incorporating the time interval into the multi-modal sequential recommendation and investigating the internal influences across multiple modalities. Firstly, we treat the time interval as an independent modality and exploit a Time Interval Encoder (TIE) to quantify the time interval in the sequences. Secondly, we consider the heterogeneity between multiple modalities and design a two-stage fusion strategy. The first stage injects time intervals into each modality (ID, image, and text) and the second stage aggregates the user preferences from each modality to get the user’s sequence preferences. To this end, we build a model-agnostic framework for the multi-modal sequential recommendation, namely the Time-interval-aware Multi-Modal Sequential Recommender (TMMSRec) . The empirical validations on four real-world datasets prove the effectiveness of our method, where the maximum performance improvement against several state-of-the-art baselines achieves up to 16.05%. Our code is available at https://github.com/wangpp0602/TMMSRec . Yuanbo Xu, Yiheng Jiang, Hangtong Xu, Fuzhen Zhuang |
ACM Trans. Inf. Syst. | 3 |
| 2025 | Auto Encoding Neural Process for Multi-interest RecommendationabstractMulti-interest recommendation constantly aspires to an oracle individual preference modeling approach, that satisfies the diverse and dynamic properties. Fueled by the deep learning technology, existing neural network (NN)-based recommender systems employ single-point or multi-point interest representation strategy to realize preference modeling,and boost the recommendation performance with a remarkable margin. However, as parameterized approximate functions, NN-based methods remain deficiencies with respect to the adaptability towards distinctive preference patterns cross different users and the calibration over the individual current intent. In this paper, we revisit multi-interest recommendation with the lens of stochastic process and Bayesian inference. Specifically, we propose to learn a distribution over functions to depict the individual diverse preferences rather than a unified function to approximate preference. Subsequently, the recommendation is encouraged with the uncertainty estimation which conforms to the dynamic shifting intent. Along these lines, we establish the connection between multi-interest recommendation and neural processes by proposing NP-Rec, which realizes the flexible multiple interests modeling and uncertainty estimation, simultaneously. Empirical study on 4 real world datasets demonstrates that our NP-Rec attains superior recommendation performances to several state-of-the-art baselines, where the average improvement achieves up to 13.94%. Yiheng Jiang, Yuanbo Xu, Yongjian Yang 0001, Funing Yang, Pengyang Wang, Chaozhuo Li |
AAAI | 1 |
| 2025 | A Small-footprint Acoustic Echo Cancellation Solution for Mobile Full-Duplex Speech InteractionsabstractIn full-duplex speech interaction systems, effective Acoustic Echo Cancellation (AEC) is crucial for recovering echo-contaminated speech. This paper presents a neural network-based AEC solution to address challenges in mobile scenarios with varying hardware, nonlinear distortions and long latency. We first incorporate diverse data augmentation strategies to enhance the model’s robustness across various environments. Moreover, progressive learning is employed to incrementally improve AEC effectiveness, resulting in a considerable improvement in speech quality. To further optimize AEC’s downstream applications, we introduce a novel post-processing strategy employing tailored parameters designed specifically for tasks such as Voice Activity Detection (VAD) and Automatic Speech Recognition (ASR), thus enhancing their overall efficacy. Finally, our method employs a small-footprint model with streaming inference, enabling seamless deployment on mobile devices. Empirical results demonstrate effectiveness of the proposed method in Echo Return Loss Enhancement and Perceptual Evaluation of Speech Quality, alongside significant improvements in both VAD and ASR results. Yiheng Jiang |
ICASSP | 1 |
| 2025 | Where and When: Predict Next POI and Its Explicit Timestamp in Sequential RecommendationabstractSequential point-of-interest (POI) recommendation aims to recommend the next POI for users in accordance with their historical check-in information. However, few attempts treat timestamps of check-ins as a core factor for sequence models, leading to insufficient insight into user behavior and subsequently suboptimal recommendations. To address these limitations, we propose to assign equal importance to both POIs and their timestamps, shifting the point of view to recommend the next POI and predict the corresponding timestamp. Along these lines, we present the Time-Aware POI Recommender with Timestamp Prediction (TAPT), a multi-task learning framework for explainable POI recommendations. Specifically, we begin by decoupling timestamps into multi-dimensional vectors and propose a timestamp encoding module to explicitly encode these vectors. Additionally, we design a specialized timestamp prediction module built on the traditional sequence-based POI recommender backbone, effectively learning the strong correlation between POIs and their corresponding timestamps through these two modules. We evaluated the proposed model with three real-world LBSN datasets and demonstrated that TAPT achieves comparable or superior performance in POI recommendation compared to the baseline backbone. Besides, TAPT can not only recommend the next POI, but predict the corresponding timestamp in the future. Yuanbo Xu, Hongxu Shen, Yiheng Jiang, En Wang |
IJCAI | 3 |
| 2025 | Pushing the Frontiers of Self-Distillation Prototypes Network with Dimension Regularization and Score Normalization
Yafeng Chen, Chong Deng, Hui Wang 0030, Yiheng Jiang, Han Yin, Qian Chen 0003, Wen Wang 0001 |
INTERSPEECH | 4 |
| 2025 | Exploring Efficient Directional and Distance Cues for Regional Speech Separation
Yiheng Jiang, Haoxu Wang, Yafeng Chen, Gang Qiao |
INTERSPEECH | 1 |
| 2025 | FLASepformer: Efficient Speech Separation with Gated Focused Linear Attention Transformer
Haoxu Wang, Yiheng Jiang, Gang Qiao, Pengteng Shi |
INTERSPEECH | 2 |
| 2025 | Unlocking the Power of Diffusion Models in Sequential Recommendation: A Simple and Effective ApproachabstractIn this paper, we focus on the often-overlooked issue of embedding collapse in existing diffusion-based sequential recommendation models and propose ADRec, an innovative framework designed to mitigate this problem. Diverging from previous diffusion-based methods, ADRec applies an independent noise process to each token and performs diffusion across the entire target sequence during training. ADRec captures token interdependency through auto-regression while modeling per-token distributions through token-level diffusion. This dual approach enables the model to effectively capture both sequence dynamics and item representations, overcoming the limitations of existing methods. To further mitigate embedding collapse, we propose a three-stage training strategy: (1) pre-training the embedding weights(2) aligning these weights with the ADRec backbone, and (3) fine-tuning the model. During inference, ADRec applies the denoising process only to the last token, ensuring that the meaningful patterns in historical interactions are preserved. Our comprehensive empirical evaluation across six datasets underscores the effectiveness of ADRec in enhancing both the accuracy and efficiency of diffusion-based sequential recommendation systems. Jialei Chen 0004, Yuanbo Xu, Yiheng Jiang |
KDD (2) | 3 |
| 2025 | Transformer-Based Channel Autoregressive with Space-to-Channel Context Ordering for Learned Image CompressionabstractIn recent years, significant advancements have been made in learned image compression (LIC). In LIC frameworks, the entropy model plays a vital role in predicting the probability distribution of latent features. The transformer–based channel autoregressive entropy model has recently demonstrated impressive performance. However, it solely exploits channel–wise context while ignoring rich spatial dependencies. To address this limitation, we propose a space–to–channel context ordering approach. Specifically, we introduce a space–to–channel operation that enables transformers to jointly utilize spatial and channel context. For more effective context modeling, we further analyze the context correlations between latent feature groups and present a mixed space–channel context ordering strategy. Moreover, we incorporate zero–centered quantization and optimize the entropy coder to further enhance performance. Experimental results demonstrate that our method achieves approximately 4% and 6% rate saving on the Kodak and the Tecnick datasets compared to the original transformer–based channel autoregressive entropy model. Haotian Zhang 0009, Xiongzhuang Liang, Yiheng Jiang, Li Li 0040, Dong Liu 0002 |
VCIP | 5 |
| 2025 | Age-of-Information Analysis for Blockchain-Based Mobile Edge ComputingabstractMobile edge computing (MEC) has emerged as a disruptive paradigm that facilitates effective offloading from clouds and enables processing tasks near users. With the surge of mobile data traffic and escalating demands for responsive wireless services, it becomes imperative to enhance trust, security, and efficiency within MEC environments. Blockchain technology, renowned for its immutability, transparency, and security, has proven to be a compelling solution for securing data, enhancing supervision, and fostering trusted collaborations among heterogeneous MEC stakeholders. Despite these advantages, integrating blockchain with MEC also introduces a substantial efficiency bottleneck. Specifically, the blockchain consensus process can result in the aging of critical MEC system information, causing users to perform suboptimal service decisions, thus risking a degradation in service performance. To analyze this bottleneck, we first explore a blockchain-based MEC model that ensures secure task processing across diverse stakeholders. We employ the practical Byzantine fault tolerance (PBFT) consensus to validate and share service statuses, providing critical on-chain references for users to select their preferred target MEC servers. We then introduce the age-of-information (AoI) as a metric of freshness to characterize the aging of status reports, identifying variable consensus delay as a key factor affecting AoI. We reveal the critical impact of AoI on MEC service performance through analysis, challenging the notion that a lower AoI always leads to better performance. Finally, we validate our analysis and findings through comprehensive simulations. Yuwei Le, Yiheng Jiang, Xintong Ling, Jiaheng Wang 0001, Derrick Wing Kwan Ng, Yongming Huang 0001 |
IEEE Trans. Commun. | 3 |
| 2025 | IMOFC: Identity-Level Metric Optimized Feature Compression for Identification TasksabstractFeature compression has attracted much attention in recent years due to its promising applications in scenarios where features are transmitted and analyzed by machine vision. However, existing research mainly focuses on coarse-grained features extracted from recognition tasks such as classification and detection, neglecting fine-grained features extracted from identification tasks. In this paper, we make a pioneering attempt to study fine-grained feature compression in the context of identification tasks. Our main focus is on the distortion metric, given its critical importance in optimizing the performance of a compression network. We initiate our discussion by reviewing the instance-level metrics in existing literature, highlighting their oversight of the inter-feature relationships. The inter-feature relationships are especially important for identification tasks as they involve similarity comparison among different identities. To address this problem, we propose to consider inter-feature relationships from the perspective of identity information. Specifically, we propose an identity-level metric to incorporate both intra-identity similarity and inter-identity discriminability. The intra-identity similarity constraint aims to cluster features from the same identity, while the inter-identity discriminability constraint ensures that features from different identities deviate from each other. We implement the identity-level metric on four different feature compression networks designed based on feature characteristics. Experimental results show the effectiveness of the proposed identity-level metric on person re-identification and face verification tasks. Changsheng Gao, Yiheng Jiang, Li Li 0040, Dong Liu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Sparse Point Clouds Assisted Learned Image CompressionabstractIn the field of autonomous driving, a variety of sensor data types exist, each representing different modalities of the same scene. Therefore, it is feasible to utilize data from other sensors to facilitate image compression. However, few techniques have explored the potential benefits of utilizing inter-modality correlations to enhance the image compression performance. In this paper, motivated by the recent success of learned image compression, we propose a new framework that uses sparse point clouds to assist in learned image compression in the autonomous driving scenario. We first project the 3D sparse point cloud onto a 2D plane, resulting in a sparse depth map. Utilizing this depth map, we proceed to predict camera images. Subsequently, we use these predicted images to extract multi-scale structural features. These features are then incorporated into learned image compression pipeline as additional information to improve the compression performance. Our proposed framework is compatible with various mainstream learned image compression models, and we validate our approach using different existing image compression methods. The experimental results show that incorporating point cloud assistance into the compression pipeline consistently enhances the performance. Yiheng Jiang, Haotian Zhang 0009, Li Li 0040, Dong Liu 0002, Zhu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Algorithms for mean-field variational inference via polyhedral optimization in the Wasserstein spaceabstractWe develop a theory of finite-dimensional polyhedral subsets over the Wasserstein space and optimization of functionals over them via first-order methods. Our main application is to the problem of mean-field variational inference, which seeks to approximate a distribution $\pi$ over $\mathbb{R}^d$ by a product measure $\pi^\star$. When $\pi$ is strongly log-concave and log-smooth, we provide (1) approximation rates certifying that $\pi^\star$ is close to the minimizer $\pi^\star_\diamond$ of the KL divergence over a \emph{polyhedral} set $\mathcal{P}_\diamond$, and (2) an algorithm for minimizing $\text{KL}(\cdot\|\pi)$ over $\mathcal{P}_\diamond$ with accelerated complexity $O(\sqrt \kappa \log(\kappa d/\varepsilon^2))$, where $\kappa$ is the condition number of $\pi$. Yiheng Jiang, Sinho Chewi, Aram-Alexandre Pooladian |
COLT | 1 |
| 2024 | Skeletal Triangulation for 3D Human Pose Estimation
Yiheng Jiang, Yunlong Zhao 0001, Yang Li 0122 |
ICPR (18) | 1 |
| 2024 | DMOFC: Discrimination Metric-Optimized Feature CompressionabstractFeature compression, as an important branch of video coding for machines (VCM), has attracted significant attention and exploration. However, the existing methods mainly focus on intra-feature similarity, such as the Mean Squared Error (MSE) between the reconstructed and original features, while neglecting the importance of inter-feature relationships. In this paper, we analyze the inter-feature relationships, focusing on feature discriminability in machine vision and underscoring its significance in feature compression. To maintain the feature discriminability of reconstructed features, we introduce a dis-crimination metric for feature compression. The discrimination metric is designed to ensure that the distance between features of the same category is smaller than the distance between features of different categories. Furthermore, we explore the relationship between the discrimination metric and the discriminability of the original features. Experimental results confirm the effectiveness of the proposed discrimination metric and reveal there exists a tradeoff between the discrimination metric and the discrim-inability of the original features. Changsheng Gao, Yiheng Jiang, Li Li 0040, Dong Liu 0002 |
PCS | 2 |
| 2024 | Spatial-Temporal Interval Aware Individual Future Trajectory PredictionabstractThe past flourishing years of sequential location-based services began with the introduction of the Self-Attention Network (SAN), which quickly superseded CNN or RNN as the state-of-the-art backbone. Recent works utilize modified attention mechanisms or neural network layers to process spatial-temporal factors to realize fine-grained individual behavior pattern modeling. However, we argue these methods can be further improved due to the significant increase in the model's parameter scale or computational burden. In this paper, we first exploit two lightweight approaches, Rotary Time Aware Position Encoder (RoTAPE) and multi-head Interval Aware Attention Block (IAAB), to impel SAN by efficiently and effectively capturing spatial-temporal intervals among the user's visited locations, which require neither extra parameters nor a high computational cost. On the one hand, RoTAPE encodes the day- and hour-level timestamps into sequence representation simultaneously via a sinusoidal encoding matrix, and the corresponding time intervals can be explicitly captured by SAN. Specifically, the multi-level temporal differences are mutually independent to reflect the periodical pattern and jointly complete to measure the absolute time interval. On the other hand, IAAB, point- wise injecting the historical spatial-temporal intervals into the attention map, can promote SAN attaching importance to the spatial relations under the constraints of time conditions. Then, we design a novel MLP-based module, Spatial-Temporal Relation Memory (STR Memory), implemented with fully connected linear layers and matrix transpose operations. STR Memory, endowing the interactions inside historical intervals along different directions, can convert the historical intervals into spatial-temporal relations in future trajectories for accurate predictions. To this end, we propose an end-to-end mobility trajectory prediction framework, namely STiSAN$^+$, employing RoTAPE, stacking multiple layers of IAAB-based encoder-decoder architecture, and coupling with STR Memory. We conducted numerous experiments on six public LBSN datasets to evaluate our proposed algorithm. From Next Location Recommendation to Multi-location Future Trajectory Prediction, our STiSAN$^+$gains average 15.05% and 18.35% improvements against several state-of-the-art sequential models, respectively. Ablation studies demonstrate the effectiveness of RoTAPE, IAAB, and STR Memory under our framework. Moreover, we separately validate the extensibility and interpretability of RoTAPE and IAAB through non-sampled metric evaluation and visualization. Yiheng Jiang, Yongjian Yang 0001, Yuanbo Xu, En Wang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | TriMLP: A Foundational MLP-Like Architecture for Sequential RecommendationabstractIn this work, we present TriMLP as a foundational MLP-like architecture for the sequential recommendation, simultaneously achieving computational efficiency and promising performance. First, we empirically study the incompatibility between existing purely MLP-based models and sequential recommendation, that the inherent fully-connective structure endows historical user–item interactions (referred as tokens) with unrestricted communications and overlooks the essential chronological order in sequences. Then, we propose the MLP-based Triangular Mixer to establish ordered contact among tokens and excavate the primary sequential modeling capability under the standard auto-regressive training fashion. It contains (1) a global mixing layer that drops the lower-triangle neurons in MLP to block the anti-chronological connections from future tokens and (2) a local mixing layer that further disables specific upper-triangle neurons to split the sequence as multiple independent sessions. The mixer serially alternates these two layers to support fine-grained preferences modeling, where the global one focuses on the long-range dependency in the whole sequence, and the local one calls for the short-term patterns in sessions. Experimental results on 12 datasets of different scales from 4 benchmarks elucidate that TriMLP consistently attains favorable accuracy/efficiency tradeoff over all validated datasets, where the average performance boost against several state-of-the-art baselines achieves up to 14.88%, and the maximum reduction of inference time reaches 23.73%. The intriguing properties render TriMLP a strong contender to the well-established RNN-, CNN-, and Transformer-based sequential recommenders. Code is available at https://github.com/jiangyiheng1/TriMLP . Yiheng Jiang, Yuanbo Xu, Yongjian Yang 0001, Funing Yang, Pengyang Wang, Chaozhuo Li, Fuzhen Zhuang, Hui Xiong 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2023 | Padding-Aware Learned Image CompressionabstractFor current learned image compression methods, padding input images is necessary to meet the resolution requirements of down-sampling layers. However, the impact of padding has not been studied thoroughly. Most previous studies ignore padded images in the training process. In this paper, we analyze the impact of padding on compression performance. Then, we propose a padding-aware training (PAT) strategy, handling the padding effect during the training. Specifically, our PAT strategy calculates the loss of pre-padding image through a masking operation. Finally, according to our systematic experimental results, we find that images with different resolutions tend to favor different padding modes. Therefore, we further propose to conduct padding mode decision in the encoding process for rate-distortion optimization. Experiments demonstrate that our proposed PAT strategy and padding mode decision effectively compensate for the performance drop caused by padding. Haotian Zhang 0009, Junqi Liao, Yiheng Jiang, Li Li 0040, Dong Liu 0002 |
ISCAS | 3 |
| 2023 | Zone-Enhanced Spatio-Temporal Representation Learning for Urban POI RecommendationabstractPoints-of-interest (POIs) recommendation plays a vital role in location-based social networks (LBSNs) by introducing unexplored POIs to consumers and has drawn extensive attention from academia and industry. Existing POI recommender systems usually learn fixed latent vectors to represent both consumers and POIs from historical check-ins and make recommendations under the spatio-temporal constraints. However, we argue that the existing works still suffer from the challenges of explaining consumers’ complicated check-in actions. To this end, we first explore the interpretability of recommendations from the POI aspect, i.e., for a specific POI, its function usually changes over time, so representing a POI with a single fixed latent vector is not sufficient to describe the dynamic nature of POIs. Besides, check-in actions to a POI are also affected by the zone where it is located. In other words, the zone's embedding learned from POI distributions, road segments, and historical check-ins could be jointly utilized to enhance POI embeddings. Along this line, we propose aTime-zone-spacePOI embedding model (ToP), which integrates multi-knowledge graphs and topic model to introduce not only spatio-temporal effects but also sentiment constraints into POI embeddings for strengthening interpretability of recommendation. Specifically, ToP learns multiple latent vectors for a POI in a different period with spatial constraints via knowledge graph learning. To add sentiment constraints, ToP jointly combines these vectors with the zone's representations learned by topic models to make explainable recommendations. ToP considers the time, space, and sentiment of POI in a unified embedding framework, which benefits the POI recommendations. Extensive experiments on real-world Changchun city datasets demonstrate that ToP achieves state-of-the-art performance in terms of common metrics and provides more insights for consumers’ POI check-in actions. En Wang, Yuanbo Xu, Yongjian Yang 0001, Yiheng Jiang, Fukang Yang, Jie Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Spatial-Temporal Interval Aware Sequential POI RecommendationabstractThe past flourishing years of sequential point-of-interest (POI) recommendation began with the introduction of Self-Attention Network (SAN), which quickly superseded CNN or RNN as the state-of-the-art backbone. To realize the fine-grained users' behavior patterns modeling, recent works utilize modified attention mechanisms or neural network layers to process spatial-temporal factors. However, due to the significant increase on either model's parameter scale or computational burden, we argue that these methods can be further improved. In this paper, we exploit two lightweight approaches, Time Aware Position Encoder (TAPE) and Interval Aware Attention Block (IAAB), to impel SAN by considering the spatial-temporal intervals among POIs separately, where requiring neither extra parameters nor high computational cost. On the one hand, TAPE, adjusting the positions in sequences based on the timestamps dynamically and generating positional representations with sinusoidal transformation, can enhance sequence representations to reflect both the absolute order and relative temporal proximity among all POIs. On the other hand, IAAB, point-wise adding the scaled spatial-temporal intervals to the attention map, can promote the attention mechanism attaching importance to the spatial relation among all POIs under the constraints of time conditions and providing more explainable recommendation. We integrate these two modules into SAN and propose a Spatial-Temporal Interval-Aware sequential POI recommender, namely STiSAN, as an end-to-end deployment. Experimental results based on three public LBSN datasets and one real-world city transportation dataset demonstrate STiSAN's superior performance (average 13.01% improvement against the strongest baseline). Moreover, we validate the extensibility and interpretability of TAPE and IAAB through metric evaluation and visualization separately. En Wang, Yiheng Jiang, Yuanbo Xu, Liang Wang 0017, Yongjian Yang 0001 |
ICDE | 2 |
| 2021 | ToP: Time-dependent Zone-enhanced Points-of-interest Embedding-based Explainable Recommender systemabstractPoints-of-interest (POIs) recommendation plays a vital role by introducing unexplored POIs to consumers and has drawn extensive attention from both academia and industry. Existing POI recommender systems usually learn latent vectors to represent both consumers and POIs from historical check-ins and make recommendations under the spatiotemporal constraints. However, we argue that the existing works still suffer from the challenges of explaining consumers complicated check-in actions. In this paper, we first explore the interpretability of recommendations from the POI aspect, i.e., for a specific POI, its function usually changes over time, so representing a POI with a single fixed latent vector is not sufficient to describe POIs dynamic function. Besides, check-in actions to a POI is also affected by the zone it belongs to. In other words, the zone's embedding learned from POI distributions, road segments, and historical check-ins could be jointly utilized to enhance the accuracy of POI recommendations. Along this line, we propose a Time-dependent Zone-enhanced POI embedding model (ToP), a recommender system that integrates knowledge graph and topic model to introduce the spatiotemporal effects into POI embeddings for strengthening interpretability of recommendation. Specifically, ToP learns multiple latent vectors for a POI in different time to capture its dynamic functions. Jointly combining these vectors with zones representations, ToP enhances the spatiotemporal interpretability of POI recommendations. With this hybrid architecture, some existing POI recommender systems can be treated as special cases of ToP. Extensive experiments on real-world Changchun city datasets demonstrate that ToP not only achieves state-of-the-art performance in terms of common metrics, but also provides more insights for consumers POI check-in actions. En Wang, Yuanbo Xu, Yongjian Yang 0001, Fukang Yang, Yiheng Jiang |
INFOCOM | 6 |
| 2020 | An Effective Speaker Recognition Method Based on Joint Identification and Verification SupervisionsabstractDeep embedding learning based speaker verification methods have attracted significant recent research interest due to their superior performance. Existing methods mainly focus on designing frame-level feature extraction structures, utterance-level aggregation methods and loss functions to learn discriminative speaker embeddings. The scores of verification trials are then computed using cosine distance or Probabilistic Linear Discriminative Analysis (PLDA) classifiers. This paper proposes an effective speaker recognition method which is based on joint identification and verification supervisions, inspired by multi-task learning frameworks. Specifically, a deep architecture with convolutional feature extractor, attentive pooling and two classifier branches is presented. The first, an identification branch, is trained with additive margin softmax loss (AM-Softmax) to classify the speaker identities. The second, a verification branch, trains a discriminator with binary cross entropy loss (BCE) to optimize a new triplet-based mutual information. To balance the two losses during different training stages, a ramp-up/ramp-down weighting scheme is employed. Furthermore, an attentive bilinear pooling method is proposed to improve the effectiveness of embeddings. Extensive experiments have been conducted on VoxCeleb1 to evaluate the proposed method, demonstrating results that relatively reduce the equal error rate (EER) by 22% compared to the baseline system using identification supervision only. Yan Song 0001, Yiheng Jiang, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001 |
INTERSPEECH | 3 |
| 2019 | Improving Aggregation and Loss Function for Better Embedding Learning in End-to-End Speaker Verification System
Zhifu Gao, Yan Song 0001, Ian McLoughlin 0001, Yiheng Jiang, Li-Rong Dai 0001 |
INTERSPEECH | 5 |
| 2019 | An Effective Deep Embedding Learning Architecture for Speaker Verification
Yiheng Jiang, Yan Song 0001, Ian McLoughlin 0001, Zhifu Gao, Li-Rong Dai 0001 |
INTERSPEECH | 1 |