Taoxing Pan

dblp:255/9328 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0003-1577-0200ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 50% Efficient and distributed learning · 33% Optimization for machine learning · 17%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.912025
Mutual Information-aware Knowledge Distillation for Short Video Recommendation · KDD (1) 2025
Recommender systems › video recommendation
short-video recommendation
0.912025
Mutual Information-aware Knowledge Distillation for Short Video Recommendation · KDD (1) 2025
Machine learning › Reinforcement learning
exploration
0.712023
Efficient Exploration in Resource-Restricted Reinforcement Learning · AAAI 2023
Machine learning › Reinforcement learning › exploration
exploration bonus
0.712023
Efficient Exploration in Resource-Restricted Reinforcement Learning · AAAI 2023
Machine learning › Optimization for machine learning
stochastic gradient methods
0.412020
D-SPIDER-SFO: A Decentralized Optimization Algorithm with Faster Convergence Rate for Nonconvex Problems · AAAI 2020
Mathematical optimization › distributed optimization
decentralized optimization
0.412020
D-SPIDER-SFO: A Decentralized Optimization Algorithm with Faster Convergence Rate for Nonconvex Problems · AAAI 2020
Mathematical optimization
nonconvex optimization
0.412020
D-SPIDER-SFO: A Decentralized Optimization Algorithm with Faster Convergence Rate for Nonconvex Problems · AAAI 2020
Recommender systems › user modeling
user feedback modeling
0.312025
Mutual Information-aware Knowledge Distillation for Short Video Recommendation · KDD (1) 2025
Distributed systems › distributed system architecture
decentralized network
0.112020
D-SPIDER-SFO: A Decentralized Optimization Algorithm with Faster Convergence Rate for Nonconvex Problems · AAAI 2020

Methods — techniques the papers use, named apart from their topics

mutual information estimation · 1.7knowledge distillation · 1.7variance reduction · 1.3gradient estimation · 1.3SPIDER-SFO · 1.3soft actor-critic · 0.7
YearPublicationVenuePosition
2025 Mutual Information-aware Knowledge Distillation for Short Video Recommendation
abstract
Short-video sharing platforms engaging billions of users have attracted intense interest recently. A key insight is that user feedback on these platforms is heavily influenced by preceding exposed videos in the same request, called context cumulative effects. For example, multiple repeated videos in a request often cause user fatigue and influence user feedback. However, related factors, such as the other exposed items in the same request, are available during model training but not accessible during online serving. Vanilla distillation methods mitigate the training-inference inconsistency, struggling to capture the dynamic dependence between context cumulative effects and user feedback. To address this problem, we propose the Mutual Information-aware Knowledge Distillation (MIKD) framework, which fuses such effects and user-item matching degrees by evaluating their impacts on user feedback based on mutual information estimation. Rigorous analysis and extensive experiments demonstrate that MIKD precisely extracts personal interests and consistently improves performance. We conduct online A/B testing on a leading short-video sharing mobile app, and the results demonstrate the effectiveness of the proposed method. MIKD has been successfully deployed online to serve the main traffic and optimize user experiences.
Taoxing Pan
KDD (1)2
2023 Efficient Exploration in Resource-Restricted Reinforcement Learning
abstract
In many real-world applications of reinforcement learning (RL), performing actions requires consuming certain types of resources that are non-replenishable in each episode. Typical applications include robotic control with limited energy and video games with consumable items. In tasks with non-replenishable resources, we observe that popular RL methods such as soft actor critic suffer from poor sample efficiency. The major reason is that, they tend to exhaust resources fast and thus the subsequent exploration is severely restricted due to the absence of resources. To address this challenge, we first formalize the aforementioned problem as a resource-restricted reinforcement learning, and then propose a novel resource-aware exploration bonus (RAEB) to make reasonable usage of resources. An appealing feature of RAEB is that, it can significantly reduce unnecessary resource-consuming trials while effectively encouraging the agent to explore unvisited states. Experiments demonstrate that the proposed RAEB significantly outperforms state-of-the-art exploration strategies in resource-restricted reinforcement learning environments, improving the sample efficiency by up to an order of magnitude.
Taoxing Pan, Qi Zhou 0008, Jie Wang 0005
AAAI2
2020 D-SPIDER-SFO: A Decentralized Optimization Algorithm with Faster Convergence Rate for Nonconvex Problems
abstract
Decentralized optimization algorithms have attracted intensive interests recently, as it has a balanced communication pattern, especially when solving large-scale machine learning problems. Stochastic Path Integrated Differential Estimator Stochastic First-Order method (SPIDER-SFO) nearly achieves the algorithmic lower bound in certain regimes for nonconvex problems. However, whether we can find a decentralized algorithm which achieves a similar convergence rate to SPIDER-SFO is still unclear. To tackle this problem, we propose a decentralized variant of SPIDER-SFO, called decentralized SPIDER-SFO (D-SPIDER-SFO). We show that D-SPIDER-SFO achieves a similar gradient computation cost—that is, O(ε−3) for finding an ϵ-approximate first-order stationary point—to its centralized counterpart. To the best of our knowledge, D-SPIDER-SFO achieves the state-of-the-art performance for solving nonconvex optimization problems on decentralized networks in terms of the computational cost. Experiments on different network configurations demonstrate the efficiency of the proposed method.
Taoxing Pan, Jun Liu 0003, Jie Wang 0005
AAAI1