Kai Li 0022

dblp:181/2853-22 · DBLP profile ↗
← Back
46ranked-venue papers
7as first author
36since 2021 · last 2026
0000-0003-3840-3270ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 5 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 8 since 2021Computer networks · 6 · 6 since 2021
YearPublicationVenuePosition
2026 Deep (Predictive) Discounted Counterfactual Regret Minimization
abstract
Counterfactual regret minimization (CFR) is a family of algorithms for effectively solving imperfect-information games. To enhance CFR's applicability in large games, researchers use neural networks to approximate its behavior. However, existing methods are mainly based on vanilla CFR and struggle to effectively integrate more advanced CFR variants. In this work, we propose an efficient model-free neural CFR algorithm, overcoming the limitations of existing methods in approximating advanced CFR variants. At each iteration, it collects variance-reduced sampled advantages based on a value network, fits cumulative advantages by bootstrapping, and applies discounting and clipping operations to simulate the update mechanisms of advanced CFR variants. Experimental results show that, compared with model-free neural algorithms, it exhibits faster convergence in typical imperfect-information games and demonstrates stronger adversarial performance in a large poker game.
Hang Xu 0006, Kai Li 0022, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
AAAI2
2026 Timely Requesting for Time-Critical Content Users in Decentralized F-RANs
abstract
With the rising demand for high-rate and timely communications, fog radio access networks (F-RANs) offer a promising solution. This work investigates age of information (AoI) performance in F-RANs, consisting of multiple content users (CUs), enhanced remote radio heads (eRRHs), and content providers (CPs). Time-critical CUs need rapid content updates from CPs but cannot communicate directly with them; instead, eRRHs act as intermediaries. CUs decide whether to request content from a CP and which eRRH to send the request to, while eRRHs decide whether to command CPs to update content or use cached content. We study two broad classes of policies: (i) oblivious policies, where decision-making is independent of historical information, and (ii) non-oblivious policies, where decisions are influenced by historical information. We first derive closed-form expressions for the average AoI of eRRHs under both policy types. Due to the complexity of calculating closed-form expressions for CUs, we then derive general upper bounds for their average AoI. Next, we identify optimal policies for both types. Under both optimal policies, each CU requests content from each CP at an equal rate. When demand is low or resources are limited, all requests are consolidated to a single eRRH; when demand is high and resources are ample, requests are evenly distributed among eRRHs. eRRHs command content from each CP at an equal rate under an optimal oblivious policy, while prioritize the CP with the highest age under an optimal non-oblivious policy. Our numerical results validate these theoretical findings. We further extend our analytical framework to two generalized scenarios, and simulations confirm the validity of our conclusions.
Xingran Chen, Kai Li 0022, Kun Yang 0001
IEEE Trans. Netw.2
2025 An Open-Ended Learning Framework for Opponent Modeling
abstract
Opponent Modeling (OM) aims to enhance decision-making by modeling other agents in multi-agent environments. Existing works typically learn opponent models against a pre-designated fixed set of opponents during training. However, this will cause poor generalization when facing unknown opponents during testing, as previously unseen opponents can exhibit out-of-distribution (OOD) behaviors that the learned opponent models cannot handle. To tackle this problem, we introduce a novel Open-Ended Opponent Modeling (OEOM) framework, which continuously generates opponents with diverse strengths and styles to reduce the possibility of OOD situations occurring during testing. Founded on population-based training and information-theoretic trajectory space diversity regularization, OEOM generates a dynamic set of opponents. This set is then fed to any OM approaches to train a potentially generalizable opponent model. Upon this, we further propose a simple yet effective OM approach that naturally fits within the OEOM framework. This approach is based on in-context reinforcement learning and learns a Transformer that dynamically recognizes and responds to opponents based on their trajectories. Extensive experiments in cooperative, competitive, and mixed environments demonstrate that OEOM is an approach-agnostic framework that improves generalizability compared to training against a fixed set of opponents, regardless of OM approaches or testing opponent settings. The results also indicate that our proposed approach generally outperforms existing OM baselines.
Yuheng Jing, Kai Li 0022, Bingyun Liu, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
AAAI2
2025 Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
abstract
Recovering a spectrum of diverse policies from a set of expert trajectories is an important research topic in imitation learning. After determining a latent style for a trajectory, previous diverse polices recovering methods usually employ a vanilla behavioral cloning learning objective conditioned on the latent style, treating each state-action pair in the trajectory with equal importance. Based on an observation that in many scenarios, behavioral styles are often highly relevant with only a subset of state-action pairs, this paper presents a new principled method in diverse polices recovering. In particular, after inferring or assigning a latent style for a trajectory, we enhance the vanilla behavioral cloning by incorporating a weighting mechanism based on pointwise mutual information. This additional weighting reflects the significance of each state-action pair's contribution to learning the style, thus allowing our method to focus on state-action pairs most representative of that style. We provide theoretical justifications for our new objective, and extensive empirical evaluations confirm the effectiveness of our method in recovering diverse polices from expert data.
Jian Yao 0008, Weiming Liu 0004, Hanmin Qin, Hansheng Kong, Kirk Tang, Jiechao Xiong, Chao Yu 0004, Kai Li 0022, Junliang Xing, Hongwu Chen, Juchao Zhuo, Qiang Fu 0016, Haobo Fu
ICLR10
2025 Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement Learning
abstract
Offline multi-task reinforcement learning aims to learn a unified policy capable of solving multiple tasks using only pre-collected task-mixed datasets, without requiring any online interaction with the environment. However, it faces significant challenges in effectively sharing knowledge across tasks. Inspired by the efficient knowledge abstraction observed in human learning, we propose Goal-Oriented Skill Abstraction (GO-Skill), a novel approach designed to extract and utilize reusable skills to enhance knowledge transfer and task performance. Our approach uncovers reusable skills through a goal-oriented skill extraction process and leverages vector quantization to construct a discrete skill library. To mitigate class imbalances between broadly applicable and task-specific skills, we introduce a skill enhancement phase to refine the extracted skills. Furthermore, we integrate these skills using hierarchical policy learning, enabling the construction of a high-level policy that dynamically orchestrates discrete skills to accomplish specific tasks. Extensive experiments on diverse robotic manipulation tasks within the MetaWorld benchmark demonstrate the effectiveness and versatility of GO-Skill.
Jinmin He, Kai Li 0022, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
ICML2
2025 Offline Opponent Modeling with Truncated Q-driven Instant Policy Refinement
abstract
Offline Opponent Modeling (OOM) aims to learn an adaptive autonomous agent policy that dynamically adapts to opponents using an offline dataset from multi-agent games. Previous work assumes that the dataset is optimal. However, this assumption is difficult to satisfy in the real world. When the dataset is suboptimal, existing approaches struggle to work. To tackle this issue, we propose a simple and general algorithmic improvement framework, Truncated Q-driven Instant Policy Refinement (TIPR), to handle the suboptimality of OOM algorithms induced by datasets. The TIPR framework is plug-and-play in nature. Compared to original OOM algorithms, it requires only two extra steps: (1) Learn a horizon-truncated in-context action-value function, namely Truncated Q, using the offline dataset. The Truncated Q estimates the expected return within a fixed, truncated horizon and is conditioned on opponent information. (2) Use the learned Truncated Q to instantly decide whether to perform policy refinement and to generate policy after refinement during testing. Theoretically, we analyze the rationale of Truncated Q from the perspective of No Maximization Bias probability. Empirically, we conduct extensive comparison and ablation experiments in four representative competitive environments. TIPR effectively improves various OOM algorithms pretrained with suboptimal datasets.
Yuheng Jing, Kai Li 0022, Bingyun Liu, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
ICML2
2025 Pattern Extraction Learning for Cooperative Multi-Agent Reinforcement Learning
abstract
Multi-agent reinforcement learning (MARL) has demonstrated its superiority in addressing complex decision-making tasks involving multiple agents. However, the intricate and dynamic interactions among agents make this problem exceptionally challenging. Existing MARL methods often simplify the problem by implicitly decomposing shared rewards into individual utilities, neglecting the underlying interconnections between relevant entities. To overcome these limitations, we propose a novel framework of Multi-Agent Pattern Extraction (MAPE), which captures cooperation patterns from both agent-level and global perspectives to enhance decision-making and collaboration efficiency. Specifically, MAPE introduces two key modules: the Agent Pattern Extractor (APE) and the Global Pattern Extractor (GPE) that focus on specific interactions between the entities of interests from individual and global perspectives, respectively. The APE module focuses on computing each agent’s attention to other entities across different interaction patterns, providing this information to the GPE module. The GPE module then integrates the agent-specific pattern information and state features to identify the overall interaction pattern of the multi-agent system. By filtering out irrelevant interactions between unrelated entities and highlighting meaningful relationships, MAPE fosters more focused cooperation and facilitates more efficient learning. Extensive experiments on the StarCraft II micromanagement benchmark showcase the effectiveness of MAPE in improving efficiency in complex multi-agent environments.
Yifan Zang 0001, Jinmin He, Kai Li 0022, Junliang Xing, Jian Cheng 0001
IJCNN3
2024 Multiagent Gumbel MuZero: Efficient Planning in Combinatorial Action Spaces
abstract
AlphaZero and MuZero have achieved state-of-the-art (SOTA) performance in a wide range of domains, including board games and robotics, with discrete and continuous action spaces. However, to obtain an improved policy, they often require an excessively large number of simulations, especially for domains with large action spaces. As the simulation budget decreases, their performance drops significantly. In addition, many important real-world applications have combinatorial (or exponential) action spaces, making it infeasible to search directly over all possible actions. In this paper, we extend AlphaZero and MuZero to learn and plan in more complex multiagent (MA) Markov decision processes, where the action spaces increase exponentially with the number of agents. Our new algorithms, MA Gumbel AlphaZero and MA Gumbel MuZero, respectively without and with model learning, achieve superior performance on cooperative multiagent control problems, while reducing the number of environmental interactions by up to an order of magnitude compared to model-free approaches. In particular, we significantly improve prior performance when planning with much fewer simulation budgets. The code and appendix are available at https://github.com/tjuHaoXiaotian/MA-MuZero.
Xiaotian Hao, Jianye Hao, Chenjun Xiao, Kai Li 0022, Dong Li 0016, Yan Zheng 0002
AAAI4
2024 Not All Tasks Are Equally Difficult: Multi-Task Deep Reinforcement Learning with Dynamic Depth Routing
abstract
Multi-task reinforcement learning endeavors to accomplish a set of different tasks with a single policy. To enhance data efficiency by sharing parameters across multiple tasks, a common practice segments the network into distinct modules and trains a routing network to recombine these modules into task-specific policies. However, existing routing approaches employ a fixed number of modules for all tasks, neglecting that tasks with varying difficulties commonly require varying amounts of knowledge. This work presents a Dynamic Depth Routing (D2R) framework, which learns strategic skipping of certain intermediate modules, thereby flexibly choosing different numbers of modules for each task. Under this framework, we further introduce a ResRouting method to address the issue of disparate routing paths between behavior and target policies during off-policy training. In addition, we design an automatic route-balancing mechanism to encourage continued routing exploration for unmastered tasks without disturbing the routing of mastered ones. We conduct extensive experiments on various robotics manipulation tasks in the Meta-World benchmark, where D2R achieves state-of-the-art performance with significantly improved learning efficiency.
Jinmin He, Kai Li 0022, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
AAAI2
2024 Towards Offline Opponent Modeling with In-context Learning
abstract
Opponent modeling aims at learning the opponent's behaviors, goals, or beliefs to reduce the uncertainty of the competitive environment and assist decision-making. Existing work has mostly focused on learning opponent models online, which is impractical and inefficient in practical scenarios. To this end, we formalize an Offline Opponent Modeling (OOM) problem with the objective of utilizing pre-collected offline datasets to learn opponent models that characterize the opponent from the viewpoint of the controlled agent, which aids in adapting to the unknown fixed policies of the opponent. Drawing on the promises of the Transformers for decision-making, we introduce a general approach, Transformer Against Opponent (TAO), for OOM. Essentially, TAO tackles the problem by harnessing the full potential of the supervised pre-trained Transformers' in-context learning capabilities. The foundation of TAO lies in three stages: an innovative offline policy embedding learning stage, an offline opponent-aware response policy training stage, and a deployment stage for opponent adaptation with in-context learning. Theoretical analysis establishes TAO's equivalence to Bayesian posterior sampling in opponent modeling and guarantees TAO's convergence in opponent policy recognition. Extensive experiments and ablation studies on competitive environments with sparse and dense rewards demonstrate the impressive performance of TAO. Our approach manifests remarkable prowess for fast adaptation, especially in the face of unseen opponent policies, confirming its in-context learning potency.
Yuheng Jing, Kai Li 0022, Bingyun Liu, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
ICLR2
2024 Dynamic Discounted Counterfactual Regret Minimization
abstract
Counterfactual regret minimization (CFR) is a family of iterative algorithms showing promising results in solving imperfect-information games. Recent novel CFR variants (e.g., CFR+, DCFR) have significantly improved the convergence rate of the vanilla CFR. The key to these CFR variants’ performance is weighting each iteration non-uniformly, i.e., discounting earlier iterations. However, these algorithms use a fixed, manually-specified scheme to weight each iteration, which enormously limits their potential. In this work, we propose Dynamic Discounted CFR (DDCFR), the first equilibrium-finding framework that discounts prior iterations using a dynamic, automatically-learned scheme. We formalize CFR’s iteration process as a carefully designed Markov decision process and transform the discounting scheme learning problem into a policy optimization problem within it. The learned discounting scheme dynamically weights each iteration on the fly using information available at runtime. Experimental results across multiple games demonstrate that DDCFR’s dynamic discounting scheme has a strong generalization ability and leads to faster convergence with improved performance. The code is available at https://github.com/rpSebastian/DDCFR.
Hang Xu 0006, Kai Li 0022, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
ICLR2
2024 Minimizing Weighted Counterfactual Regret with Optimistic Online Mirror Descent
Hang Xu 0006, Kai Li 0022, Bingyun Liu, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
IJCAI2
2024 Efficient Multi-task Reinforcement Learning with Cross-Task Policy Guidance
abstract
Multi-task reinforcement learning endeavors to efficiently leverage shared information across various tasks, facilitating the simultaneous learning of multiple tasks. Existing approaches primarily focus on parameter sharing with carefully designed network structures or tailored optimization procedures. However, they overlook a direct and complementary way to exploit cross-task similarities: the control policies of tasks already proficient in some skills can provide explicit guidance for unmastered tasks to accelerate skills acquisition. To this end, we present a novel framework called Cross-Task Policy Guidance (CTPG), which trains a guide policy for each task to select the behavior policy interacting with the environment from all tasks' control policies, generating better training trajectories. In addition, we propose two gating mechanisms to improve the learning efficiency of CTPG: one gate filters out control policies that are not beneficial for guidance, while the other gate blocks tasks that do not necessitate guidance. CTPG is a general framework adaptable to existing parameter sharing approaches. Empirical evaluations demonstrate that incorporating CTPG with these approaches significantly enhances performance in manipulation and locomotion benchmarks.
Jinmin He, Kai Li 0022, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
NeurIPS2
2024 Opponent Modeling with In-context Search
abstract
Opponent modeling is a longstanding research topic aimed at enhancing decision-making by modeling information about opponents in multi-agent environments. However, existing approaches often face challenges such as having difficulty generalizing to unknown opponent policies and conducting unstable performance. To tackle these challenges, we propose a novel approach based on in-context learning and decision-time search named Opponent Modeling with In-context Search (OMIS). OMIS leverages in-context learning-based pretraining to train a Transformer model for decision-making. It consists of three in-context components: an actor learning best responses to opponent policies, an opponent imitator mimicking opponent actions, and a critic estimating state values. When testing in an environment that features unknown non-stationary opponent agents, OMIS uses pretrained in-context components for decision-time search to refine the actor's policy. Theoretically, we prove that under reasonable assumptions, OMIS without search converges in opponent policy recognition and has good generalization properties; with search, OMIS provides improvement guarantees, exhibiting performance stability. Empirically, in competitive, cooperative, and mixed environments, OMIS demonstrates more effective and stable adaptation to opponents than other approaches. See our project website at https://sites.google.com/view/nips2024-omis.
Yuheng Jing, Bingyun Liu, Kai Li 0022, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
NeurIPS3
2024 Selection Strategy Based on Proper Pareto Optimality in Evolutionary Multi-objective Optimization
Kai Li 0022, Kangnian Lin, Ruihao Zheng, Zhenkun Wang 0001
PPSN (4)1
2024 Automatically designing counterfactual regret minimization algorithms for solving imperfect-information games
Kai Li 0022, Hang Xu 0006, Haobo Fu, Qiang Fu 0016, Junliang Xing
Artif. Intell.1
2024 Cooperative Multiagent Transfer Learning With Coalition Pattern Decomposition
abstract
Knowledge transfer in cooperative multi-agent reinforcement learning (MARL) has drawn increasing attention in recent years. Unlike generalizing policies in single-agent tasks, it is more important to consider coordination knowledge than individual knowledge in multi-agent transfer learning. However, most of the existing methods only focus on knowledge transfer of the individual agent policy, which leads to coordination bias and finally affects the final performance in cooperative MARL. In this paper, we propose a level-adaptive MARL framework called “LA-QTransformer”, to realize the knowledge transfer on the coordination level via efficiently decomposing the agent coordination into multi-level coalition patterns for different agents. Compatible with centralized training with decentralized execution (CTDE) regime, LA-QTransformer utilizes the Level- Adaptive Transformer to generate suitable coalition patterns and then realizes the credit assignment for each agent. Besides, to deal with unexpected changes in the number of agents in the coordination transfer phase, we design a policy network called “Population invariant agent with Transformer (PIT)” to adapt dynamic observation and action space. We evaluate the LAQTransformer and PIT in the StarCraft II micro-management benchmark by comparing them with several state-of-the-art MARL baselines. The experimental results demonstrate the superiority of LA-QTransformer and PIT and verify the feasibility of coordination knowledge transfer.
Tianze Zhou, Fubiao Zhang, Kun Shao, Zipeng Dai, Kai Li 0022, Wenhan Huang, Weixun Wang, Bin Wang 0034, Dong Li 0016, Wulong Liu, Jianye Hao
IEEE Trans. Games5
2024 Self-Supervised MAFENN for Classifying Low-Labeled Distorted Images Over Mobile Fading Channels
abstract
Image distortion during wireless transmission presents a significant challenge for real-world artificial intelligence (AI) applications. Recent methods have attempted to address this issue by integrating neural networks into the wireless transmission system. However, these approaches often require a large volume of labeled training data, which can be expensive and time-consuming to collect. To address this issue, we propose a novel approach,Self-SupervisedMulti-AgentFeedbackEnabledNeuralNetworks (S2MAFENN). S2MAFENN is designed to improve the efficiency of labeled data in wireless image transmission. It incorporates a Feedbacker agent that emulates the error correction mechanisms observed in primate brains and employs self-supervised contrastive learning to extract representations from unlabeled distorted images independently. From a theoretical perspective, we model the training process of S2MAFENN as a three-player Stackelberg game and provide evidence that S2MAFENN can achieve exponential convergence rates. We then empirically validate our approach by assessing the representations learned through S2MAFENN. We use varied labeled CIFAR10 and CIFAR100 data to simulate real image transmissions over the Rayleigh fading and 5G channels. Our results show that S2MAFENN matches or even surpasses the performance of state-of-the-art self-supervised training methods, even when only 50% of labels are used. Moreover, S2MAFENN yields average accuracy gains of 5.11%, 5.8%, and 4.58% with only 0.1, 0.2, and 0.5 of the labels transmitted over the 5G channel, respectively. For the downstream task of semantic segmentation over the 5G channel, S2MAFENN exhibits significant advancements on the ADE20K dataset. It achieves enhancements of approximately 7% and 8.7% in Mean IoU and DICE metrics, respectively, surpassing the performance of current state-of-the-art methods.
Yang Li 0116, Fanglei Sun, Jingchen Hu, Fan Wu 0006, Kai Li 0022, Ying Wen 0001, Zheng Tian 0002, Yaodong Yang 0001, Jiangcheng Zhu, Jun Wang 0012, Yang Yang 0001
IEEE Trans. Mob. Comput.6
2024 OpenHoldem: A Benchmark for Large-Scale Imperfect-Information Game Research
abstract
Owing to the unremitting efforts from a few institutes, researchers have recently made significant progress in designing superhuman artificial intelligence (AI) in no-limit Texas hold'em (NLTH), the primary testbed for large-scale imperfect-information game research. However, it remains challenging for new researchers to study this problem since there are no standard benchmarks for comparing with existing methods, which hinders further developments in this research area. This work presents OpenHoldem, an integrated benchmark for large-scale imperfect-information game research using NLTH. OpenHoldem makes three main contributions to this research direction: 1) a standardized evaluation protocol for thoroughly evaluating different NLTH AIs; 2) four publicly available strong baselines for NLTH AI; and 3) an online testing platform with easy-to-use APIs for public NLTH AI evaluation. We will publicly release OpenHoldem and hope it facilitates further studies on the unsolved theoretical and computational issues in this area and cultivates crucial research problems like opponent modeling and human-computer interactive learning.
Kai Li 0022, Hang Xu 0006, Enmin Zhao, Junliang Xing
IEEE Trans. Neural Networks Learn. Syst.1
2023 Pseudo Value Network Distillation for High-Performance Exploration
abstract
Solving hard exploration tasks with sparse rewards is notoriously challenging in reinforcement learning (RL), which needs to address two key issues simultaneously: exploiting past successful experiences and exploring the unknown environment. Many prior works take expert demonstrations as successful experiences and learn to imitate them directly. However, these demonstrations are often not available in practice. Recently, curiosity-driven RL methods provide intrinsic rewards, encouraging the agent to explore states with high novelty. Nonetheless, they lack a mechanism for leveraging past good experiences effectively. This work presents a Pseudo Value Network Distillation (PVND) framework to balance the RL agent's exploitative and exploratory behaviors effectively and automatically. In particular, PVND learns to set high exploitation bonuses to the critical states in rewarded trajectories from past experiences and high exploration bonuses to the novel states that agents rarely visit during exploration. We theoretically demonstrate that PVND gives larger positive intrinsic rewards to more critical states. Furthermore, PVND automatically finds meaningful and critical hierarchical sub-tasks for agents to accomplish the final goal progressively. Competitive results in several hard exploration sparse reward problems have verified its effectiveness and efficiency.
Enmin Zhao, Junliang Xing, Kai Li 0022, Yongxin Kang, Pin Tao
IJCNN3
2023 Automatic Grouping for Efficient Cooperative Multi-Agent Reinforcement Learning
abstract
Grouping is ubiquitous in natural systems and is essential for promoting efficiency in team coordination. This paper proposes a novel formulation of Group-oriented Multi-Agent Reinforcement Learning (GoMARL), which learns automatic grouping without domain knowledge for efficient cooperation. In contrast to existing approaches that attempt to directly learn the complex relationship between the joint action-values and individual utilities, we empower subgroups as a bridge to model the connection between small sets of agents and encourage cooperation among them, thereby improving the learning efficiency of the whole team. In particular, we factorize the joint action-values as a combination of group-wise values, which guide agents to improve their policies in a fine-grained fashion. We present an automatic grouping mechanism to generate dynamic groups and group action-values. We further introduce a hierarchical control for policy learning that drives the agents in the same group to specialize in similar policies and possess diverse strategies for various groups. Experiments on the StarCraft II micromanagement tasks and Google Research Football scenarios verify our method's effectiveness. Extensive component studies show how grouping works and enhances performance.
Yifan Zang 0001, Jinmin He, Kai Li 0022, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
NeurIPS3
2022 AutoCFR: Learning to Design Counterfactual Regret Minimization Algorithms
abstract
Counterfactual regret minimization (CFR) is the most commonly used algorithm to approximately solving two-player zero-sum imperfect-information games (IIGs). In recent years, a series of novel CFR variants such as CFR+, Linear CFR, DCFR have been proposed and have significantly improved the convergence rate of the vanilla CFR. However, most of these new variants are hand-designed by researchers through trial and error based on different motivations, which generally requires a tremendous amount of efforts and insights. This work proposes to meta-learn novel CFR algorithms through evolution to ease the burden of manual algorithm design. We first design a search language that is rich enough to represent many existing hand-designed CFR variants. We then exploit a scalable regularized evolution algorithm with a bag of acceleration techniques to efficiently search over the combinatorial space of algorithms defined by this language. The learned novel CFR algorithm can generalize to new IIGs not seen during training and performs on par with or better than existing state-of-the-art CFR variants. The code is available at https://github.com/rpSebastian/AutoCFR.
Hang Xu 0006, Kai Li 0022, Haobo Fu, Qiang Fu 0016, Junliang Xing
AAAI2
2022 AlphaHoldem: High-Performance Artificial Intelligence for Heads-Up No-Limit Poker via End-to-End Reinforcement Learning
abstract
Heads-up no-limit Texas hold’em (HUNL) is the quintessential game with imperfect information. Representative priorworks like DeepStack and Libratus heavily rely on counter-factual regret minimization (CFR) and its variants to tackleHUNL. However, the prohibitive computation cost of CFRiteration makes it difficult for subsequent researchers to learnthe CFR model in HUNL and apply it in other practical applications. In this work, we present AlphaHoldem, a high-performance and lightweight HUNL AI obtained with an end-to-end self-play reinforcement learning framework. The proposed framework adopts a pseudo-siamese architecture to directly learn from the input state information to the output actions by competing the learned model with its different historical versions. The main technical contributions include anovel state representation of card and betting information, amultitask self-play training loss function, and a new modelevaluation and selection metric to generate the final model.In a study involving 100,000 hands of poker, AlphaHoldemdefeats Slumbot and DeepStack using only one PC with threedays training. At the same time, AlphaHoldem only takes 2.9milliseconds for each decision-making using only a singleGPU, more than 1,000 times faster than DeepStack. We release the history data among among AlphaHoldem, Slumbot,and top human professionals in the author’s GitHub repository to facilitate further studies in this direction.
Enmin Zhao, Renye Yan, Jinqiu Li, Kai Li 0022, Junliang Xing
AAAI4
2022 DBM: Delay-sensitive Buffering Mechanism for DNN Offloading Services
abstract
DNN offloading has become an important supporting technology for edge intelligence. However, most of the existing works do not consider thread scheduling, which can achieve the parallelism of multiple threads in the practical distributed DNN inference system. To address this issue, we discuss the thread scheduling of the computing units participating in offloading in this paper, considering a single-core Central Processing Unit (CPU) and the Round Robin Scheduling (RRS). We deduce the relationship between the blocking of DNN inference-related threads and the Average Task Delay (ATD) and prove that an appropriate buffer setting can reduce blocking times. Theoretical analysis verifies that the buffering mechanism (DBM) can reduce the ATD significantly, and experimental results demonstrate that the DBM-improved DNN offloading can achieve a delay reduction of 14%-71%.
Guoliang Gao, Liantao Wu, Yang Yang 0001, Kai Li 0022
APCC4
2022 Actor-Critic Policy Optimization in a Large-Scale Imperfect-Information Game
Haobo Fu, Weiming Liu 0004, Kai Li 0022, Junliang Xing, Bin Li 0025, Qiang Fu 0016, Wei Yang 0032
ICLR6
2022 Greedy when Sure and Conservative when Uncertain about the Opponents
abstract
We develop a new approach, named Greedy when Sure and Conservative when Uncertain (GSCU), to competing online against unknown and nonstationary opponents. GSCU improves in four aspects: 1) introduces a novel way of learning opponent policy embeddings offline; 2) trains offline a single best response (conditional additionally on our opponent policy embedding) instead of a finite set of separate best responses against any opponent; 3) computes online a posterior of the current opponent policy embedding, without making the discrete and ineffective decision which type the current opponent belongs to; and 4) selects online between a real-time greedy policy and a fixed conservative policy via an adversarial bandit algorithm, gaining a theoretically better regret than adhering to either. Experimental studies on popular benchmarks demonstrate GSCU’s superiority over the state-of-the-art methods. The code is available online at \url{https://github.com/YeTianJHU/GSCU}.
Haobo Fu, Hongxiang Yu, Weiming Liu 0004, Jiechao Xiong, Ying Wen 0001, Kai Li 0022, Junliang Xing, Qiang Fu 0016, Wei Yang 0032
ICML8
2022 Towards a Unified Benchmark for Reinforcement Learning in Sparse Reward Environments
Yongxin Kang, Enmin Zhao, Yifan Zang 0001, Kai Li 0022, Junliang Xing
ICONIP (4)4
2022 L2E: Learning to Exploit Your Opponent
abstract
Opponent modeling is essential to exploit sub-optimal opponents in strategic interactions. Most previous works focus on building explicit models to predict the opponents' styles or strategies, which require a large amount of data to train the model and lack adaptability to unknown opponents. In this work, we propose a novel Learning to Exploit (L2E) framework for implicit opponent modeling. L2E acquires the ability to exploit opponents through a few interactions with different opponents during training of a neural network and can quickly adapt to new opponents with unknown styles during testing. To automatically produce challenging and diverse opponents for training, we further present a novel opponent strategy generation algorithm. We evaluate L2E on two poker games and one grid soccer game, which are the commonly used benchmarks for opponent modeling. Comprehensive experimental results indicate that L2E rapidly adapts to diverse styles of unknown opponents.
Kai Li 0022, Hang Xu 0006, Yifan Zang 0001, Bo An 0001, Junliang Xing
IJCNN2
2022 Multiagent Q-learning with Sub-Team Coordination
abstract
In many real-world cooperative multiagent reinforcement learning (MARL) tasks, teams of agents can rehearse together before deployment, but then communication constraints may force individual agents to execute independently when deployed. Centralized training and decentralized execution (CTDE) is increasingly popular in recent years, focusing mainly on this setting. In the value-based MARL branch, credit assignment mechanism is typically used to factorize the team reward into each individual’s reward — individual-global-max (IGM) is a condition on the factorization ensuring that agents’ action choices coincide with team’s optimal joint action. However, current architectures fail to consider local coordination within sub-teams that should be exploited for more effective factorization, leading to faster learning. We propose a novel value factorization framework, called multiagent Q-learning with sub-team coordination (QSCAN), to flexibly represent sub-team coordination while honoring the IGM condition. QSCAN encompasses the full spectrum of sub-team coordination according to sub-team size, ranging from the monotonic value function class to the entire IGM function class, with familiar methods such as QMIX and QPLEX located at the respective extremes of the spectrum. Experimental results show that QSCAN’s performance dominates state-of-the-art methods in matrix games, predator-prey tasks, the Switch challenge in MA-Gym. Additionally, QSCAN achieves comparable performances to those methods in a selection of StarCraft II micro-management tasks.
Wenhan Huang, Kai Li 0022, Kun Shao, Tianze Zhou, Matthew E. Taylor, Jun Luo 0009, Dongge Wang 0001, Hangyu Mao, Jianye Hao, Jun Wang 0012, Xiaotie Deng
NeurIPS2
2022 Multi-Agent Feedback Enabled Neural Networks for Intelligent Communications
abstract
In the intelligent communication field, deep learning (DL) has attracted much attention due to its strong fitting ability and data-driven learning capability. Compared with the typical DL feedforward network structures, an enhancement structure with direct data feedback have been studied and proved to have better performance than the feedfoward networks. However, due to the above simple feedback methods lack sufficient analysis and learning ability on the feedback data, it is inadequate to deal with more complicated nonlinear systems and therefore the performance is limited for further improvement. In this paper, a novel multi-agent feedback enabled neural network (MAFENN) framework is proposed, consisting of three fully cooperative intelligent agents, which make the framework have stronger feedback learning capabilities and more intelligence on feature abstraction, denoising or generation, etc. Furthermore, the MAFENN frame work is theoretically formulated into a three-player Feedback Stackelberg game, and the game is proved to converge to the Feedback Stackelberg equilibrium. The design of MAFENN framework and algorithm are dedicated to enhance the learning capability of the feedfoward DL networks or their variations with the simple data feedback. To verify the MAFENN framework’s feasibility in wireless communications, a multi-agent MAFENN based equalizer (MAFENN-E) is developed for wireless fading channels with inter-symbol interference (ISI). Experimental results show that when the quadrature phase-shift keying (QPSK) modulation scheme is adopted, the SER performance of our proposed method outperforms that of the traditional equalizers by about 2 dB in linear channels. When in nonlinear channels, the SER performance of our proposed method outperforms that of either traditional or DL based equalizers more significantly, which shows the effectiveness and robustness of our proposal in the complex channel environment.
Fanglei Sun, Yang Li 0116, Ying Wen 0001, Jingchen Hu, Jun Wang 0012, Yang Yang 0001, Kai Li 0022
IEEE Trans. Wirel. Commun.7
2021 Exploration via State influence Modeling
abstract
This paper studies the challenging problem of reinforcement learning (RL) in hard exploration tasks with sparse rewards. It focuses on the exploration stage before the agent gets the first positive reward, in which case, traditional RL algorithms with simple exploration strategies often work poorly. Unlike previous methods using some attribute of a single state as the intrinsic reward to encourage exploration, this work leverages the social influence between different states to permit more efficient exploration. It introduces a general intrinsic reward construction method to evaluate the social influence of states dynamically. Three kinds of social influence are introduced for a state: conformity, power, and authority. By measuring the state’s social influence, agents quickly find the focus state during the exploration process. The proposed RL framework with state social influence evaluation works well in hard exploration task. Extensive experimental analyses and comparisons in Grid Maze and many hard exploration Atari 2600 games demonstrate its high exploration efficiency.
Yongxin Kang, Enmin Zhao, Kai Li 0022, Junliang Xing
AAAI3
2021 SFDIC: Spatial Features Distributed Interference Coordination for Massive MIMO Systems
abstract
In 5G massive multiple input multiple output (MIMO) system, the main challenges to mitigate inter-cell interference (ICI) are overhead of information exchange and computational complexity. In this paper, we propose an interference approximation method based on spatial features, which can cover the major channel information by low overhead. And based on this method, a novel distributed low-complexity interference coordination algorithm called SFDIC is proposed, which is based on the idea of leader-follower game to avoid strong ICI. The experimental results show that the proposed interference approximation method is strongly consistent with the traditional channel matrix based interference calculation method on the trend, whose correlation coefficient is 0.9098 and Kullback-Leibler (KL) divergence is close to 0. In low and medium-speed scenarios, the SFDIC increases system and edge throughput by more than 20% and 107% than joint space division multiplexing (JSDM) respectively. And these scenarios reduce the sharing overhead by more than 50% simultaneously. In addition, the new channel predicted module based on Koopman operator is incorporated to improve practical feasibility and system performance loss causing by delay for the first time.
Kai Li 0022, Yang Yang 0001, Liantao Wu, Fanglei Sun, Jinhan Guo
APCC2
2021 MAFENN: Multi-Agent Feedback Enabled Neural Network for Wireless Channel Equalization
abstract
Feedback mechanism has been widely used in wireless communication such as channel equalization and resource allocation. In recent years, deep learning (DL) has made great progress in the field of wireless communication. There is now some work that attempts to introduce plain feedback mechanisms into DL algorithm to solve wireless communication problems. However, the improvement of plain feedback DL methods is limited in complex situations due to those methods lack sufficient learning ability on feedback information. In this paper, we propose a Multi-Agent Feedback Enabled Neural Network (MAFENN) equalizer, which consists of a specific learnable feedback agent and two feed-forward agents. Three fully cooperative intelligent agents help the system improve the ability to remove wireless inter-symbol interference (ISI) in receiving ends. We further formulate it into a three-player Stackelberg Game, which helps us to optimize and train this model more efficiently. To verify the feasibility of our proposed MAFENN system and the Stackelberg Game optimization, we conduct a series of experiments to compare the symbol error rate (SER) performance of the MAFENN equalizer and the other methods which utilizes quadrature phase-shift keying (QPSK) modulation scheme. Our performance outperforms that of the other equalizers at different signal-to-noise ratio (SNR) settings for both linear and nonlinear channels.
Yang Li 0116, Fanglei Sun, Weiqin Zu, Wenbin Song, Ying Wen 0001, Jun Wang 0012, Yang Yang 0001, Kai Li 0022, Liantao Wu
GLOBECOM8
2021 FSST: Frequency-Space Signal Transformation of Massive MIMO Channels
abstract
High overhead of sharing and feedback and high computational complexity are common problems in multi-cell processing. In this paper, a novel framework for bidirectional signal transformation between space and frequency domains of massive MIMO channels is proposed to reduce system processing overhead and complexity. We design new space and frequency features and build the framework by two off-line trained neural networks (NN). Moreover, the uniqueness of spatial features is proved. Average errors of uni- and bi-directional transformation are 7.6% and 7.3%. When applying the framework to inter-cell interference coordination (ICIC), the system and edge throughput are both increased compared to the traditional scheme with low information sharing overhead.
Guoliang Gao, Kai Li 0022, Yang Yang 0001, Liantao Wu, Fanglei Sun
GLOBECOM3
2021 Retrospective Thinking based Multi-Agent System for Wireless Video Transmissions
abstract
Benefiting from the breakthrough development of the fifth generation (5G), beyond 5G (B5G) wireless communication networks and Artificial Intelligence (AI) in recent years, the artificial intelligence of things (AIoT) is a new trend in the future. AIoT devices often have high-quality wireless video transmission requirements. However, the propagating signals at millimeter wave suffer from high propagation loss and sensitivity to blockage, resulting in the received video is vulnerable to be interfered. Due to the ability of Deep Learning (DL) to discover and learn good representations, some DL methods have achieved breakthrough performance in video recovery. However, most of these methods cannot exploit information from the higher to lower level to refine themselves. In this paper, we propose a novel retrospective thinking based multi-agent (ReTMA) system to solve the interference problem experienced on wireless channels. Compared with other plain feedback models, we add a retrospective agent on the feedback loop, which makes the entire system have stronger capabilities to learn good representative features. We further formulate it as a Stackelberg game to analyze the dependency relationship between the agents and facilitate the complex training issue of the multiple agents. To verify the feasibility of ReTMA system, we randomly add masks to simulate the severe interference received by the video frames in wireless transmissions. Experimental results show that the performances of similarity index measure (SSIM), peak signal-to-noise ratio (PSNR) and classification accuracy all achieve significant gains compared with those of other plain feedback models at different mask ratios.
Yang Li 0116, Fanglei Sun, Wenbin Song, Ying Wen 0001, Kai Li 0022, Jun Wang 0012, Yang Yang 0001
ICC5
2021 Learning to Play Hard Exploration Games Using Graph-Guided Self-Navigation
abstract
This work considers the problem of deep reinforcement learning (RL) with long time dependencies and sparse rewards, as are found in many hard exploration games. A graph-based representation is proposed to allow an agent to perform self-navigation for environmental exploration. The graph representation not only effectively models the environment structure, but also efficiently traces the agent state changes and the corresponding actions. By encouraging the agent to earn a new influence-based curiosity reward for new game observations, the whole exploration task is divided into sub-tasks, which are effectively solved using a unified deep RL model. Experimental evaluations on hard exploration Atari Games demonstrate the effectiveness of the proposed method. The source code and learned models will be released to facilitate further studies on this problem.
Enmin Zhao, Renye Yan, Kai Li 0022, Lijuan Li 0002, Junliang Xing
IJCNN3
2020 Potential Driven Reinforcement Learning for Hard Exploration Tasks
abstract
Experience replay plays a crucial role in Reinforcement Learning (RL), enabling the agent to remember and reuse experience from the past. Most previous methods sample experience transitions using simple heuristics like uniformly sampling or prioritizing those good ones. Since humans can learn from both good and bad experiences, more sophisticated experience replay algorithms need to be developed. Inspired by the potential energy in physics, this work introduces the artificial potential field into experience replay and develops Potentialized Experience Replay (PotER) as a new and effective sampling algorithm for RL in hard exploration tasks with sparse rewards. PotER defines a potential energy function for each state in experience replay and helps the agent to learn from both good and bad experiences using intrinsic state supervision. PotER can be combined with different RL algorithms as well as the self-imitation learning algorithm. Experimental analyses and comparisons on multiple challenging hard exploration environments have verified its effectiveness and efficiency.
Enmin Zhao, Shihong Deng, Yifan Zang 0001, Yongxin Kang, Kai Li 0022, Junliang Xing
IJCAI5
2019 Parallel Scheduling of Multiple Tasks in Heterogeneous Fog Networks
abstract
Fog computing has been promoted to support delay-sensitive applications in future Internet of Things (IoT) and wireless networks. For a general heterogeneous fog network consisting of many dispersive Fog Nodes (FNs) with diverse resources and capabilities, some of them have delay-sensitive tasks to process, i.e., Task Nodes (TNs), while some have spare resources to help their neighboring TNs to process tasks, i.e., Helper Nodes (HNs). How to effectively map multiple tasks or TNs into multiple HNs to minimize every task's service delay in a distributed manner is a fundamental challenge, which is key to reap the full benefits of fog computing. The problem becomes more challenging when tasks can be divided into multiple subtasks to further reduce the service delay via distributed computing. To tackle this challenge, in this paper, a generalized nash equilibrium (NE) game called Parallel Scheduling of Multiple Tasks (PSMT) is formulated and studied. The structure properties of the problem are deduced and thus the existence of NE is proven by the fixed point theorem. Further, the corresponding distributed task scheduling algorithm/mechanism is developed via Gauss-Seidel-type method. Simulation results show that the proposed PSMT algorithm can converge in a fast way and offer much better performance in system average delay and number of beneficial TNs, comparing to the Paired Offloading of Multiple Tasks (POMT) solution to the counterpart problem not supporting distributed computing.
Zening Liu, Kunlun Wang 0001, Kai Li 0022, Ming-Tuo Zhou, Yang Yang 0001
APCC3
2019 Spatial Temporal Attentional Glimpse for Human Activity Classification in Video
abstract
Recently, the Convolutional Networks (ConvNet) has become the dominated approach to the human activity classification problem. We investigate current standard ConvNet architectures and pinpoint one of their main limitations: the spatial-temporal dependency is simply captured by global pooling operation, which may not well capture the complex long term spatial-temporal relationships in videos. For this work, we propose a Spatial Temporal Attentional Glimpse (STAG) module to overcome this shortcoming. Specifically, the input to this STAG module is a 3D tensor which is first processed by a spatial-temporal attention block. Spatial Temporal Glimpse block decomposes the resulting tensor into two low dimensional tensors and then fuses their operation results. The proposed STAG module is pluggable, easy to learn, and effective in computation. We conduct extended ablation studies to show that our model incorporated with the STAG block substantially improves the performance over the state-of-the-art. All the experimental results, the trained models, and the complete source codes will be released to facilitate further studies on this problem.
Jiangtao Kong, Rongchao Xu, Junliang Xing, Kai Li 0022
ICIP4
2018 Deep Cost-Sensitive and Order-Preserving Feature Learning for Cross-Population Age Estimation
abstract
Facial age estimation from a face image is an important yet very challenging task in computer vision, since humans with different races and/or genders, exhibit quite different patterns in their facial aging processes. To deal with the influence of race and gender, previous methods perform age estimation within each population separately. In practice, however, it is often very difficult to collect and label sufficient data for each population. Therefore, it would be helpful to exploit an existing large labeled dataset of one (source) population to improve the age estimation performance on another (target) population with only a small labeled dataset available. In this work, we propose a Deep Cross-Population (DCP) age estimation model to achieve this goal. In particular, our DCP model develops a two-stage training strategy. First, a novel cost-sensitive multitask loss function is designed to learn transferable aging features by training on the source population. Second, a novel order-preserving pair-wise loss function is designed to align the aging features of the two populations. By doing so, our DCP model can transfer the knowledge encoded in the source population to the target population. Extensive experiments on the two of the largest benchmark datasets show that our DCP model outperforms several strong baseline methods and many state-of-the-art methods.
Kai Li 0022, Junliang Xing, Chi Su, Weiming Hu 0004, Stephen J. Maybank
CVPR1
2017 Weakly supervised multiscale-inception learning for web-scale face recognition
abstract
Supervised deep learning models like convolutional neural network (CNN) have shown very promising results for the face recognition problem, which often require a huge number of labeled face images. Since manually labeling a large training set is a very difficult and time-consuming task, it is very beneficial if the deep model can be trained from face samples with only weak annotations. In this paper, we propose a general framework to train a deep CNN model with weakly labeled facial images that are available on the Internet. Specifically, we first design a deep Multiscale-Inception CNN (MICNN) architecture to exploit the multi-scale information for face recognition. Then, we train an initial MICNN model with only a limited number of labeled samples. After that, we propose a dual-level sample selection strategy to further fine-tune the MICNN model with the weakly labeled samples from both the sample level and class level, which aims to skip outliers and select more samples from confusing class pairs during training. Extensive experimental results on the LFW and YTF benchmarks demonstrate the effectiveness of the proposed method.
Junliang Xing, Youji Feng, Xiaohu Shao, Kai Li 0022
ICIP6
2017 A Novel Network Optimization Method for Cooperative Massive MIMO Systems
abstract
Capacity-coverage tradeoff balancing is a classical but critical problem for practical wireless multi-cell network optimization. Increasing the average capacity is often at the expense of decreasing the cell coverage and vice versa. It becomes more complex in massive multiple-input/multiple-output (MIMO) networks to balance the above two performance indicators since capacity is much highly improved while the inter-cell interference is also increased greatly under massive antenna scenarios. In this work, a novel system-level optimization parameter is proposed, namely minimum user signal to interference plus noise ratio (user-SINR) threshold, to solve the above balancing problem. By utilizing this new parameter, the corresponding user scheduling and inter-cell interference coordination scheme are further provided for cooperative massive MIMO networks. Performance analysis and numerical results show that the proposed minimum user-SINR threshold is a very effective optimization parameter to achieve the capacity-coverage tradeoff performance with lower optimization complexity, combined with the proposed scheduling and interference coordination schemes.
Kai Li 0022, Yang Yang 0001, Yu Chen 0006, Xiumei Yang, Huiyue Yi
VTC Spring1
2017 D2C: Deep cumulatively and comparatively learning for human age estimation
Kai Li 0022, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank
Pattern Recognit.1
2017 Diagnosing deep learning models for high accuracy age estimation from a single image
Junliang Xing, Kai Li 0022, Weiming Hu 0004, Chunfeng Yuan, Haibin Ling
Pattern Recognit.2
2016 Bootstrapping deep feature hierarchy for pornographic image recognition
abstract
Automatically recognizing pornographic images from the Web is a vital step to purify Internet environment. Inspired by the rapid developments of deep learning models, we present a deep architecture of convolutional neural network (CNN) for high accuracy pornographic image recognition. The proposed architecture is built upon existing CNNs which accepts input images of different sizes and incorporates features from different hierarchy to perform prediction. To effectively train the model, we propose a two-stage training strategy to learn the model parameters from scratch and end-to-end. During the training procedure, we also employ a hard negative sampling strategy to further reduce the false positive rate of the model. Experimental results on a large dataset demonstrate good performance of the proposed model and the effectiveness of our training strategies, with a considerable improvement over some traditional methods using hand-crafted features and deep learning method using mainstream CNN architecture.
Kai Li 0022, Junliang Xing, Bing Li 0001, Weiming Hu 0004
ICIP1
2015 Predicting Image Memorability by Multi-view Adaptive Regression
abstract
The images we encounter throughout our lives make different impressions on us: Some are remembered at first glance, while others are forgotten. This phenomenon is caused by the intrinsic memorability of images revealed by recent studies [5,6]. In this paper, we address the issue of automatically estimating the memorability of images by proposing a novel multi-view adaptive regression (MAR) model. The MAR model provides an effective mapping of visual features to memorability scores by taking advantage of robust feature selection and multiple feature integration. It consists of three major components: an adaptive loss function, an adaptive regularization and a multi-view modeling strategy. Moreover, we design an alternating direction method (ADM) optimization algorithm to solve the proposed objective function. Experimental results on the MIT benchmark dataset show the superiority of the proposed model compared with existing image memorability prediction methods.
Houwen Peng, Kai Li 0022, Bing Li 0001, Haibin Ling, Weihua Xiong, Weiming Hu 0004
ACM Multimedia2