EDBT 2026 Demo / reviewers in the wild / expert
Junge Zhang
dblp:06/9989
· DBLP profile ↗
93ranked-venue papers
6as first author
52since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 5 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 52 · 4 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsabstractBackdoor attacks embed malicious behaviors into Large Language Models (LLMs), enabling adversaries to trigger harmful outputs or bypass safety controls. However, the persistence of the implanted backdoors under user-driven post-deployment continual fine-tuning has been rarely examined. Most prior works evaluate the effectiveness and generalization of implanted backdoors only at releasing and empirical evidence shows that naively injected backdoor persistence degrades after updates. In this work, we study whether and how implanted backdoors persist through a multi‑stage post-deployment fine‑tuning. We propose P‑Trojan, a trigger‑based attack algorithm that explicitly optimizes for backdoor persistence across repeated updates. By aligning poisoned gradients with those of clean tasks on token embeddings, the implanted backdoor mapping is less likely to be suppressed or forgotten during subsequent updates. Theoretical analysis shows the feasibility of such persistent backdoor attacks after continual fine-tuning. And experiments conducted on the Qwen2.5 and LLaMA3 families of LLMs, as well as diverse task sequences, demonstrate that P‑Trojan achieves over \textbf{99\%} persistence while preserving clean‑task accuracy. Our findings highlight the need for persistence-aware evaluation and stronger defenses in realistic model adaptation pipelines. Jianbin Jiao, Junge Zhang |
AAAI | 4 |
| 2026 | Calibration-Aware Policy Optimization for Reasoning LLMsabstractGroup Relative Policy Optimization (GRPO) enhances LLM reasoning but often induces overconfidence, where incorrect responses yield lower perplexity than correct ones, degrading relative calibration as described by the Area Under the Curve (AUC).Existing approaches either yield limited improvements in calibration or sacrifice gains in reasoning accuracy.We first prove that this degradation in GRPO-style algorithms stems from their uncertainty-agnostic advantage estimation, which inevitably misaligns optimization gradients with calibration.This leads to improved accuracy at the expense of degraded calibration.We then propose Calibration-Aware Policy Optimization (CAPO).It adopts a logistic AUC surrogate loss that is theoretically consistent and admits regret bound, enabling uncertainty-aware advantage estimation.By further incorporating a noise masking mechanism, CAPO achieves stable learning dynamics that jointly optimize calibration and accuracy.Experiments on multiple mathematical reasoning benchmarks show that CAPO-1.5Bsignificantly improves calibration by up to 15% while achieving accuracy comparable to or better than GRPO, and further boosts accuracy on downstream inference-time scaling tasks by up to 5%.Moreover, when allowed to abstain under low-confidence conditions, CAPO achieves a Pareto-optimal precision-coverage trade-off, highlighting its practical value for hallucination mitigation. Xingzhou Lou, Meiqi Wu, Zhengqi Wen, Junge Zhang |
ACL (1) | 5 |
| 2026 | RainbowArena: A multi-agent toolkit for reinforcement learning and large language models in tabletop games
Yingzhuo Liu, Shuodi Liu, Hongsong Tang, Yubing Ma, Zikang Li, Junge Zhang, Liuyu Xiang, Zhaofeng He 0001 |
Knowl. Based Syst. | 6 |
| 2026 | Subgame Pruning: Efficiently Solving Two-Player Imperfect Information Games by Accelerated Public Tree Traversal in CFRabstractCounterfactual Regret Minimization is the state-of-the-art algorithm for solving imperfect information games, yet it struggles against scalability. While existing pruning techniques primarily focus on pruning unreachable branches, they fail to account for subgames in which strategies have temporarily converged. Continuing to update strategies within these subgames leads to unnecessary computational overhead. To address this, we propose Subgame Pruning, a novel regret-based pruning paradigm that accelerates traversal of the public tree by dynamically pruning subgames where strategies remain stable across iterations. To ensure safe and efficient pruning, we introduce two key mechanisms: Pruning Constraint Checking, which verifies whether a subgame satisfies pruning criteria, and Regret Matching Compensation, which defers regret matching and average strategy updates until the pruned subgames are revisited. Our method significantly reduces computational overhead while preserving the theoretical convergence guarantees of CFR. In particular, the exploitability of the average strategy profile remains bounded by$\mathcal {O}(1/\sqrt{T})$. Experimental results across five benchmark games demonstrate substantial reductions in the number of traversed nodes, with exploitability comparable to that of standard CFR. Shenkai Zhang, Peipei Yang, Zekeng Zeng, Junge Zhang |
IEEE Trans. Games | 4 |
| 2025 | Sequential Preference Optimization: Multi-Dimensional Preference Alignment with Implicit Reward ModelingabstractHuman preference alignment is critical in building powerful and reliable large language models (LLMs). However, current methods either ignore the multi-dimensionality of human preferences (e.g. helpfulness and harmlessness) or struggle with the complexity of managing multiple reward models. To address these issues, we propose Sequential Preference Optimization (SPO), a method that sequentially fine-tunes LLMs to align with multiple dimensions of human preferences. SPO avoids explicit reward modeling, directly optimizing the models to align with nuanced human preferences. We theoretically derive closed-form optimal SPO policy and loss function. Gradient analysis is conducted to show how SPO manages to fine-tune the LLMs while maintaining alignment on previously optimized dimensions. Empirical results on LLMs of different size and multiple evaluation datasets demonstrate that SPO successfully aligns LLMs across multiple dimensions of human preferences and significantly outperforms the baselines. Xingzhou Lou, Junge Zhang, Lifeng Liu, Kaiqi Huang |
AAAI | 2 |
| 2025 | EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement LearningabstractLarge Language Models (LLMs) have shown impressive reasoning capabilities in well-defined problems with clear solutions, such as mathematics and coding. However, they still struggle with complex real-world scenarios like business negotiations, which require strategic reasoning—an ability to navigate dynamic environments and align long-term goals amidst uncertainty.Existing methods for strategic reasoning face challenges in adaptability, scalability, and transferring strategies to new contexts.To address these issues, we propose explicit policy optimization (EPO) for strategic reasoning, featuring an LLM that provides strategies in open-ended action space and can be plugged into arbitrary LLM agents to motivate goal-directed behavior.To improve adaptability and policy transferability, we train the strategic reasoning model via multi-turn reinforcement learning (RL), utilizing process rewards and iterative self-play.Experiments across social and physical domains demonstrate EPO’s ability of long-term goal alignment through enhanced strategic reasoning, achieving state-of-the-art performance on social dialogue and web navigation tasks. Our findings reveal various collaborative reasoning mechanisms emergent in EPO and its effectiveness in generating novel strategies, underscoring its potential for strategic reasoning in real-world applications. Code and data are available at https://github.com/lxqpku/EPO. Yongbin Li 0001, Yuchuan Wu, Aobo Kong, Fei Huang 0002, Jianbin Jiao, Junge Zhang |
ACL (1) | 9 |
| 2025 | IDEA-Bench: How Far are Generative Models from Professional Designing?abstractRecent advancements in image generation models enable the creation of high-quality images and targeted modifications based on textual instructions. Some models even support multimodal complex guidance and demonstrate robust task generalization capabilities. However, they still fall short of meeting the nuanced, professional demands of designers. To bridge this gap, we introduce IDEA-Bench, a comprehensive benchmark designed to advance image generation models toward applications with robust task generalization. IDEA-Bench comprises 100 professional image generation tasks and 275 specific cases, categorized into five major types based on the current capabilities of existing models. Furthermore, we provide a representative subset of 18 tasks with enhanced evaluation criteria to facilitate more nuanced and reliable evaluations using Multimodal Large Language Models (MLLMs). By assessing models’ ability to comprehend and execute novel, complex tasks, IDEA-Bench paves the way toward the development of generative models with autonomous and versatile visual generation capabilities. Lianghua Huang, Jingwu Fang, Huanzhang Dou, Wei Wang 0354, Zhi-Fan Wu, Yupeng Shi, Junge Zhang, Xin Zhao 0012, Yu Liu 0063 |
CVPR | 8 |
| 2025 | Uncertainty-Aware Diffusion-Guided Refinement of 3D ScenesabstractReconstructing 3D scenes from a single image is a fundamentally ill-posed task due to the severely under-constrained nature of the problem. Consequently, when the scene is rendered from novel camera views, existing single image to 3D reconstruction methods render incoherent and blurry views. This problem is exacerbated when the unseen regions are far away from the input camera. In this work, we address these inherent limitations in existing single image-to-3D scene feedforward networks. To alleviate the poor performance due to insufficient information beyond the input image's view, we leverage a strong generative prior in the form of a pre-trained latent video diffusion model, for iterative refinement of a coarse scene represented by optimizable Gaussian parameters. To ensure that the style and texture of the generated images align with that of the input image, we incorporate on-the-fly Fourier-style transfer between the generated images and the input image. Additionally, we design a semantic uncertainty quantification module that calculates the per-pixel entropy and yields uncertainty maps used to guide the refinement process from the most confident pixels while discarding the remaining highly uncertain ones. We conduct extensive experiments on real-world scene datasets, including in-domain RealEstate-10K and out-of-domain KITTI-v2, showing that our approach can provide more realistic and high-fidelity novel view synthesis results compared to existing state-of-the-art methods. Sarosij Bose, Arindam Dutta, Sayak Nag, Junge Zhang, Konstantinos Karydis, Amit K. Roy-Chowdhury |
ICCV | 4 |
| 2025 | RainbowArena: A Multi-Agent Toolkit for Reinforcement Learning and Large Language Models in Competitive Tabletop Games
Yingzhuo Liu, Shuodi Liu, Hongsong Tang, Yubing Ma, Zikang Li, Junge Zhang, Liuyu Xiang, Zhaofeng He 0001 |
AAMAS | 6 |
| 2025 | Distribution Optimization Under Gaussian Hypothesis for Domain Adaptive Semantic SegmentationabstractDomain adaptive semantic segmentation aims to transfer a model, proficient in dense image classification, from a source domain to a target domain. While various transfer methods have been explored in previous studies, we argue that the modeling of categories within the model significantly affects its transferability. Building on the Gaussian Hypothesis, which posits that each category in the feature space adheres to a multidimensional Gaussian distribution, we propose a Class-Aware Variational Inference (CAVI) training method. This approach normalizes features of different categories into distinct multidimensional Gaussian distributions. To further learn domain-independent feature distributions, we optimize the feature space using a Gaussian-based alignment strategy and incorporate Gaussian-based contrastive learning. Experimental results demonstrate that our method achieves state-of-the-art performance on the GTAV → Cityscapes and Synthia → Cityscapes benchmarks. Xin Zhao 0012, Junyan Wang 0001, Lijun Cao, Junge Zhang |
WACV | 6 |
| 2025 | Rethinking Generalizability and Discriminability of Self-Supervised Learning from Evolutionary Game Theory Perspective
Jiangmeng Li, Zehua Zang, Qirui Ji, Chuxiong Sun, Wenwen Qiang, Junge Zhang, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001 |
Int. J. Comput. Vis. | 6 |
| 2025 | Generalizable agent modeling for agent collaboration-competition adaptation with multi-retrieval and dynamic generation
Yonggang Jin, Youpeng Zhao 0001, Zipeng Dai, Jian Zhao 0018, Liuyu Xiang, Junge Zhang, Zhaofeng He 0001 |
Neurocomputing | 8 |
| 2025 | S-NeRF++: Autonomous Driving Simulation via Neural Reconstruction and GenerationabstractAutonomous driving simulation system plays a crucial role in enhancing self-driving data and simulating complex and rare traffic scenarios, ensuring navigation safety. However, traditional simulation systems, which often heavily rely on manual modeling and 2D image editing, struggled with scaling to extensive scenes and generating realistic simulation data. In this study, we present S-NeRF++, an innovative autonomous driving simulation system based on neural reconstruction. Trained on widely-used self-driving datasets, such as nuScenes and Waymo, S-NeRF++ can generate a large number of realistic street scenes and foreground objects with high rendering quality as well as offering considerable flexibility in manipulation and simulation. Specifically, S-NeRF++ is an enhanced neural radiance field for synthesizing large-scale scenes and moving vehicles, with improved scene parameterization and camera pose learning. The system effectively utilizes noisy and sparse LiDAR data to refine training and address depth outliers, ensuring high-quality reconstruction and novel-view rendering. It also provides a diverse foreground asset bank by reconstructing and generating different foreground vehicles to support comprehensive scenario creation. Moreover, we have developed an advanced foreground-background fusion pipeline that skillfully integrates illumination and shadow effects, further enhancing the realism of our simulations. With the high-quality simulated data provided by our S-NeRF++, we found the perception methods enjoy performance boosts on several autonomous driving downstream tasks, further demonstrating our proposed simulator's effectiveness. Yurui Chen, Junge Zhang, Ziyang Xie, Wenye Li 0002, Feihu Zhang, Li Zhang 0040 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Learning Individual Potential-Based Rewards in Multiagent Reinforcement LearningabstractA great challenge for applying multiagent reinforcement learning (MARL) in the field of game artificial intelligence (AI) is to enable agents to learn diversified policies to handle different game-specific problems, while receiving only a shared team reward. At present, a common approach is reward shaping, which focuses on designing rewards for agents to guide cooperation. However, most of the existing methods require prior knowledge on the environment for reward design or alter the optimal policies after imposing extra rewards. Besides, previous MARL methods that rely on manually designed rewards can hardly generalize across different game environments. To this end, we propose a new MARL method that learns individual potential-based rewards for agents. Specifically, we learn a parameterized potential function for each agent to generate individual rewards in the discounted temporal difference form. The whole update procedure is modeled as the bilevel optimization problem, where the lower level is to optimize policies with potential-based rewards, and the upper level is to optimize parameterized potential functions toward maximizing the environment return. We theoretically prove that the individual potential-based rewards can guarantee policy invariance for agents, so that the optimization objective is consistent with the original MARL problem. We evaluate our method with a number of existing state-of-the-art MARL methods on predator–prey andStarCraft IIgame environments. Empirical results show that our proposed method significantly outperforms baseline methods and achieves better game AI that enjoys high performance and generalization. Pei Xu 0003, Junge Zhang |
IEEE Trans. Games | 3 |
| 2025 | Relation-Aware Learning for Multitask Multiagent Cooperative GamesabstractCollaboration among multiple tasks is advantageous for enhancing learning efficiency in multiagent reinforcement learning. To guide agents in cooperating with different teammates in multiple tasks, contemporary approaches encourage agents to exploit common cooperative patterns or identify the learning priorities of multiple tasks. Despite the progress made by these methods, they all assume that all cooperative tasks to be learned are related and desire similar agent policies. This is rarely the case in multiagent cooperation, where minor changes in team composition can lead to significant variations in cooperation, resulting in distinct cooperative strategies compete for limited learning resources. In this article, to tackle the challenge posed by multitask learning in potentially competing cooperative tasks, we propose a novel framework called relation-aware learning (RAL). RAL incorporates a relation awareness module in both task representation and task optimization, aiding in reasoning about task relationships and mitigating negative transfers among dissimilar tasks. To assess the performance of RAL, we conduct a comparative analysis with baseline methods in a multitaskStarCraftenvironment. The results demonstrate the superiority of RAL in multitask cooperative scenarios, particularly in scenarios involving multiple conflicting tasks. Yang Yu 0056, Likun Yang, Zhourui Guo, Qiyue Yin, Junge Zhang, Kaiqi Huang |
IEEE Trans. Games | 6 |
| 2024 | BadRL: Sparse Targeted Backdoor Attack against Reinforcement LearningabstractBackdoor attacks in reinforcement learning (RL) have previously employed intense attack strategies to ensure attack success. However, these methods suffer from high attack costs and increased detectability. In this work, we propose a novel approach, BadRL, which focuses on conducting highly sparse backdoor poisoning efforts during training and testing while maintaining successful attacks. Our algorithm, BadRL, strategically chooses state observations with high attack values to inject triggers during training and testing, thereby reducing the chances of detection. In contrast to the previous methods that utilize sample-agnostic trigger patterns, BadRL dynamically generates distinct trigger patterns based on targeted state observations, thereby enhancing its effectiveness. Theoretical analysis shows that the targeted backdoor attack is always viable and remains stealthy under specific assumptions. Empirical results on various classic RL tasks illustrate that BadRL can substantially degrade the performance of a victim agent with minimal poisoning efforts (0.003% of total training steps) during training and infrequent attacks during testing. Code is available at: https://github.com/7777777cc/code. Yuzhe Ma, Jianbin Jiao, Junge Zhang |
AAAI | 5 |
| 2024 | TAPE: Leveraging Agent Topology for Cooperative Multi-Agent Policy GradientabstractMulti-Agent Policy Gradient (MAPG) has made significant progress in recent years. However, centralized critics in state-of-the-art MAPG methods still face the centralized-decentralized mismatch (CDM) issue, which means sub-optimal actions by some agents will affect other agent's policy learning. While using individual critics for policy updates can avoid this issue, they severely limit cooperation among agents. To address this issue, we propose an agent topology framework, which decides whether other agents should be considered in policy gradient and achieves compromise between facilitating cooperation and alleviating the CDM issue. The agent topology allows agents to use coalition utility as learning objective instead of global utility by centralized critics or local utility by individual critics. To constitute the agent topology, various models are studied. We propose Topology-based multi-Agent Policy gradiEnt (TAPE) for both stochastic and deterministic MAPG methods. We prove the policy improvement theorem for stochastic TAPE and give a theoretical explanation for the improved cooperation among agents. Experiment results on several benchmarks show the agent topology is able to facilitate agent cooperation and alleviate CDM issue respectively to improve performance of TAPE. Finally, multiple ablation studies and a heuristic graph search algorithm are devised to show the efficacy of the agent topology. Xingzhou Lou, Junge Zhang, Timothy J. Norman, Kaiqi Huang, Yali Du 0001 |
AAAI | 2 |
| 2024 | ProAgent: Building Proactive Cooperative Agents with Large Language ModelsabstractBuilding agents with adaptive behavior in cooperative tasks stands as a paramount goal in the realm of multi-agent systems. Current approaches to developing cooperative agents rely primarily on learning-based methods, whose policy generalization depends heavily on the diversity of teammates they interact with during the training phase. Such reliance, however, constrains the agents' capacity for strategic adaptation when cooperating with unfamiliar teammates, which becomes a significant challenge in zero-shot coordination scenarios. To address this challenge, we propose ProAgent, a novel framework that harnesses large language models (LLMs) to create proactive agents capable of dynamically adapting their behavior to enhance cooperation with teammates. ProAgent can analyze the present state, and infer the intentions of teammates from observations. It then updates its beliefs in alignment with the teammates' subsequent actual behaviors. Moreover, ProAgent exhibits a high degree of modularity and interpretability, making it easily integrated into various of coordination scenarios. Experimental evaluations conducted within the Overcooked-AI environment unveil the remarkable performance superiority of ProAgent, outperforming five methods based on self-play and population-based training when cooperating with AI agents. Furthermore, in partnered with human proxy models, its performance exhibits an average improvement exceeding 10% compared to the current state-of-the-art method. For more information about our project, please visit https://pku-proagent.github.io. Ceyao Zhang, Kaijie Yang, Siyi Hu 0001, Guanghe Li, Yihang Sun, Zhaowei Zhang 0001, Anji Liu, Song-Chun Zhu, Xiaojun Chang, Junge Zhang, Feng Yin 0001, Yitao Liang, Yaodong Yang 0001 |
AAAI | 12 |
| 2024 | NeRF-LiDAR: Generating Realistic LiDAR Point Clouds with Neural Radiance FieldsabstractLabelling LiDAR point clouds for training autonomous driving is extremely expensive and difficult. LiDAR simulation aims at generating realistic LiDAR data with labels for training and verifying self-driving algorithms more efficiently. Recently, Neural Radiance Fields (NeRF) have been proposed for novel view synthesis using implicit reconstruction of 3D scenes. Inspired by this, we present NeRF-LIDAR, a novel LiDAR simulation method that leverages real-world information to generate realistic LIDAR point clouds. Different from existing LiDAR simulators, we use real images and point cloud data collected by self-driving cars to learn the 3D scene representation, point cloud generation and label rendering. We verify the effectiveness of our NeRF-LiDAR by training different 3D segmentation models on the generated LiDAR point clouds. It reveals that the trained models are able to achieve similar accuracy when compared with the same model trained on the real LiDAR data. Besides, the generated data is capable of boosting the accuracy through pre-training which helps reduce the requirements of the real labeled data. Code is available at https://github.com/fudan-zvg/NeRF-LiDAR Junge Zhang, Feihu Zhang, Shaochen Kuang |
AAAI | 1 |
| 2024 | Information Bottleneck Based Data Correction in Continual Learning
Mingyi Zhang 0004, Junge Zhang, Kaiqi Huang |
ECCV (87) | 3 |
| 2024 | Task-Wise Prompt Query Function for Rehearsal-Free Continual LearningabstractContinual learning (CL) aims to enable a model to retain knowledge of old tasks while learning new ones. One effective approach to CL is based on data rehearsal method. However, this approach increases the cost of storing data and cannot be used when data from old tasks is unavailable for some reason. Recently, with the emergence of large- scale pre-trained transformer models, prompt-based methods have become an alternative to data rehearsal. These methods rely on a query mechanism to generate prompts and have demonstrated resistance to forgetting in CL scenarios without rehearsal. However, these methods generate prompts in a task-wise way while queries for samples in an instance-wise way, and usually directly use pre-trained models as the encoding function for generating queries. This may lead to data retrieval errors and failure to match the correct prompts. In contrast, we propose building a new task-wise prompt query function that can continuously learn as the task progresses, thereby avoiding the issue of pre-trained models being unable to correctly match appropriate sample-prompt pairs. Our approach improves the effectiveness of the current state-of- the-art methods and has been verified on a series of datasets through our experimental results. Mingyi Zhang 0004, Junge Zhang, Kaiqi Huang |
ICASSP | 3 |
| 2024 | Position: Foundation Agents as the Paradigm Shift for Decision MakingabstractDecision making demands intricate interplay between perception, memory, and reasoning to discern optimal policies. Conventional approaches to decision making face challenges related to low sample efficiency and poor generalization. In contrast, foundation models in language and vision have showcased rapid adaptation to diverse new tasks. Therefore, we advocate for the construction of foundation agents as a transformative shift in the learning paradigm of agents. This proposal is underpinned by the formulation of foundation agents with their fundamental characteristics and challenges motivated by the success of large language models (LLMs). Moreover, we specify the roadmap of foundation agents from large interactive data collection or generation, to self-supervised pretraining and adaptation, and knowledge and value alignment with LLMs. Lastly, we pinpoint critical research questions derived from the formulation and delineate trends for foundation agents supported by real-world use cases, addressing both technical and theoretical aspects to propel the field towards a more comprehensive and impactful future. Xingzhou Lou, Jianbin Jiao, Junge Zhang |
ICML | 4 |
| 2024 | Computing Approximate Nash Equilibrium in Two-Team Zero-Sum Games by NashConv Descent
Zekeng Zeng, Youzhi Zhang 0001, Peipei Yang, Junge Zhang |
ICONIP (4) | 5 |
| 2024 | ADMN: Agent-Driven Modular Network for Dynamic Parameter Sharing in Cooperative Multi-Agent Reinforcement Learning
Yang Yu 0056, Qiyue Yin, Junge Zhang, Pei Xu 0003, Kaiqi Huang |
IJCAI | 3 |
| 2024 | Population-Based Diverse Exploration for Sparse-Reward Multi-Agent Tasks
Pei Xu 0003, Junge Zhang, Kaiqi Huang |
IJCAI | 2 |
| 2024 | Cross-modal misalignment-robust feature fusion for crowd counting
Weihang Kong, Zepeng Yu, He Li 0053, Junge Zhang |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Softmax-Free Linear Transformers
Junge Zhang, Xiatian Zhu, Jianfeng Feng, Tao Xiang 0002, Li Zhang 0040 |
Int. J. Comput. Vis. | 2 |
| 2024 | Leveraging Joint-Action Embedding in Multiagent Reinforcement Learning for Cooperative GamesabstractState-of-the-art multi-agent policy gradient (MAPG) methods have demonstrated convincing capability in many cooperative games. However, the exponentially growing joint-action space severely challenges the critic's value evaluation and hinders performance of MAPG methods. To address this issue, we augment Central-Q policy gradient with a joint-action embedding function and propose Mutual-information Maximization MAPG (M3APG). The joint-action embedding function makes joint-actions contain information of state transitions, which will improve the critic's generalization over the joint-action space by allowing it to infer joint-actions' outcomes. We theoretically prove that with a fixed joint-action embedding function, the convergence of M3APG is guaranteed. Experiment results on the StarCraft Multi-Agent Challenge (SMAC) demonstrate that M3APG gives evaluation results with better accuracy and outperform other MAPG basic models across various maps of multiple difficulty levels. We empirically show that our joint-action embedding model can be extended to value-based multi-agent reinforcement learning methods and state-of-the-art MAPG methods. Finally, we run ablation study to show that the usage of mutual information in our method is necessary and effective. Xingzhou Lou, Junge Zhang, Yali Du 0001, Chao Yu 0004, Zhaofeng He 0001, Kaiqi Huang |
IEEE Trans. Games | 2 |
| 2024 | Contrastive Correlation Preserving Replay for Online Continual LearningabstractOnline Continual Learning (OCL), as a core step towards achieving human-level intelligence, aims to incrementally learn and accumulate novel concepts from streaming data that can be seen only once, while alleviating catastrophic forgetting on previously acquired knowledge. Under this mode, the model needs to learn new classes or tasks in an online manner, and the data distribution may change over time. Moreover, task boundaries and identities are not available during training and evaluation. To balance the stability and plasticity of networks, in this work, we propose a replay-based framework for OCL, named Contrastive Correlation Preserving Replay (CCPR), which focuses on not only instances but also correlations between multiple instances. Specifically, besides the previous raw samples, the corresponding representations are stored in the memory and used to construct correlations for the past and the current model. To better capture correlation and higher-order dependencies, we maximize the low bound of mutual information between the past correlation and the current correlation by leveraging contrastive objectives. Furthermore, to improve the performance, we propose a new memory update strategy, which simultaneously encourages the balance and diversity of samples within the memory. With limited memory slots, it allows less redundant and more representative samples for later replay. We conduct extensive evaluations on several popular CL datasets, and experiments show that our method consistently outperforms the state-of-the-art methods and can effectively consolidate knowledge to alleviate forgetting. Mingyi Zhang 0004, Mantian Li, Fusheng Zha, Junge Zhang, Lining Sun, Kaiqi Huang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Subspace-Aware Exploration for Sparse-Reward Multi-Agent TasksabstractExploration under sparse rewards is a key challenge for multi-agent reinforcement learning problems. One possible solution to this issue is to exploit inherent task structures for an acceleration of exploration. In this paper, we present a novel exploration approach, which encodes a special structural prior on the reward function into exploration, for sparse-reward multi-agent tasks. Specifically, a novel entropic exploration objective which encodes the structural prior is proposed to accelerate the discovery of rewards. By maximizing the lower bound of this objective, we then propose an algorithm with moderate computational cost, which can be applied to practical tasks. Under the sparse-reward setting, we show that the proposed algorithm significantly outperforms the state-of-the-art algorithms in the multiple-particle environment, the Google Research Football and StarCraft II micromanagement tasks. To the best of our knowledge, on some hard tasks (such as 27m_vs_30m}) which have relatively larger number of agents and need non-trivial strategies to defeat enemies, our method is the first to learn winning strategies under the sparse-reward setting. Pei Xu 0003, Junge Zhang, Qiyue Yin, Chao Yu 0004, Yaodong Yang 0001, Kaiqi Huang |
AAAI | 2 |
| 2023 | Safe-NORA: Safe Reinforcement Learning-based Mobile Network Resource Allocation for Diverse User DemandsabstractAs mobile communication technologies advance, mobile networks become increasingly complex, and user requirements become increasingly diverse. To satisfy the diverse demands of users while improving the overall performance of the network system, the limited wireless network resources should be efficiently and dynamically allocated to them based on the magnitude of their demands and their relative location to the base stations. We separated the problem into four constrained subproblems, which we then solved using a safe reinforcement learning method. In addition, we design a reward mechanism to encourage agent cooperation in distributed training environments. We test our methodology in a simulated scenario with thousands of users and hundreds of base stations. According to experimental findings, our method guarantees that over 95% of user demands are satisfied while also maximizing the overall system throughput. Wenzhen Huang, Tong Li 0013, Yuting Cao, Zhe Lyu, Yanping Liang, Depeng Jin, Junge Zhang, Yong Li 0008 |
CIKM | 8 |
| 2023 | Improved Training Of Mixture-Of-Experts Language GANsabstractDespite the dramatic success in image generation, Generative Adversarial Networks (GANs) still face great challenges in text generation. The difficulty in generator training arises from the limited representation capacity and uninformative learning signals obtained from the discriminator. In this work, we (1) first empirically show that the multi-generator approach is able to enhance the representation capacity of the generator for sequence GANs and (2) harness the Feature Statistics Alignment (FSA) paradigm to render fine-grained learning signals to advance the generator training. Specifically, FSA forces the mean statistics of the distribution of fake data to approach that of real samples as close as possible in the finite-dimensional feature space. Empirical study on synthetic and real benchmarks shows the superior performance in quantitative evaluation and demonstrates the effectiveness of our approach to adversarial text generation. Yekun Chai, Qiyue Yin, Junge Zhang |
ICASSP | 3 |
| 2023 | S-NeRF: Neural Radiance Fields for Street Views
Ziyang Xie, Junge Zhang, Wenye Li 0002, Feihu Zhang, Li Zhang 0001 |
ICLR | 2 |
| 2023 | Exploration via Joint Policy Diversity for Sparse-Reward Multi-Agent TasksabstractExploration under sparse rewards is a key challenge for multi-agent reinforcement learning problems. Previous works argue that complex dynamics between agents and the huge exploration space in MARL scenarios amplify the vulnerability of classical count-based exploration methods when combined with agents parameterized by neural networks, resulting in inefficient exploration. In this paper, we show that introducing constrained joint policy diversity into a classical count-based method can significantly improve exploration when agents are parameterized by neural networks. Specifically, we propose a joint policy diversity to measure the difference between current joint policy and previous joint policies, and then use a filtering-based exploration constraint to further refine the joint policy diversity. Under the sparse-reward setting, we show that the proposed method significantly outperforms the state-of-the-art methods in the multiple-particle environment, the Google Research Football, and StarCraft II micromanagement tasks. To the best of our knowledge, on the hard 3s_vs_5z task which needs non-trivial strategies to defeat enemies, our method is the first to learn winning strategies without domain knowledge under the sparse-reward setting. Pei Xu 0003, Junge Zhang, Kaiqi Huang |
IJCAI | 2 |
| 2023 | Neural Text Classification by Jointly Learning to Cluster and AlignabstractDistributional text clustering delivers semantically informative representations and captures the relevance between each word and semantic clustering centroids. We extend the neural text clustering approach to text classification tasks by inducing cluster centers via a variational autoencoder and interacting with distributional word embeddings, to enrich the text representation and measure the relatedness between tokens and each learnable cluster centroid. The proposed method jointly learns word clustering centroids and cluster-token alignments, achieving competitive results on multiple benchmark datasets and proving that the proposed cluster-token alignment mechanism is favorable to text classification. Notably, the learned text representations are well-clustered, which matches the ground-truth categories. Experimental results show that our model can also improve the classification performance on top of BERT representations. To the best of our knowledge, we are the first adopting the variational autoencoder to update clustering centroids for text classification. Yekun Chai, Qiyue Yin, Junge Zhang |
IJCNN | 4 |
| 2023 | Explicitly Learning Policy Under Partial Observability in Multiagent Reinforcement LearningabstractWe explore explicit solutions for multiagent reinforcement learning (MARL) under the constraint of partial observability. With a general framework of centralized training with decentralized execution (CTDE), existing methods implicitly alleviate partial observability by introducing global information during centralized training. However, such implicit solution cannot well address partial observability and shows low sample efficiency in many MARL problems. In this paper, we focus on the influence of partial observability on the policy of agents, and formally derive an ideal form of policy that maximizes MARL objective under partial observability. Furthermore, we develop a new method named Explicitly Learning Policy (ELP), which adopts a novel teacher-student structure and utilizes knowledge distillation to explicitly learn individual policy under partial observability for each agent. Compared to prior methods, ELP presents a more general and interpretable training process, and the procedure of ELP can be easily extended to existing methods for performance boost. Our empirical experiments on StarCraft II micromanagement benchmark show that ELP significantly outperforms prevailing state-of-the-art baselines, which demonstrates the advantage of ELP in addressing partial observability and improving sample efficiency. Guangkai Yang, Hao Chen 0103, Junge Zhang |
IJCNN | 4 |
| 2023 | Underexplored Subspace Mining for Sparse-Reward Cooperative Multi-Agent Reinforcement LearningabstractLearning cooperation in sparse-reward multi-agent reinforcement learning is challenging, since agents need to explore in the large joint-state space with sparse feedback. However, in cooperative games, the cooperative target is often related to partial attributes, hence there is no need to treat the whole state space equally. Therefore, we propose Underexplored Subspace Mining (USM), a novel type of intrinsic reward that encourages agents to selectively explore partial attributes instead of wasting time on the whole state space to accelerate learning. Specially, considering that the target-related attributes are varying in different games and hard to predefine, we choose to focus on the underexplored subspace as an alternative, which is an automatic aggregation of the underexplored bottom-level dimensions without any human design or learning parameters. We evaluate our method in cooperative games with discrete and continuous state space separately. Results demonstrate that USM consistently outperforms existing state-of-the-art methods, and becomes the only method that has succeeded in sparse-reward games evaluated with larger state space or more complicated cooperation dynamics. Yang Yu 0056, Qiyue Yin, Junge Zhang, Hao Chen 0103, Kaiqi Huang |
IJCNN | 3 |
| 2023 | Dynamic Equilibrium-Based Continual Learning Model with Disentangled Meta-featuresabstractThe field of artificial intelligence research has witnessed remarkable advancements in recent decades. How-ever, conventional approaches in AI research primarily depend on fixed datasets and stationary settings, which have limited applicability to real-world scenarios. In order to address this limitation, there is an increasing need to develop and study algorithms and methods for continual learning, which enables artificial systems to learn from a continuous stream of data. One of the key challenges in continual learning is to strike a balance between transfer and interference, and to identify an equilibrium solution that can effectively learn from non-stationary data. This paper presents a canonical model that is specifically designed for continual learning, utilizing the derivative of the loss function to evaluate parameter changes between tasks and achieve dynamic equilibrium. Additionally, to improve the efficiency of limited training samples in continual tasks, a feature learning method based on meta-feature disentangling is proposed. By leveraging the second derivative term of the canonical model, the parameter vector can be decoupled and meta-features can be discovered. Experimental results demon-strate the superiority of the proposed method over state-of-the-art methods in continual lifelong supervised learning benchmarks. The validity of the proposed canonical model is further supported by these experimental results. As the demands of available settings become increasingly stringent, the advantages of disentangling meta-features become more prominent, resulting in a significant performance gap with other continual learning methods. Mingyi Zhang 0004, Junge Zhang |
SMC | 2 |
| 2023 | Direction-aware attention aggregation for single-stage hazy-weather crowd counting
Weihang Kong, Jienan Shen, He Li 0053, Junge Zhang |
Expert Syst. Appl. | 5 |
| 2023 | CSA-Net: Cross-modal scale-aware attention-aggregated network for RGB-T crowd counting
He Li 0053, Junge Zhang, Weihang Kong, Jienan Shen, Yuguang Shao |
Expert Syst. Appl. | 2 |
| 2022 | Achieving Consensus to Learn an Efficient and Robust Communication via Reinforcement Learning
Wei Qing, Zhaofeng He 0001, Junge Zhang, Luzhan Yuan, Wei Wang 0353 |
CogSci | 3 |
| 2022 | Multi-Agent Uncertainty Sharing for Cooperative Multi-Agent Reinforcement LearningabstractCooperative multi-agent reinforcement learning has been considered promising to complete many complex cooperative tasks in the real world such as coordination of robot swarms and self-driving. To promote multi-agent cooperation, Centralized Training with Decentralized Execution emerges as a popular learning paradigm due to partial observability and communication constraints during execution and computational complexity in training. Value decomposition has been known to produce competitive performance to other methods in complex environment within this paradigm such as VDN and QMIX, which approximates the global joint Q-value function with multiple local individual Q-value functions. However, existing works often neglect the uncertainty of multiple agents resulting from the partial observability and very large action space in the multi-agent setting and can only obtain the sub-optimal policy. To alleviate the limitations above, building upon the value decomposition, we propose a novel method called multi-agent uncertainty sharing (MAUS). This method utilizes the Bayesian neural network to explicitly capture the uncertainty of all agents and combines with Thompson sampling to select actions for policy learning. Besides, we impose the uncertainty-sharing mechanism among agents to stabilize training as well as coordinate the behaviors of all the agents for multi-agent cooperation. Extensive experiments on the StarCraft Multi-Agent Challenge (SMAC) environment demonstrate that our approach achieves significant performance to exceed the prior baselines and verify the effectiveness of our method. Hao Chen 0103, Guangkai Yang, Junge Zhang, Qiyue Yin, Kaiqi Huang |
IJCNN | 3 |
| 2022 | RACA: Relation-Aware Credit Assignment for Ad-Hoc Cooperation in Multi-Agent Deep Reinforcement LearningabstractIn recent years, reinforcement learning has faced several challenges in the multi-agent domain, such as the credit assignment issue. Value function factorization emerges as a promising way to handle the credit assignment issue under the centralized training with decentralized execution (CTDE) paradigm. However, existing value function factorization methods cannot deal with ad-hoc cooperation, that is, adapting to new configurations of teammates at test time. Specifically, these methods do not explicitly utilize the relationship between agents and cannot adapt to different sizes of inputs. To address these limitations, we propose a novel method, called Relation-Aware Credit Assignment (RACA), which achieves zero-shot generalization in ad-hoc cooperation scenarios. RACA takes advantage of a graph-based relation encoder to encode the topological structure between agents. Furthermore, RACA utilizes an attention-based observation abstraction mechanism that can generalize to an arbitrary number of teammates with a fixed number of parameters. Experiments demonstrate that our method outperforms baseline methods on the StarCraftII micromanagement benchmark and ad-hoc cooperation scenarios. Hao Chen 0103, Guangkai Yang, Junge Zhang, Qiyue Yin, Kaiqi Huang |
IJCNN | 3 |
| 2022 | FGA-NAS: Fast Resource-Constrained Architecture Search by Greedy-ADMM AlgorithmabstractDifferentiable architecture search has demonstrated promising results in automatically designing neural network architectures with desired properties, such as high accuracy and low FLOPs. However, it suffers from a cumbersome training process, and the injection of constraints in the search phase often relies on some hand-crafted heuristic regularizers, the design of which typically requires tremendous human effort. In this paper, to address these critical challenges, we present FGA-NAS, an efficient method for resource-constrained architecture search. First, to reduce the computational cost and improve search flexibility, we propose a novel condensed search space that merges multiple parallel-placed candidates into a single one. Second, to enable the gradient-based optimization for neural architecture search (NAS) under multiple combinatorial constraints, we decompose the constrained NAS into a few simple sub-problems without introducing any heuristics by using the ADMM algorithm [1]. Then, the constrained NAS can be resolved by alternately solving the simple sub-problems. Experimental results on ImageNet show that our method can discover efficient and accurate neural network architectures that achieve the state-of-the-art by only using 0.2 GPU days. Junge Zhang, Qiaozhe Li, Hao Chen 0103, Kaiqi Huang |
IJCNN | 2 |
| 2022 | Offline reinforcement learning with representations for actions
Xingzhou Lou, Qiyue Yin, Junge Zhang, Chao Yu 0004, Zhaofeng He 0001, Nengjie Cheng, Kaiqi Huang |
Inf. Sci. | 3 |
| 2022 | Deep Reinforcement Learning With Part-Aware Exploration Bonus in Video GamesabstractReinforcement learning algorithms rely on carefully engineering environment rewards that are extrinsic to agents. However, environments with dense rewards are rare, motivating the need for developing reward functions that are intrinsic to agents. Curiosity is a type of successful intrinsic reward function, which uses the prediction error as an reward signal. In prior work, the prediction problem used to generate intrinsic rewards is optimized in the pixel space rather than a learnable feature space to avoid randomness caused by feature changes. However, these methods ignore small but important elements of the states that are often associated with locations of the character, which makes it impossible to generate accurate internal rewards for efficient exploration. In this article, we first demonstrate the effectiveness of introducing prior learned features for existing prediction-based exploration methods. Then, an attention map mechanism is designed to discretize learned features, thereby updating the learned feature and meanwhile reducing the impact of randomness on intrinsic rewards caused by the learning process of features. We verify our method on some video games from the standard reinforcement learning Atari benchmark, achieving clear improvements over random network distillation, which is one of the most advanced exploration methods, in almost all Atari games. Pei Xu 0003, Qiyue Yin, Junge Zhang, Kaiqi Huang |
IEEE Trans. Games | 3 |
| 2021 | Learning to Reweight Imaginary Transitions for Model-Based Reinforcement Learning
Wenzhen Huang, Qiyue Yin, Junge Zhang, Kaiqi Huang |
AAAI | 3 |
| 2021 | CCF-Net: Composite Context Fusion Network with Inter-Slice Correlative Fusion for Multi-Disease Lesion DetectionabstractDetecting lesions from computed tomography (CT) scans relies on two aspects of the input: intra-slice texture information from the key slice and inter-slice structural context information from the adjacent slices. However, most existing methods ignore the correlation and complementarity between texture and structural information resulting in unexpected loss of performance. In this paper, a novel Composite Context Fusion Network (CCF-Net) is proposed to jointly model intra-slice and inter-slice features so as to prove the effectiveness of the two-steam framework. To extract both texture and structural information, two streams of 2D and 3D convolutional modules are employed in each stage. Moreover, a Composite Fusion architecture equipped with Inter-slice Correlative Fusion (ICF) modules is proposed to achieve stage-by-stage feature fusion in order to excavate and exchange information between texture-aware and context-aware features. Extensive experiments show that the proposed CCF-Net is able to achieve state-of-the-art detection performance on the multi-disease CT lesion detection task and significantly surpass the baseline methods.1 Jiechao Ma, Shu Zhang 0001, Yemin Shi 0001, Junge Zhang, Kaiqi Huang, Yizhou Yu |
ICIP | 5 |
| 2021 | SOFT: Softmax-free Transformer with Linear ComplexityabstractVision transformers (ViTs) have pushed the state-of-the-art for various visual recognition tasks by patch-wise image tokenization followed by self-attention. However, the employment of self-attention modules results in a quadratic complexity in both computation and memory usage. Various attempts on approximating the self-attention computation with linear complexity have been made in Natural Language Processing. However, an in-depth analysis in this work shows that they are either theoretically flawed or empirically ineffective for visual recognition. We further identify that their limitations are rooted in keeping the softmax self-attention during approximations. Specifically, conventional self-attention is computed by normalizing the scaled dot-product between token feature vectors. Keeping this softmax operation challenges any subsequent linearization efforts. Based on this insight, for the first time, a softmax-free transformer or SOFT is proposed. To remove softmax in self-attention, Gaussian kernel function is used to replace the dot-product similarity without further normalization. This enables a full self-attention matrix to be approximated via a low-rank matrix decomposition. The robustness of the approximation is achieved by calculating its Moore-Penrose inverse using a Newton-Raphson method. Extensive experiments on ImageNet show that our SOFT significantly improves the computational efficiency of existing ViT variants. Crucially, with a linear complexity, much longer token sequences are permitted in SOFT, resulting in superior trade-off between accuracy and complexity. Jinghan Yao, Junge Zhang, Xiatian Zhu, Hang Xu 0004, Weiguo Gao, Chunjing Xu, Tao Xiang 0002, Li Zhang 0040 |
NeurIPS | 3 |
| 2021 | Coordinated Proximal Policy OptimizationabstractWe present Coordinated Proximal Policy Optimization (CoPPO), an algorithm that extends the original Proximal Policy Optimization (PPO) to the multi-agent setting. The key idea lies in the coordinated adaptation of step size during the policy update process among multiple agents. We prove the monotonicity of policy improvement when optimizing a theoretically-grounded joint objective, and derive a simplified optimization objective based on a set of approximations. We then interpret that such an objective in CoPPO can achieve dynamic credit assignment among agents, thereby alleviating the high variance issue during the concurrent update of agent policies. Finally, we demonstrate that CoPPO outperforms several strong baselines and is competitive with the latest multi-agent PPO method (i.e. MAPPO) under typical multi-agent settings, including cooperative matrix games and the StarCraft II micromanagement tasks. Zifan Wu, Chao Yu 0004, Deheng Ye, Junge Zhang, Haiyin Piao, Hankui Zhuo |
NeurIPS | 4 |
| 2021 | Consistency Regularization for Ensemble Model Based Reinforcement Learning
Ruonan Jia, Qingming Li, Wenzhen Huang, Junge Zhang, Xiu Li 0001 |
PRICAI (3) | 4 |
| 2021 | Universal adversarial perturbations against object detection
Debang Li, Junge Zhang, Kaiqi Huang |
Pattern Recognit. | 2 |
| 2020 | Composing Good Shots by Exploiting Mutual RelationsabstractFinding views with a good composition from an input image is a common but challenging problem. There are usually at least dozens of candidates (regions) in an image, and how to evaluate these candidates is subjective. Most existing methods only use the feature corresponding to each candidate to evaluate the quality. However, the mutual relations between the candidates from an image play an essential role in composing a good shot due to the comparative nature of this problem. Motivated by this, we propose a graph-based module with a gated feature update to model the relations between different candidates. The candidate region features are propagated on a graph that models mutual relations between different regions for mining the useful information such that the relation features and region features are adaptively fused. We design a multi-task loss to train the model, especially, a regularization term is adopted to incorporate the prior knowledge about the relations into the graph. A data augmentation method is also developed by mixing nodes from different graphs to improve the model generalization ability. Experimental results show that the proposed model performs favorably against state-of-the-art methods, and comprehensive ablation studies demonstrate the contribution of each module and graph-based inference of the proposed method. Debang Li, Junge Zhang, Kaiqi Huang, Ming-Hsuan Yang 0001 |
CVPR | 2 |
| 2020 | Learning to Learn Cropping Models for Different Aspect Ratio RequirementsabstractImage cropping aims at improving the framing of an image by removing its extraneous outer areas, which is widely used in the photography and printing industry. In some cases, the aspect ratio of cropping results is specified depending on some conditions. In this paper, we propose a meta-learning (learning to learn) based aspect ratio specified image cropping method called Mars, which can generate cropping results of different expected aspect ratios. In the proposed method, a base model and two meta-learners are obtained during the training stage. Given an aspect ratio in the test stage, a new model with new parameters can be generated from the base model. Specifically, the two meta-learners predict the parameters of the base model based on the given aspect ratio. The learning process of the proposed method is learning how to learn cropping models for different aspect ratio requirements, which is a typical meta-learning process. In the experiments, the proposed method is evaluated on three datasets and outperforms most state-of-the-art methods in terms of accuracy and speed. In addition, both the intermediate and final results show that the proposed model can predict different cropping windows for an image depending on different aspect ratio requirements. Debang Li, Junge Zhang, Kaiqi Huang |
CVPR | 2 |
| 2019 | Bootstrap Estimated Uncertainty of the Environment Model for Model-Based Reinforcement LearningabstractModel-based reinforcement learning (RL) methods attempt to learn a dynamics model to simulate the real environment and utilize the model to make better decisions. However, the learned environment simulator often has more or less model error which would disturb making decision and reduce performance. We propose a bootstrapped model-based RL method which bootstraps the modules in each depth of the planning tree. This method can quantify the uncertainty of environment model on different state-action pairs and lead the agent to explore the pairs with higher uncertainty to reduce the potential model errors. Moreover, we sample target values from their bootstrap distribution to connect the uncertainties at current and subsequent time-steps and introduce the prior mechanism to improve the exploration efficiency. Experiment results demonstrate that our method efficiently decreases model error and outperforms TreeQN and other stateof-the-art methods on multiple Atari games. Wenzhen Huang, Junge Zhang, Kaiqi Huang |
AAAI | 2 |
| 2019 | Transductive Zero-Shot Learning via Visual Center AdaptationabstractIn this paper, we propose a Visual Center Adaptation Method (VCAM) to address the domain shift problem in zero-shot learning. For the seen classes in the training data, VCAM builds an embedding space by learning the mapping from semantic space to some visual centers. While for unseen classes in the test data, the construction of embedding space is constrained by a symmetric Chamfer-distance term, aiming to adapt the distribution of the synthetic visual centers to that of the real cluster centers. Therefore the learned embedding space can generalize the unseen classes well. Experiments on two widely used datasets demonstrate that our model significantly outperforms state-of-the-art methods. Ziyu Wan, Yan Li 0043, Junge Zhang |
AAAI | 4 |
| 2019 | Few-Shot Image Recognition With Knowledge TransferabstractHuman can well recognize images of novel categories just after browsing few examples of these categories. One possible reason is that they have some external discriminative visual information about these categories from their prior knowledge. Inspired from this, we propose a novel Knowledge Transfer Network architecture (KTN) for few-shot image recognition. The proposed KTN model jointly incorporates visual feature learning, knowledge inferring and classifier learning into one unified framework for their optimal compatibility. First, the visual classifiers for novel categories are learned based on the convolutional neural network with the cosine similarity optimization. To fully explore the prior knowledge, a semantic-visual mapping network is then developed to conduct knowledge inference, which enables to infer the classifiers for novel categories from base categories. Finally, we design an adaptive fusion scheme to infer the desired classifiers by effectively integrating the above knowledge and visual information. Extensive experiments are conducted on two widely-used Mini-ImageNet and ImageNet Few-Shot benchmarks to evaluate the effectiveness of the proposed method. The results compared with the state-of-the-art approaches show the encouraging performance of the proposed method, especially on 1-shot and 2-shot tasks. Zhimao Peng, Zechao Li, Junge Zhang, Guo-Jun Qi, Jinhui Tang 0001 |
ICCV | 3 |
| 2019 | SparseMask: Differentiable Connectivity Learning for Dense Image PredictionabstractIn this paper, we aim at automatically searching an efficient network architecture for dense image prediction. Particularly, we follow the encoder-decoder style and focus on designing a connectivity structure for the decoder. To achieve that, we design a densely connected network with learnable connections, named Fully Dense Network, which contains a large set of possible final connectivity structures. We then employ gradient descent to search the optimal connectivity from the dense connections. The search process is guided by a novel loss function, which pushes the weight of each connection to be binary and the connections to be sparse. The discovered connectivity achieves competitive results on two segmentation datasets, while runs more than three times faster and requires less than half parameters compared to the state-of-the-art methods. An extensive experiment shows that the discovered connectivity is compatible with various backbones and generalizes well to other dense image prediction tasks. Huikai Wu, Junge Zhang, Kaiqi Huang |
ICCV | 2 |
| 2019 | MVP-Net: Multi-view FPN with Position-Aware Attention for Deep Universal Lesion Detection
Shu Zhang 0001, Junge Zhang, Kaiqi Huang, Yizhou Wang 0001, Yizhou Yu |
MICCAI (6) | 3 |
| 2019 | GP-GAN: Towards Realistic High-Resolution Image BlendingabstractIt is common but challenging to address high-resolution image blending in the automatic photo editing application. In this paper, we would like to focus on solving the problem of high-resolution image blending, where the composite images are provided. We propose a framework called Gaussian-Poisson Generative Adversarial Network (GP-GAN) to leverage the strengths of the classical gradient-based approach and Generative Adversarial Networks. To the best of our knowledge, it's the first work that explores the capability of GANs in high-resolution image blending task. Concretely, we propose Gaussian-Poisson Equation to formulate the high-resolution image blending problem, which is a joint optimization constrained by the gradient and color information. Inspired by the prior works, we obtain gradient information via applying gradient filters. To generate the color information, we propose a Blending GAN to learn the mapping between the composite images and the well-blended ones. Compared to the alternative methods, our approach can deliver high-resolution, realistic images with fewer bleedings and unpleasant artifacts. Experiments confirm that our approach achieves the state-of-the-art performance on Transient Attributes dataset. A user study on Amazon Mechanical Turk finds that the majority of workers are in favor of the proposed method. The source code is available in \urlhttps://github.com/wuhuikai/GP-GAN, and there's also an online demo in \urlhttp://wuhuikai.me/DeepJS. Huikai Wu, Shuai Zheng 0001, Junge Zhang, Kaiqi Huang |
ACM Multimedia | 3 |
| 2019 | Transductive Zero-Shot Learning with Visual Structure ConstraintabstractTo recognize objects of the unseen classes, most existing Zero-Shot Learning (ZSL) methods first learn a compatible projection function between the common semantic space and the visual space based on the data of source seen classes, then directly apply it to the target unseen classes. However, in real scenarios, the data distribution between the source and target domain might not match well, thus causing the well-known domain shift problem. Based on the observation that visual features of test instances can be separated into different clusters, we propose a new visual structure constraint on class centers for transductive ZSL, to improve the generality of the projection function (\ie alleviate the above domain shift problem). Specifically, three different strategies (symmetric Chamfer-distance,Bipartite matching distance, and Wasserstein distance) are adopted to align the projected unseen semantic centers and visual cluster centers of test instances. We also propose a new training strategy to handle the real cases where many unrelated images exist in the test dataset, which is not considered in previous methods. Experiments on many widely used datasets demonstrate that the proposed visual structure constraint can bring substantial performance gain consistently and achieve state-of-the-art results. Ziyu Wan, Dongdong Chen 0001, Yan Li 0043, Xingguang Yan, Junge Zhang, Yizhou Yu, Jing Liao 0001 |
NeurIPS | 5 |
| 2019 | Semi-supervised Lesion Detection with Reliable Label Propagation and Missing Label Mining
Shu Zhang 0001, Junge Zhang, Kaiqi Huang |
PRCV (2) | 4 |
| 2019 | Mixed Supervised Object Detection with Robust Objectness TransferabstractIn this paper, we consider the problem of leveraging existing fully labeled categories to improve the weakly supervised detection (WSD) of new object categories, which we refer to as mixed supervised detection (MSD). Different from previous MSD methods that directly transfer the pre-trained object detectors from existing categories to new categories, we propose a more reasonable and robust objectness transfer approach for MSD. In our framework, we first learn domain-invariant objectness knowledge from the existing fully labeled categories. The knowledge is modeled based on invariant features that are robust to the distribution discrepancy between the existing categories and new categories; therefore the resulting knowledge would generalize well to new categories and could assist detection models to reject distractors (e.g., object parts) in weakly labeled images of new categories. Under the guidance of learned objectness knowledge, we utilize multiple instance learning (MIL) to model the concepts of both objects and distractors and to further improve the ability of rejecting distractors in weakly labeled images. Our robust objectness transfer approach outperforms the existing MSD methods, and achieves state-of-the-art results on the challenging ILSVRC2013 detection dataset and the PASCAL VOC datasets. Yan Li 0043, Junge Zhang, Kaiqi Huang, Jianguo Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Multi-view clustering via joint feature selection and partially constrained cluster label learning
Qiyue Yin, Junge Zhang, Hexi Li |
Pattern Recognit. | 2 |
| 2019 | Fast A3RL: Aesthetics-Aware Adversarial Reinforcement Learning for Image CroppingabstractImage cropping aims at improving the quality of images by removing unwanted outer areas, which is widely used in the photography and printing industry. Most previous cropping methods that don't need bounding box supervision rely on the sliding window mechanism. The sliding window method results in fixed aspect ratios and limits the shape of the cropping region. Moreover, the sliding window method usually produces lots of candidates on the input image, which is very time-consuming. Motivated by these challenges, we formulate image cropping as a sequential decision-making process and propose a reinforcement learning based framework to address this problem, namely Fast Aesthetics-Aware Adversarial Reinforcement Learning (Fast A3RL). Particularly, the proposed method develops an aesthetics-aware reward function, which is dedicated for image cropping. Similar to human's decisionmaking process, we use a comprehensive state representation including both the current observation and historical experience. We train the agent using the actor-critic architecture in an end-to-end manner. The adversarial learning process is also applied during the training stage. The proposed method is evaluated on several popular cropping datasets, in which the images are unseen during training. Experiment results show that our method achieves state-of-the-art performance with much fewer candidate windows and much less time compared with related methods. Debang Li, Huikai Wu, Junge Zhang, Kaiqi Huang |
IEEE Trans. Image Process. | 3 |
| 2018 | Deep Semantic Structural Constraints for Zero-Shot LearningabstractZero-shot learning aims to classify unseen image categories by learning a visual-semantic embedding space. In most cases, the traditional methods adopt a separated two-step pipeline that extracts image features are utilized to learn the embedding space. It leads to the lack of specific structural semantic information of image features for zero-shot learning task. In this paper, we propose an end-to-end trainable Deep Semantic Structural Constraints model to address this issue. The proposed model contains the Image Feature Structure constraint and the Semantic Embedding Structure constraint, which aim to learn structure-preserving image features and endue the learned embedding space with stronger generalization ability respectively. With the assistance of semantic structural information, the model gains more auxiliary clues for zero-shot learning. The state-of-the-art performance certifies the effectiveness of our proposed method. Yan Li 0043, Junge Zhang, Kaiqi Huang, Tieniu Tan |
AAAI | 3 |
| 2018 | DF2Net: Discriminative Feature Learning and Fusion Network for RGB-D Indoor Scene ClassificationabstractThis paper focuses on the task of RGB-D indoor scene classification. It is a very challenging task due to two folds. 1) Learning robust representation for indoor scene is difficult because of various objects and layouts. 2) Fusing the complementary cues in RGB and Depth is nontrivial since there are large semantic gaps between the two modalities. Most existing works learn representation for classification by training a deep network with softmax loss and fuse the two modalities by simply concatenating the features of them. However, these pipelines do not explicitly consider intra-class and inter-class similarity as well as inter-modal intrinsic relationships. To address these problems, this paper proposes a Discriminative Feature Learning and Fusion Network (DF2Net) with two-stage training. In the first stage, to better represent scene in each modality, a deep multi-task network is constructed to simultaneously minimize the structured loss and the softmax loss. In the second stage, we design a novel discriminative fusion network which is able to learn correlative features of multiple modalities and distinctive features of each modality. Extensive analysis and experiments on SUN RGB-D Dataset and NYU Depth Dataset V2 show the superiority of DF2Net over other state-of-the-art methods in RGB-D indoor scene classification task. Yabei Li, Junge Zhang, Yanhua Cheng, Kaiqi Huang, Tieniu Tan |
AAAI | 2 |
| 2018 | A2-RL: Aesthetics Aware Reinforcement Learning for Image CroppingabstractImage cropping aims at improving the aesthetic quality of images by adjusting their composition. Most weakly supervised cropping methods (without bounding box supervision) rely on the sliding window mechanism. The sliding window mechanism requires fixed aspect ratios and limits the cropping region with arbitrary size. Moreover, the sliding window method usually produces tens of thousands of windows on the input image which is very time-consuming. Motivated by these challenges, we firstly formulate the aesthetic image cropping as a sequential decision-making process and propose a weakly supervised Aesthetics Aware Reinforcement Learning (A2-RL) framework to address this problem. Particularly, the proposed method develops an aesthetics aware reward function which especially benefits image cropping. Similar to human's decision making, we use a comprehensive state representation including both the current observation and the historical experience. We train the agent using the actor-critic architecture in an end-to-end manner. The agent is evaluated on several popular unseen cropping datasets. Experiment results show that our method achieves the state-of-the-art performance with much fewer candidate windows and much less time compared with previous weakly supervised methods. Debang Li, Huikai Wu, Junge Zhang, Kaiqi Huang |
CVPR | 3 |
| 2018 | Discriminative Learning of Latent Features for Zero-Shot RecognitionabstractZero-shot learning (ZSL) aims to recognize unseen image categories by learning an embedding space between image and semantic representations. For years, among existing works, it has been the center task to learn the proper mapping matrices aligning the visual and semantic space, whilst the importance to learn discriminative representations for ZSL is ignored. In this work, we retrospect existing methods and demonstrate the necessity to learn discriminative representations for both visual and semantic instances of ZSL. We propose an end-to-end network that is capable of 1) automatically discovering discriminative regions by a zoom network; and 2) learning discriminative semantic representations in an augmented space introduced for both user-defined and latent attributes. Our proposed method is tested extensively on two challenging ZSL datasets, and the experiment results show that the proposed method significantly outperforms state-of-the-art methods. Yan Li 0043, Junge Zhang, Jianguo Zhang 0001, Kaiqi Huang |
CVPR | 2 |
| 2018 | Fast End-to-End Trainable Guided FilterabstractImage processing and pixel-wise dense prediction have been advanced by harnessing the capabilities of deep learning. One central issue of deep learning is the limited capacity to handle joint upsampling. We present a deep learning building block for joint upsampling, namely guided filtering layer. This layer aims at efficiently generating the high-resolution output given the corresponding low-resolution one and a high-resolution guidance map. The proposed layer is composed of a guided filter, which is reformulated as a fully differentiable block. To this end, we show that a guided filter can be expressed as a group of spatial varying linear transformation matrices. This layer could be integrated with the convolutional neural networks (CNNs) and jointly optimized through end-to-end training. To further take advantage of end-to-end training, we plug in a trainable transformation function that generates task-specific guidance maps. By integrating the CNNs and the proposed layer, we form deep guided filtering networks. The proposed networks are evaluated on five advanced image processing tasks. Experiments on MIT-Adobe FiveK Dataset demonstrate that the proposed approach runs 10-100× faster and achieves the state-of-the-art performance. We also show that the proposed guided filtering layer helps to improve the performance of multiple pixel-wise dense prediction tasks. The code is available at https://github.com/wuhuikai/DeepGuidedFilter. Huikai Wu, Shuai Zheng 0001, Junge Zhang, Kaiqi Huang |
CVPR | 3 |
| 2018 | ACM: Learning Dynamic Multi-agent Cooperation via Attentional Communication Model
Hongping Yan, Junge Zhang, Lingfeng Wang 0002 |
ICANN (2) | 3 |
| 2017 | Encyclopedia enhanced semantic embedding for zero-shot learningabstractThere are tremendous object categories in the real world besides those in image datasets. Zero-shot learning aims to recognize image categories which are unseen in the training set. A large number of previous zero-shot learning models use word vectors of the class labels directly as category prototypes in the semantic embedding space. But word vectors cannot obtain the global knowledge of an image category sufficiently. In this paper, we propose a new encyclopedia enhanced semantic embedding model to promote the discriminative capability of word vector prototypes with the global knowledge of each image category. The proposed model extracts the TF-IDF key words from encyclopedia articles to acquire the global knowledge of each category. The convex combination of the key words' word vectors acts as the prototypes of the object categories. The prototypes of seen and unseen classes build up the embedding space where the nearest neighbour search is implemented to recognize the unseen images. The experiments show that the proposed method achieves the state-of-the-art performance on the challenging ImageNet Fall 2011 1k2hop dataset. Junge Zhang, Kaiqi Huang, Tieniu Tan |
ICIP | 2 |
| 2017 | Semantics-guided multi-level RGB-D feature fusion for indoor semantic segmentationabstractIndoor RGB-D semantic segmentation is a new and challenging problem. Traditional methods usually apply two-stream convolutional neural networks (CNNs) to represent RGB and depth images respectively, and fuse the two streams on a specific layer. In this paper, we explore several fusion strategies based on this two-stream-CNN framework and point out such a single-layer fusion method cannot exploit the complementary RGB and depth cues well for semantic segmentation. To address this problem, we propose a novel Semantics-guided Multi-level feature fusion approach, which first learns deep feature representation from bottom to up, and then gradually fuses the RGB and depth features from high level to low level under the guidance of the semantic cues. Experimental results on SUN RGB-D dataset demonstrate the advantages of the proposed method over the state of the arts. Yabei Li, Junge Zhang, Yanhua Cheng, Kaiqi Huang, Tieniu Tan |
ICIP | 2 |
| 2017 | Local structured representation for generic object detection
Junge Zhang, Kaiqi Huang, Tieniu Tan, Zhaoxiang Zhang 0001 |
Frontiers Comput. Sci. | 1 |
| 2017 | GRMA: Generalized Range Move Algorithms for the Efficient Optimization of MRFs
Junge Zhang, Peipei Yang, Stephen J. Maybank, Kaiqi Huang |
Int. J. Comput. Vis. | 2 |
| 2017 | ORGM: Occlusion Relational Graphical Model for Human Pose EstimationabstractArticulated human pose estimation from monocular image is a challenging problem in computer vision. Occlusion is a main challenge for human pose estimation, which is largely ignored in popular tree structured models. The tree structured model is simple and convenient for exact inference, but short in modeling the occlusion coherence especially in the case of self-occlusion. We propose an occlusion relational graphical model, which is able to model both self-occlusion and occlusion by the other objects simultaneously. The proposed model can encode the interactions between human body parts and objects, and enables it to learn occlusion coherence from data discriminatively. We evaluate our model on several public benchmarks for human pose estimation, including challenging subsets featuring significant occlusion. The experimental results show that our method is superior to the previous state-of-the-arts, and is robust to occlusion for 2D human pose estimation. Lianrui Fu, Junge Zhang, Kaiqi Huang |
IEEE Trans. Image Process. | 2 |
| 2016 | FastLCD: Fast Label Coordinate Descent for the Efficient Optimization of 2D Label MRFs
Junge Zhang, Peipei Yang, Kaiqi Huang |
IJCAI | 2 |
| 2015 | GRSA: Generalized range swap algorithm for the efficient optimization of MRFsabstractMarkov Random Field (MRF) is an important tool and has been widely used in many vision tasks. Thus, the optimization of MRFs is a problem of fundamental importance. Recently, Veskler and Kumar et. al propose the range move algorithms, which are one of the most successful solvers to this problem. However, two problems have limited the applicability of previous range move algorithms: 1) They are limited in the types of energies they can handle (i.e. only truncated convex functions); 2) These algorithms tend to be very slow compared to other graph-cut based algorithms (e.g. α-expansion and αβ-swap). In this paper, we propose a generalized range swap algorithm (GRSA) for efficient optimization of MRFs. To address the first problem, we extend the GRSA to arbitrary semimetric energies by restricting the chosen labels in each move so that the energy is submodular on the chosen subset. Furthermore, to feasibly choose the labels satisfying the submodular condition, we provide a sufficient condition of the submodularity. For the second problem, unlike previous range move algorithms which execute the set of all possible range moves, we dynamically obtain the iterative moves by solving a set cover problem, which greatly reduces the number of moves during the optimization. Experiments show that the GRSA offers a great speedup over previous range swap algorithms, while it obtains competitive solutions. Junge Zhang, Peipei Yang, Kaiqi Huang |
CVPR | 2 |
| 2015 | Beyond Tree Structure Models: A New Occlusion Aware Graphical Model for Human Pose EstimationabstractOcclusion is a main challenge for human pose estimation, which is largely ignored in popular tree structure models. The tree structure model is simple and convenient for exact inference, but short in modeling the occlusion coherence especially in the case of self-occlusion. We propose an occlusion aware graphical model which is able to model both self-occlusion and occlusion by the other objects simultaneously. The proposed model structure can encodes the interactions between human body parts and objects, and hence enables it to learn occlusion coherence from data discriminatively. We evaluate our model on several public benchmarks for human pose estimation including challenging subsets featuring significant occlusion. The experimental results show that our method obtains comparable accuracy with the state-of-the-arts, and is robust to occlusion for 2D human pose estimation. Lianrui Fu, Junge Zhang, Kaiqi Huang |
ICCV | 2 |
| 2015 | Context aware model for articulated human pose estimationabstractSimple tree model prevails for 2D pose estimation for its simplicity and efficiency. However, the limited kinetic constraints often lead to double-counting and damage the accuracy of leaf parts, and this is largely ignored in previous work. In this paper, we propose a novel enhanced tree model which incorporates both local kinetic constraints and global contextual constraints among non-adjacent parts. By introducing virtual parts, we are able to model richer constraints within a tree structure and dynamic programming can be utilized for efficient inference. Experiments on public benchmarks show that our method is more effective in tackling double counting problem and can improve the localization accuracy, especially for the challenging lower limbs. Lianrui Fu, Junge Zhang, Kaiqi Huang |
ICIP | 2 |
| 2015 | ISEE Smart Home (ISH): Smart video analysis for home security
Junge Zhang, Yanhu Shan, Kaiqi Huang |
Neurocomputing | 1 |
| 2015 | Large-Scale Weakly Supervised Object Localization via Latent Category LearningabstractLocalizing objects in cluttered backgrounds is challenging under large-scale weakly supervised conditions. Due to the cluttered image condition, objects usually have large ambiguity with backgrounds. Besides, there is also a lack of effective algorithm for large-scale weakly supervised localization in cluttered backgrounds. However, backgrounds contain useful latent information, e.g., the sky in the aeroplane class. If this latent information can be learned, object-background ambiguity can be largely reduced and background can be suppressed effectively. In this paper, we propose the latent category learning (LCL) in large-scale cluttered conditions. LCL is an unsupervised learning method which requires only image-level class labels. First, we use the latent semantic analysis with semantic object representation to learn the latent categories, which represent objects, object parts or backgrounds. Second, to determine which category contains the target object, we propose a category selection strategy by evaluating each category's discrimination. Finally, we propose the online LCL for use in large-scale conditions. Evaluation on the challenging PASCAL Visual Object Class (VOC) 2007 and the large-scale imagenet large-scale visual recognition challenge 2013 detection data sets shows that the method can improve the annotation precision by 10% over previous methods. More importantly, we achieve the detection precision which outperforms previous results by a large margin and can be competitive to the supervised deformable part model 5.0 baseline on both data sets. Kaiqi Huang, Weiqiang Ren, Junge Zhang, Stephen J. Maybank |
IEEE Trans. Image Process. | 4 |
| 2014 | Deformable Object Matching via Deformation Decomposition Based 2D Label MRFabstractDeformable object matching, which is also called elastic matching or deformation matching, is an important and challenging problem in computer vision. Although numerous deformation models have been proposed in different matching tasks, not many of them investigate the intrinsic physics underlying deformation. Due to the lack of physical analysis, these models cannot describe the structure changes of deformable objects very well. Motivated by this, we analyze the deformation physically and propose a novel deformation decomposition model to represent various deformations. Based on the physical model, we formulate the matching problem as a two-mensional label Markov Random Field. The MRF energy function is derived from the deformation decomposition model. Furthermore, we propose a two-stage method to optimize the MRF energy function. To provide a quantitative benchmark, we build a deformation matching database with an evaluation criterion. Experimental results show that our method outperforms previous approaches especially on complex deformations. Junge Zhang, Kaiqi Huang, Tieniu Tan |
CVPR | 2 |
| 2014 | Improved Optimization Based on Graph Cuts for Discrete Energy MinimizationabstractDiscrete energy optimization is a NP hard problem. Recent years, the graph cuts based algorithms especially the a-expansion and aß-swap, become more and more popular. Both the a-expansion and a/3-swap have been widely used in many applications, and they perform extremely well for the Potts energies. However, since all pixels only have a choice of two labels in one move, both the expansion and swap algorithm get approximate solution by a series of iterations and they do not perform well for more general energies, such as the truncated convex energies [1]. In this paper, we analyze the problems of both the expansion and swap algorithms. The expansion algorithm usually encourages more pixels to get the label fa, since all pixels are only allowed to change their current labels to fa. In contrast, the swap move sometimes cannot swap the labels of pixels reasonably. Based on the analysis, we propose the Interleaved Expansion-Swap Algorithm (IESA) by combining the expansion and swap moves effectively. To prove the effectiveness of the algorithm, we test it on both image restoration and stereo correspondence. The experimental evaluations show that our algorithm gets better optimization compared with both a-expansion and aß-swap. Junge Zhang, Kaiqi Huang |
ICPR | 2 |
| 2014 | Learning Convolutional Nonlinear Features for K Nearest Neighbor Image ClassificationabstractLearning low-dimensional feature representations is a crucial task in machine learning and computer vision. Recently the impressive breakthrough in general object recognition made by large scale convolutional networks shows that convolutional networks are able to extract discriminative hierarchical features in large scale object classification task. However, for vision tasks other than end-to-end classification, such as K Nearest Neighbor classification, the learned intermediate features are not necessary optimal for the specific problem. In this paper, we aim to exploit the power of deep convolutional networks and optimize the output feature layer with respect to the task of K Nearest Neighbor (kNN) classification. By directly optimizing the kNN classification error on training data, we in fact learn convolutional nonlinear features in a data-driven and task-driven way. Experimental results on standard image classification benchmarks show that the proposed method is able to learn better feature representations than other general end-to-end classification methods on kNN classification task. Weiqiang Ren, Yinan Yu, Junge Zhang, Kaiqi Huang |
ICPR | 3 |
| 2014 | Robust Object Recognition via Visual Pathway FeedbackabstractObject recognition, which consists of classification and detection, has two important attributes for robustness: (1) Closeness: detection windows should be close to object locations, and (2) Adaptiveness: object matching should be adaptive to object variations in classification. It is difficult to satisfy both attributes by considering classification and detection separately, thus recent studies combine them based on confidence contextualization and foreground modeling. However, these combinations neglect feature saliency and object structure, which are important for recognition. In fact, object recognition originates in the mechanism of "what" and "where" pathways in human visual systems, and more importantly, these pathways have feedback to each other, which provides a probable way to improve closeness and adaptiveness. Inspired by the feedback, we propose a robust object recognition framework by designing a computational model of the feedback mechanism. In the "what" feedback, the feature saliency from classification is exploited to rectify detection windows for better closeness, while in the "where" feedback, object parts from detection are used to model object matching of object structure for better adaptiveness. Experiments show that the "what" and "where" feedback can be effective to improve closeness and adaptiveness for robust object recognition, and encouraging results are obtained on the challenging PASCAL VOC 2007 dataset. Junge Zhang, Peipei Yang, Kaiqi Huang |
ICPR | 2 |
| 2013 | Recent Progress on Object Classification and Detection
Tieniu Tan, Yongzhen Huang, Junge Zhang |
CIARP (2) | 3 |
| 2013 | Exploring the Power of Kernel in Feature Representation for Object Categorization
Weiqiang Ren, Yinan Yu, Junge Zhang, Kaiqi Huang |
ICONIP (3) | 3 |
| 2012 | Data Decomposition and Spatial Mixture Modeling for Part Based Model
Junge Zhang, Yongzhen Huang, Kaiqi Huang, Zifeng Wu, Tieniu Tan |
ACCV (1) | 1 |
| 2012 | Interest Point Selection with Spatio-temporal Context for Realistic Action RecognitionabstractSpatio-Temporal Interest Point (STIP) has been widely used for human action recognition. However, the performance of the STIP based methods are still limited in realistic datasets which often include large variations in illuminations, viewpoints and camera motions. One reason of the low performance is that the STIPs only reflect the local change in videos, which is not enough to obtain stable informative features for action representation in realistic scene. To tackle the problem, we proposed an approach to selecting the "stable STIPs" with the spatio-temporal distribution of STIPs in neighbor region. Then, BoW feature is constructed to represent actions with these selected points. The experimental results on KTH dataset and HMDB (the largest realistic human action dataset) demonstrate that the proposed approach has obvious effect on improving the recognition rates of realistic data. Yanhu Shan, Zhang Zhang 0001, Junge Zhang, Kaiqi Huang, Oh Se Hyun |
AVSS | 3 |
| 2012 | Semantic windows mining in sliding window based object detection
Junge Zhang, Xin Zhao 0012, Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
ICPR | 1 |
| 2011 | Boosted local structured HOG-LBP for object localizationabstractObject localization is a challenging problem due to variations in object's structure and illumination. Although existing part based models have achieved impressive progress in the past several years, their improvement is still limited by low-level feature representation. Therefore, this paper mainly studies the description of object structure from both feature level and topology level. Following the bottom-up paradigm, we propose a boosted Local Structured HOG-LBP based object detector. Firstly, at feature level, we propose Local Structured Descriptor to capture the object's local structure, and develop the descriptors from shape and texture information, respectively. Secondly, at topology level, we present a boosted feature selection and fusion scheme for part based object detector. All experiments are conducted on the challenging PASCAL VOC2007 datasets. Experimental results show that our method achieves the state-of-the-art performance. Junge Zhang, Kaiqi Huang, Yinan Yu, Tieniu Tan |
CVPR | 1 |
| 2011 | Robust view transformation model for gait recognitionabstractRecent gait recognition systems often suffer from the challenges including viewing angle variation and large intra-class variations. In order to address these challenges, this paper presents a robust View Transformation Model for gait recognition. Based on the gait energy image, the proposed method establishes a robust view transformation model via robust principal component analysis. Partial least square is used as feature selection method. Compared with the existing methods, the proposed method finds out a shared linear correlated low rank subspace, which brings the advantages that the view transformation model is robust to viewing angle variation, clothing and carrying condition changes. Conducted on the CASIA gait dataset, experimental results show that the proposed method outperforms the other existing methods. Shuai Zheng 0001, Junge Zhang, Kaiqi Huang, Ran He 0001, Tieniu Tan |
ICIP | 2 |