EDBT 2026 Demo / reviewers in the wild / expert
Liuyu Xiang
dblp:242/7959
· DBLP profile ↗
28ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0001-8486-6255ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 10 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CO-EVO: Co-evolving Semantic Anchoring and Style Diversification for Federated DG-ReIDabstractFederated domain generalization for person reidentification (FedDG-ReID) aims to collaboratively train a pedestrian retrieval model across multiple decentralized source domains such that it can generalize to unseen target environments without compromising raw data privacy.However, this task is significantly challenged by the inherent stylistic gaps across decentralized clients.Without global supervision, models easily succumb to shortcut learning where representations overfit to domain specific camera biases rather than universal identity features.We propose CO-EVO, a novel federated framework that resolves this semantic-style conflict through a co-evolutionary mechanism.On the semantic side, Camera-Invariant Semantic Anchoring (CSA) learns identity prompts with cross-camera consistency to establish purified and domain-agnostic anchors that filter out local imaging noise.On the visual side, Global Style Diversification (GSD), powered by a Global Camera-Style Bank (GCSB), synthesizes realistic perturbations to expand the visual boundaries of training data.The core of CO-EVO is its co-evolutionary loop where purified anchors act as gravitational centers to guide the image encoder toward robust anatomical attributes amidst diverse style variations.Extensive experiments demonstrate that CO-EVO achieves state-of-the-art (SOTA) performance, proving that the synergy between semantic purification and style expansion is essential for robust cross-domain generalization. Fengchun Zhang, Liuyu Xiang, Jinshan Lai, Tingxuan Huang |
ACL (1) | 3 |
| 2026 | Knowledge-guided policy arbitration: A hierarchical cognitive framework for safety-critical decision-making under dynamic conflicting objectives
Fuqing Bie, Xingyang Chang, Leyan Wang, Dehua Ma, Songfu Xu, Shuodi Liu, Yingzhuo Liu, Liuyu Xiang, Zhaofeng He 0001 |
Expert Syst. Appl. | 8 |
| 2026 | RainbowArena: A multi-agent toolkit for reinforcement learning and large language models in tabletop games
Yingzhuo Liu, Shuodi Liu, Hongsong Tang, Yubing Ma, Zikang Li, Junge Zhang, Liuyu Xiang, Zhaofeng He 0001 |
Knowl. Based Syst. | 7 |
| 2026 | Symmetric Image-Text Tuning With Entropy-Guided Fusion for Online Continual Learning in Non-Stationary Visual StreamsabstractOnline continual learning studies how models learn from continuous and non-stationary data streams. In this paper, we observe that CLIP models exhibit an asymmetric image-text interaction under online continual learning. Specifically, text features of previously seen classes may introduce unfavorable supervision when paired with visual features of newly observed data, leading to catastrophic forgetting. To alleviate this issue, we propose a simple yet effective symmetric image-text tuning (SIT) strategy that removes such asymmetric text supervision during online learning. We further introduce an entropy-guided fusion (EGF) mechanism that adaptively combines predictions from the pretrained and finetuned branches based on their relative uncertainty. This design allows the model to recover pretrained knowledge when the finetuned branch becomes unreliable, while still preserving plasticity on recently observed classes when confidence is high. In addition, we present MiD-Blurry, an online continual learning benchmark that combines multiple class distribution patterns to better reflect realistic data streams with blurred temporal boundaries. Extensive experiments on standard continual learning benchmarks and the MiD-Blurry setting evaluate inference-at-any-time performance and generalization to future data. The results show that the proposed approach maintains a practical balance between adapting to new data and preserving previously learned information in realistic online learning scenarios. Leyuan Wang, Liuyu Xiang, Yiwei Ru, Yunlong Wang 0003, Zhaofeng He 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | Rethinking Class-Incremental Learning From a Dynamic Imbalanced Learning PerspectiveabstractDeep neural networks suffer from catastrophic forgetting when continually learning new concepts. In this paper, we analyze this problem from a data imbalance point of view. We argue that the imbalance between old task and new task data contributes to forgetting of the old tasks. Moreover, the increasing imbalance ratio during incremental learning further aggravates the problem. To address the dynamic imbalance issue, we propose Uniform Prototype Contrastive Learning (UPCL), where uniform and compact features are learned. Specifically, we generate a set of non-learnable uniform prototypes before each task starts. Then we assign these uniform prototypes to each class and guide the feature learning through prototype contrastive learning. We also dynamically adjust the relative margin between old and new classes so that the feature distribution will be maintained balanced and compact. Finally, we demonstrate through extensive experiments that the proposed method achieves state-of-the-art performance on several benchmark including CIFAR-100, ImageNet-100, TinyImageNet, Food-101, and CUB-200. Experimental results show that our approach not only effectively addresses the issue of imbalanced old data in memory but also tackles the problem of imbalanced new data distributions. Leyuan Wang, Liuyu Xiang, Yunlong Wang 0003, Huijia Wu, Huafeng Yang, Jingqian Liu, Zhaofeng He 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language ModelsabstractLarge language models (LLMs) have demonstrated remarkable reasoning and planning capabilities, driving extensive research into task decomposition.Existing task decomposition methods focus primarily on memory, tool usage, and feedback mechanisms, achieving notable success in specific domains, but they often overlook the trade-off between performance and cost.In this study, we first conduct a comprehensive investigation on task decomposition, identifying six categorization schemes.Then, we perform an empirical analysis of three factors that influence the performance and cost of task decomposition: categories of approaches, characteristics of tasks, and configuration of decomposition and execution models, uncovering three critical insights and summarizing a set of practical principles.Building on this analysis, we propose the Select-Then-Decompose strategy, which establishes a closed-loop problemsolving process composed of three stages: selection, execution, and verification.This strategy dynamically selects the most suitable decomposition approach based on task characteristics and enhances the reliability of the results through a verification module.Comprehensive evaluations across multiple benchmarks show that the Select-Then-Decompose consistently lies on the Pareto frontier, demonstrating an optimal balance between performance and cost.Our code is publicly available at https://github.com/summervvind/ Select-Then-Decompose. Shuodi Liu, Yingzhuo Liu, Zi Wang 0014, Huijia Wu, Liuyu Xiang, Zhaofeng He 0001 |
EMNLP | 6 |
| 2025 | Adaptive Articulated Object Manipulation on the Fly with Foundation Model Reasoning and Part GroundingabstractArticulated objects pose diverse manipulation challenges for robots. Since their internal structures are not directly observable, robots must adaptively explore and refine actions to generate successful manipulation trajectories. While existing works have attempted cross-category generalization in adaptive articulated object manipulation, two major challenges persist: (1) the geometric diversity of real-world articulated objects complicates visual perception and understanding, and (2) variations in object functions and mechanisms hinder the development of a unified adaptive manipulation strategy. To address these challenges, we propose AdaRPG, a novel framework that leverages foundation models to extract object parts, which exhibit greater local geometric similarity than entire objects, thereby enhancing visual affordance generalization for functional primitive skills. To support this, we construct a part-level affordance annotation dataset to train the affordance model. Additionally, AdaRPG utilizes the common knowledge embedded in foundation models to reason about complex mechanisms and generate high-level control codes that invoke primitive skill functions based on part affordance inference. Simulation and real-world experiments demonstrate AdaRPG's strong generalization ability across novel articulated object categories. Yuanfei Wang, Ruihai Wu, Kunqi Xu, Yu Li 0022, Liuyu Xiang, Hao Dong 0003, Zhaofeng He 0001 |
ICCV | 6 |
| 2025 | FoodWeight1.4M: A Large-scale Multi-modal Dataset for Weight EstimationabstractLarge vision language models (VLMs) excel in visual tasks but struggle with weight estimation, hindering 3D perception and embodied intelligence. To address the lack of large-scale weight datasets, we present FoodWeight1.4M, derived from real-world supermarket scenarios. It contains 1.4 million high-quality images across 1,550 food categories, with weights precisely measured and rigorously filtered, making it the first large-scale weight estimation dataset. The weight estimation performance of current VLMs were tested and found to be unsatisfactory, which can be significantly improved by instruction tuning using Food-Weight1.4M. Moreover, we propose two strategies, Category-Guided and Reference Calibration, to enhance weight estimation without fine-tuning. Experiments confirm their effectiveness in improving multi-modal weight perception. Furthermore, experimental results show that pre-training on FoodWeight1.4M can benefit other food analysis tasks. Our dataset will be publicly available soon. Zhenbo Xu, Dehua Ma, Liuyu Xiang, Huijia Wu, Zhaofeng He 0001 |
ICME | 5 |
| 2025 | RainbowArena: A Multi-Agent Toolkit for Reinforcement Learning and Large Language Models in Competitive Tabletop Games
Yingzhuo Liu, Shuodi Liu, Hongsong Tang, Yubing Ma, Zikang Li, Junge Zhang, Liuyu Xiang, Zhaofeng He 0001 |
AAMAS | 7 |
| 2025 | GraphDiffusion: A Graph-conditioned Diffusion Model for Chip PlacementabstractPlacement is a crucial and time-intensive step in the modern Electronic Design Automation (EDA) chip design process, involving the allocation of millions of modules on a chip canvas. Previous studies have demonstrated the efficacy of machine learning methods, especially reinforcement learning (RL) in chip placement. However, existing RL-based methods still face challenges, including lengthy placement times due to placing only a single module at each step, and a lack of generalization capability. In response to these challenges, we propose a Graph Convolutional Network (GCN)-based diffusion model called GraphDiffusion. Our model presents a novel approach for chip placement by representing the positioning of movable nodes as a conditional denoising diffusion process. The GCN is utilized to extract feature information from the circuit, which then serves as an embedding to guide the diffusion model in generating the chip layout. We exploit the diffusion model’s capabilities to improve layout quality and generalizability. We conduct extensive experiments on the ISPD benchmark. The results demonstrate that our approach outperforms advanced placement methods while enhancing the transferability of the layout model. Upon training with expert data, GraphDiffusion can interpret the input circuit netlists and generate state-of-the-art chip layouts. Siyuan Fang, Liuyu Xiang, Wei Li 0032, Zhaofeng He 0001 |
ISCAS | 2 |
| 2025 | Can large language models independently complete tasks? A dynamic evaluation framework for multi-turn task planning and completion
Junlin Cui, Huijia Wu, Liuyu Xiang, Xiangang Li, Yaodong Yang 0001, Zhaofeng He 0001 |
Neurocomputing | 4 |
| 2025 | Generalizable agent modeling for agent collaboration-competition adaptation with multi-retrieval and dynamic generation
Yonggang Jin, Youpeng Zhao 0001, Zipeng Dai, Jian Zhao 0018, Liuyu Xiang, Junge Zhang, Zhaofeng He 0001 |
Neurocomputing | 7 |
| 2025 | DrlGoFPGA: FPGA Global Placement Considering Input-Output Buffer Based on Deep Reinforcement Learning and Gradient OptimizationabstractThe placement of the input-output buffer (IOBUF) can impact the performance and power consumption of the FPGA. The existing global placement (GP) methods lack consideration for IOBUF, resulting in a decrease in placement and routing quality. To address this issue, we propose a GP framework, DrlGoFPGA, which combines IOBUF placement based on deep reinforcement learning (DRL) with other instances placement based on gradient optimization (GO). A policy network structure with multi-action sampling is designed to accelerate the running speed of DRL, and a parallelizable reward function is designed to optimize each IOBUF placement action and avoid sparse reward problems. Then, an IOBUF line-network relationship (ILNR) graph creation method is designed to improve the agent’s ability to explore optimal solutions, and the graph features of ILNR by capturing them through a graph neural network embedded in the convolutional neural network. Finally, an IOBUF placement legalization method is designed to ensure that the IOBUF position meets the FPGA architecture. The experimental results show that compared with the state-of-the-art placement tools based on GO, DrlGoFPGA can improve GP speed by 13.2%-7×, half-perimeter wirelength by 0.2%-2.6%, and wirelength by 0.2%-1.5% and the IOBUF placement model has good generalization. Jianwang Zhai, Liuyu Xiang, Zixi Huang, Zhaofeng He 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Toward Realistic Hierarchical Object Detection: Problem, Benchmark, and SolutionabstractWith the continuous advancement of deep learning, object detection has made remarkable progress in accurately identifying a wide range of object categories, even within increasingly complex scenes. However, as the number of categories grows, visual concepts naturally organize into a label hierarchy. We contend that existing hierarchical classification and detection methods predominantly prioritize fine-grained prediction, potentially leading to inconsistencies with realistic human perception. From this perspective, we investigate the Hierarchical Object Detection (HOD) problem to better align with real-world perception. To address the lack of benchmarks in the field, we build a large-scale HOD benchmark termed RHOD with open-source datasets, comprising 740 categories. To better align the hierarchical object detectors towards realistic perception, we propose a new evaluation metric named Hierarchical Average Precision (HAP). Furthermore, we present a novel hierarchical object detection method that includes two components, Tree Soft Labeling (TSL) and Hierarchical Extension and Suppression (HES). Our method mitigates the issue of overconfidence in fine-grained predictions, which has been prevalent in previous approaches. We evaluate a range of existing methods on the RHOD benchmark, including plain, hierarchical, and open-vocabulary models. Additionally, we perform comprehensive experiments to assess the performance of our proposed method. The experimental results show that our method achieves state-of-the-art performance on the RHOD benchmark. Juexiao Feng, Yuhong Yang 0008, Mengyao Lyu, Tianxiang Hao 0001, Yi-Jie Huang, Yanchun Xie, Jungong Han, Liuyu Xiang, Guiguang Ding |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2025 | Toward Real-World Remote Sensing Image Super-Resolution: A New Benchmark and an Efficient ModelabstractSuper-resolution (SR) is a fundamental and crucial task in remote sensing. It can improve low-resolution (LR) remote sensing images and has potential benefits for downstream tasks such as remote sensing object detection and recognition. Existing remote sensing image SR (RSISR) methods are trained on simulated paired datasets, in which LR images are obtained by a simple and uniform (i.e., bicubic) degradation from corresponding high-resolution (HR) images. However, since this simulated degradation usually deviates from the real degradation, the performance of the trained model is limited when applied to real scenarios. To address this issue, we construct a novel real-world RSISR (RRSISR) dataset to model the real-world degradation, which exploits the imaging characteristics of the spectral camera to capture paired LR-HR images of the same scene. To ensure the precise alignment of the paired images, algorithms such as image registration and geometric correction are utilized. In addition, considering the vast amount of data involved in the RSISR task and its requirement for higher efficiency, we divide the image into patches with different restoration difficulties and propose a reference table-based patch exiting (RPE) method to efficiently reduce the computation of SR. Specifically, this method incorporates a predictor to estimate the performance of the current layer and a lookup table to decide whether to exit. Extensive experiments show that models trained on the proposed RRSISR dataset produce more realistic images than models with simulated datasets and generalize well to other satellites. We also demonstrate the efficiency of our RPE. Jia Wang 0038, Liuyu Xiang, Jiaochong Xu, Peipei Li 0002, Qizhi Xu, Zhaofeng He 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | FDNet: A Frequency-Aware Decomposition Network for Robust Face Super-Resolution Against Adversarial AttacksabstractFace super-resolution (FSR) is a crucial step in the face analysis pipeline, achieving remarkable progress by applying deep neural networks (DNNs). However, DNN-based FSR models are not robust enough and may suffer significant performance degradation due to subtle adversarial perturbations. In addition, the high-frequency details of images restored by existing models are insufficient, especially at large upsampling factors. In this paper, we propose a frequency-aware decomposition network (FD-Net) for robust face super-resolution, which aims to defend against adversarial attacks and obtain face images with fidelity. Observing that the noise introduced by adversarial attacks is often intricately mixed with the high-frequency information of the input image, we decompose and process the features of different frequencies separately to eliminate harmful perturbations and enhance high-frequency information. Specifically, by leveraging the frequency-aware capability of empirical mode decomposition (EMD), we propose an EMD-based multi-branch structure. The framework implicitly compels different branches to adaptively extract features from distinct frequency bands, limiting the adversarial noise into decoupled components restricted to specific branches. It also improves the recovery of high-frequency information, which is conducive to producing more credible results. Furthermore, we introduce a high-frequency noise suppressor capable of randomly eliminating imperceptible noise in the high-frequency components. Quantitative and qualitative results demonstrate the superior robustness of our proposed method against adversarial attacks, showing better fidelity in image reconstruction compared to state-of-the-art FSR methods, especially for upscaling factors of 8 and 16. Jia Wang 0038, Peipei Li 0002, Liuyu Xiang, Rui Wang 0124, Zhaofeng He 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Distributed Policy Space Response Oracles in Two-Player Zero-Sum GamesabstractPolicy space response oracle (PSRO) is a population-based algorithm that can be used to solve two-player zero-sum games. In the PSRO solution framework, optimizing policy diversity is crucial for addressing nontransitive game problems, helping the agent population avoid exploitation by unfamiliar opponents. In addition, while deep reinforcement learning is highly effective in solving complex game environments, its integration with PSRO remains fragmented and lacking in effective coordination. In this study, we propose distributed PSRO to efficiently solve complex game scenarios. To enhance diversity while managing optimization costs, we introduce TOP-K truncation, which prioritizes high-quality opponents and limits the size of the policy pool during sampling. This approach not only reduces interference from less effective strategies but also ensures computational efficiency by seamlessly integrating with our distributed training framework. We also design the distributed training framework to incorporate diversity estimation directly into the sampling process, achieving diversity optimization without incurring additional computational overhead. Furthermore, we introduce the opponent first (OF) method, which enhances decision-making by leveraging opponent information during interaction sampling. We perform experimental validation using a nontransitive mixture model and AlphaStar888 to confirm the effectiveness of the TOP-K truncation approach. Finally, we demonstrate the feasibility and efficiency of the distributed training framework and the OF approach in a Google Research Football 11 versus 11 scenario. Hongsong Tang, Yingzhuo Liu, Letian Ni, Liuyu Xiang, Yaodong Yang 0001, Ke Bi, Zhaofeng He 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Debiased Novel Category Discovering and LocalizationabstractIn recent years, object detection in deep learning has experienced rapid development. However, most existing object detection models perform well only on closed-set datasets, ignoring a large number of potential objects whose categories are not defined in the training set. These objects are often identified as background or incorrectly classified as pre-defined categories by the detectors. In this paper, we focus on the challenging problem of Novel Class Discovery and Localization (NCDL), aiming to train detectors that can detect the categories present in the training data, while also actively discover, localize, and cluster new categories. We analyze existing NCDL methods and identify the core issue: object detectors tend to be biased towards seen objects, and this leads to the neglect of unseen targets. To address this issue, we first propose an Debiased Region Mining (DRM) approach that combines class-agnostic Region Proposal Network (RPN) and class-aware RPN in a complementary manner. Additionally, we suggest to improve the representation network through semi-supervised contrastive learning by leveraging unlabeled data. Finally, we adopt a simple and efficient mini-batch K-means clustering method for novel class discovery. We conduct extensive experiments on the NCDL benchmark, and the results demonstrate that the proposed DRM approach significantly outperforms previous methods, establishing a new state-of-the-art. Juexiao Feng, Yuhong Yang 0008, Yanchun Xie, Yandong Guo, Liuyu Xiang, Guiguang Ding |
AAAI | 8 |
| 2024 | SCOMatch: Alleviating Overtrusting in Open-Set Semi-supervised Learning
Zerun Wang, Liuyu Xiang, Lang Huang 0001, Jiafeng Mao, Ling Xiao 0001, Toshihiko Yamasaki |
ECCV (51) | 2 |
| 2023 | Box-Level Active DetectionabstractActive learning selects informative samples for annotation within budget, which has proven efficient recently on object detection. However, the widely used active detection benchmarks conduct image-level evaluation, which is unrealistic in human workload estimation and biased towards crowded images. Furthermore, existing methods still perform image-level annotation, but equally scoring all targets within the same image incurs waste of budget and redundant labels. Having revealed above problems and limitations, we introduce a box-level active detection framework that controls a box-based budget per cycle, prioritizes informative targets and avoids redundancy for fair comparison and efficient application. Under the proposed box-level setting, we devise a novel pipeline, namely Complementary Pseudo Active Strategy (ComPAS). It exploits both human annotations and the model intelligence in a complementary fashion: an efficient input-end committee queries labels for informative objects only; meantime well-learned targets are identified by the model and compensated with pseudo-labels. ComPAS consistently outperforms 10 competitors under 4 settings in a unified codebase. With supervision from labeled data only, it achieves 100% supervised performance of VOC0712 with merely 19% box annotations. On the COCO dataset, it yields up to 4.3% mAP improvement over the second-best method. ComPAS also supports training with the unlabeled pool, where it surpasses 90% COCO supervised performance with 85% label reduction. Our source code is publicly available at https://github.com/lyumengyao/blad. Mengyao Lyu, Jundong Zhou, Hui Chen 0013, Dongdong Yu, Yandong Guo, Liuyu Xiang, Guiguang Ding |
CVPR | 9 |
| 2023 | Generative Iris Prior Embedded Transformer for Iris RestorationabstractIris restoration from complexly degraded iris images, aiming to improve iris recognition performance, is a challenging problem. Due to the complex degradation, directly training a convolutional neural network (CNN) without prior cannot yield satisfactory results. In this work, we propose a generative iris prior embedded Transformer model (Gformer), in which we build a hierarchical encoder-decoder network employing Transformer block and generative iris prior. First, we tame Transformer blocks to model long-range dependencies in target images. Second, we pretrain an iris generative adversarial network (GAN) to obtain the rich iris prior, and incorporate it into the iris restoration process with our iris feature modulator. Our experiments demonstrate that the proposed Gformer outperforms state-of-the-art methods. Besides, iris recognition performance has been significantly improved after applying Gformer. Jia Wang 0038, Peipei Li 0002, Liuyu Xiang, Peigang Li, Zhaofeng He 0001 |
ICME | 4 |
| 2023 | Margin-aware rectified augmentation for long-tailed recognition
Liuyu Xiang, Jungong Han, Guiguang Ding |
Pattern Recognit. | 1 |
| 2022 | ReMoNet: Recurrent Multi-Output Network for Efficient Video DenoisingabstractWhile deep neural network-based video denoising methods have achieved promising results, it is still hard to deploy them on mobile devices due to their high computational cost and memory demands. This paper aims to develop a lightweight deep video denoising method that is friendly to resource-constrained mobile devices. Inspired by the facts that 1) consecutive video frames usually contain redundant temporal coherency, and 2) neural networks are usually over-parameterized, we propose a multi-input multi-output (MIMO) paradigm to process consecutive video frames within one-forward-pass. The basic idea is concretized to a novel architecture termed Recurrent Multi-output Network (ReMoNet), which consists of recurrent temporal fusion and temporal aggregation blocks and is further reinforced by similarity-based mutual distillation. We conduct extensive experiments on NVIDIA GPU and Qualcomm Snapdragon 888 mobile platform with Gaussian noise and simulated Image-Signal-Processor (ISP) noise. The experimental results show that ReMoNet is both effective and efficient on video denoising. Moreover, we show that ReMoNet is more robust under higher noise level scenarios. Liuyu Xiang, Jundong Zhou, Jirui Liu, Zerun Wang, Haidong Huang, Jie Hu 0021, Jungong Han, Guiguang Ding |
AAAI | 1 |
| 2022 | Long-tailed visual recognition with deep models: A methodological survey and evaluation
Yu Fu 0006, Liuyu Xiang, Yumna Zahid, Guiguang Ding, Tao Mei 0001, Qiang Shen 0001, Jungong Han |
Neurocomputing | 2 |
| 2020 | PANDA: A Gigapixel-Level Human-Centric Video DatasetabstractWe present PANDA, the first gigaPixel-level humAN-centric viDeo dAtaset, for large-scale, long-term, and multi-object visual analysis. The videos in PANDA were captured by a gigapixel camera and cover real-world scenes with both wide field-of-view (~1 square kilometer area) and high-resolution details (~gigapixel-level/frame). The scenes may contain 4k head counts with over 100× scale variation. PANDA provides enriched and hierarchical ground-truth annotations, including 15,974.6k bounding boxes, 111.8k fine-grained attribute labels, 12.7k trajectories, 2.2k groups and 2.9k interactions. We benchmark the human detection and tracking tasks. Due to the vast variance of pedestrian pose, scale, occlusion and trajectory, existing approaches are challenged by both accuracy and efficiency. Given the uniqueness of PANDA with both wide FoV and high resolution, a new task of interaction-aware group detection is introduced. We design a `global-to-local zoom-in' framework, where global trajectories and local interactions are simultaneously encoded, yielding promising results. We believe PANDA will contribute to the community of artificial intelligence and praxeology by understanding human behaviors and interactions in large-scale real-world scenes. PANDA Website: http://www.panda-dataset.com. Xiya Zhang, Yinheng Zhu, Xiaoyun Yuan, Liuyu Xiang, Zerun Wang, Guiguang Ding, David J. Brady, Qionghai Dai, Lu Fang 0001 |
CVPR | 6 |
| 2020 | Learning From Multiple Experts: Self-paced Knowledge Distillation for Long-Tailed Classification
Liuyu Xiang, Guiguang Ding, Jungong Han |
ECCV (5) | 1 |
| 2019 | Adaptive Region Embedding for Text ClassificationabstractDeep learning models such as convolutional neural networks and recurrent networks are widely applied in text classification. In spite of their great success, most deep learning models neglect the importance of modeling context information, which is crucial to understanding texts. In this work, we propose the Adaptive Region Embedding to learn context representation to improve text classification. Specifically, a metanetwork is learned to generate a context matrix for each region, and each word interacts with its corresponding context matrix to produce the regional representation for further classification. Compared to previous models that are designed to capture context information, our model contains less parameters and is more flexible. We extensively evaluate our method on 8 benchmark datasets for text classification. The experimental results prove that our method achieves state-of-the-art performances and effectively avoids word ambiguity. Liuyu Xiang, Xiaoming Jin, Lan Yi, Guiguang Ding |
AAAI | 1 |
| 2019 | Incremental Few-Shot Learning for Pedestrian Attribute RecognitionabstractPedestrian attribute recognition has received increasing attention due to its important role in video surveillance applications. However, most existing methods are designed for a fixed set of attributes. They are unable to handle the incremental few-shot learning scenario, i.e. adapting a well-trained model to newly added attributes with scarce data, which commonly exists in the real world. In this work, we present a meta learning based method to address this issue. The core of our framework is a meta architecture capable of disentangling multiple attribute information and generalizing rapidly to new coming attributes. By conducting extensive experiments on the benchmark dataset PETA and RAP under the incremental few-shot setting, we show that our method is able to perform the task with competitive performances and low resource requirements. Liuyu Xiang, Xiaoming Jin, Guiguang Ding, Jungong Han, Leida Li |
IJCAI | 1 |