EDBT 2026 Demo / reviewers in the wild / expert
Caixia Yan
dblp:32/9964
· DBLP profile ↗
32ranked-venue papers
9as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SynHOI: Multi-granularity GAN synthesizer for generative zero-shot HOI detection
Caixia Yan, Yan Kou |
Neurocomputing | 1 |
| 2026 | VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection
Caixia Yan, Muyan Jiao, Nuohan Xue, Weizhan Zhang, Jiahao Wang 0004, Xiaojun Chang, Feng Tian 0002 |
Neural Networks | 1 |
| 2026 | CollabVisAdapt: Spatio-Temporal Context-Aware Adaptation of Shared Object Visualization for MR TelecollaborationabstractMixed Reality (MR) telecollaboration aims to enable users to share local objects as real-time synchronized virtual replicas to remote partners and collaborate on physical tasks as if they were co-located. However, in everyday scenarios with mobile and easy-to-setup MR devices, visualizing shared objects in a single modality, ranging from 2D images to 3D reconstruction, struggles to simultaneously optimize all the aspects of Spatiality, Fidelity, and Real-time performance. To overcome this issue, existing methods explore integrating multiple visualization modalities to leverage their respective advantages in subsets of the three aspects. However, they focus on fixed modality combinations without considering user-centered task contexts and workflow, where users may prioritize different aspects of the visualization across task phases. Moreover, they lack support for switching or require manual switching across modalities, which could become disruptive and tiring. In this paper, we propose adapting object visualization based on spatiotemporal contexts in telecollaboration. Specifically, we first couple task type with the user's relative viewing distance as the spatial context, and examine its impact on users' prioritized visualization aspects, and the corresponding switching thresholds. With differing generation speeds of modalities, we then explore temporal switching schemes when the preferred modality is not immediately available. With the obtained design choices, we implement CollabVisAdapt, a proof-of-concept prototype that supports automatic adaptation of object visualization based on spatiotemporal contexts in MR telecollaboration. A user study in remote maintenance verifies the effectiveness of the proposed workflow with adaptive visualization and the usability of the system. Xuanyu Wang 0001, Weizhan Zhang, Shuaichen Guo, Caixia Yan, Shuming Yang, Haipeng Du, Wangdu Chen, Qi Wang 0180 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | SpotActor: Training-Free Layout-Controlled Consistent Image GenerationabstractText-to-image diffusion models significantly enhance the efficiency of artistic creation with high-fidelity image generation. However, in typical application scenarios like comic book production, they can neither place each subject into its expected spot nor maintain the consistent appearance of each subject across images. For these issues, we pioneer a novel task, Layout-to-Consistent-Image (L2CI) generation, which produces consistent and compositional images in accordance with the given layout conditions and text prompts. To accomplish this challenging task, we present a new formalization of dual energy guidance with optimization in a dual semantic-latent space and thus propose a training-free pipeline, SpotActor, which features a layout-conditioned optimizing stage and a consistent sampling stage. In the optimizing stage, we innovate a nuanced layout energy function to mimic the attention activations with a sigmoid-like objective. While in the sampling stage, we design Regional Interconnection Self-Attention (RISA) and Semantic Fusion Cross-Attention (SFCA) mechanisms that allow mutual interactions across images. To evaluate the performance, we present ActorBench, a specified benchmark with hundreds of reasonable prompt-box pairs stemming from object detection datasets. Comprehensive experiments are conducted to demonstrate the effectiveness of our method. The results prove that SpotActor fulfills the expectations of this task and showcases the potential for practical applications with superior layout alignment, subject consistency, prompt conformity and background diversity. Jiahao Wang 0004, Caixia Yan, Weizhan Zhang, Haonan Lin, Mengmeng Wang 0005, Guang Dai, Tieliang Gong, Hao Sun 0015, Jingdong Wang 0001 |
AAAI | 2 |
| 2025 | Efficient Real-Time On-Mobile Video Super-Resolution with Automatic Evolutionary Neural Architecture Search
Xuncheng Liu, Weizhan Zhang, Caixia Yan, Haipeng Du |
ICANN (2) | 3 |
| 2025 | EchoShot: Multi-Shot Portrait Video GenerationabstractVideo diffusion models substantially boost the productivity of artistic workflows with high-quality portrait video generative capacity. However, prevailing pipelines are primarily constrained to single-shot creation, while real-world applications urge multiple shots with identity consistency and flexible content controllability. In this work, we propose EchoShot, a native and scalable multi-shot framework for portrait customization built upon a foundation video diffusion model. To start with, we propose shot-aware position embedding mechanisms within the video diffusion transformer architecture to model inter-shot variations and establish intricate correspondence between multi-shot visual content and their textual descriptions. This simple yet effective design enables direct training on multi-shot video data without introducing additional computational overhead. To facilitate model training within multi-shot scenarios, we construct PortraitGala, a large-scale and high-fidelity human-centric video dataset featuring cross-shot identity consistency and fine-grained captions such as facial attributes, outfits, and dynamic motions. To further enhance applicability, we extend EchoShot to perform reference image-based personalized multi-shot generation and long video synthesis with infinite shot counts. Extensive evaluations demonstrate that EchoShot achieves superior identity consistency as well as attribute-level controllability in multi-shot portrait video generation. Notably, the proposed framework demonstrates potential as a foundational paradigm for general multi-shot video modeling. Project page: https://johnneywang.github.io/EchoShot-webpage. Jiahao Wang 0004, Hualian Sheng, Sijia Cai, Weizhan Zhang, Caixia Yan, Yachuang Feng, Bing Deng, Jieping Ye |
NeurIPS | 5 |
| 2025 | Leveraging differentiable NAS and abstract genetic algorithms for optimizing on-mobile VSR performance
Xuncheng Liu, Weizhan Zhang, Tieliang Gong, Caixia Yan |
Mach. Learn. | 4 |
| 2025 | Lightweight Configuration Adaptation With Multi-Teacher Reinforcement Learning for Live Video AnalyticsabstractThe proliferation of video data and advancements in Deep Neural Networks (DNNs) have greatly boosted live video analytics, driven by the growing video capture capabilities of mobile devices. However, resource limitations necessitate the transmission of endpoint-collected videos to servers for inference. To meet real-time requirements and ensure accurate inference, it is essential to adjust video configurations at the endpoint. Traditional methods rely on deterministic strategies, posing difficulties in adapting to dynamic networks and video content. Meanwhile, emerging learning-based schemes suffer from trial-and-error exploration mechanisms, resulting in a concerning long-tail effect on upload latency. In this paper, we propose a novel lightweight and robust configuration adaptation policy (LCA), which fuses heuristic and RL-based agents using multi-teacher knowledge distillation (MKD) theory. Firstly, we propose a content-sensitive and bandwidth-adaptive RL agent and introduce a Lyapunov-based optimization agent for ensuring latency robustness. To leverage both agents' strengths, we design a feature-guided multi-teacher distillation network to transfer their advantages to the student. The experimental results across two vision tasks (pose estimation and semantic segmentation) demonstrate that LCA significantly reduces transmission latency compared to prior work (average reduction of 47.11%-89.55%, 95-percentile reduction of 27.63%-88.78%) and computational overhead while maintaining comparable inference accuracy. Yuanhong Zhang, Weizhan Zhang, Muyao Yuan, Caixia Yan, Tieliang Gong, Haipeng Du |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | SAUI: Scale-Aware Unseen Imagineer for Zero-Shot Object DetectionabstractZero-shot object detection (ZSD) aims to localize and classify unseen objects without access to their training annotations. As a prevailing solution to ZSD, generation-based methods synthesize unseen visual features by taking seen features as reference and class semantic embeddings as guideline. Although previous works continuously improve the synthesis quality, they fail to consider the scale-varying nature of unseen objects. The generation process is preformed over a single scale of object features and thus lacks scale-diversity among synthesized features. In this paper, we reveal the scale-varying challenge in ZSD and propose a Scale-Aware Unseen Imagineer (SAUI) to lead the way of a novel scale-aware ZSD paradigm. To obtain multi-scale features of seen-class objects, we design a specialized coarse-to-fine extractor to capture features through multiple scale-views. To generate unseen features scale by scale, we innovate a Series-GAN synthesizer along with three scale-aware contrastive components to imagine separable, diverse and robust scale-wise unseen features. Extensive experiments on PASCAL VOC, COCO and DIOR datasets demonstrate SAUI's better performance in different scenarios, especially for scale-varying and small objects. Notably, SAUI achieves the new state-of-the art performance on COCO and DIOR. Jiahao Wang 0004, Caixia Yan, Weizhan Zhang, Hao Sun 0015 |
AAAI | 2 |
| 2024 | Masked Distillation Advances Self-Supervised Transformer Architecture SearchabstractTransformer architecture search (TAS) has achieved remarkable progress in automating the neural architecture design process of vision transformers. Recent TAS advancements have discovered outstanding transformer architectures while saving tremendous labor from human experts. However, it is still cumbersome to deploy these methods in real-world applications due to the expensive costs of data labeling under the supervised learning paradigm. To this end, this paper proposes a masked image modelling (MIM) based self-supervised neural architecture search method specifically designed for vision transformers, termed as MaskTAS, which completely avoids the expensive costs of data labeling inherited from supervised learning. Based on the one-shot NAS framework, MaskTAS requires to train various weight-sharing subnets, which can easily diverged without strong supervision in MIM-based self-supervised learning. For this issue, we design the search space of MaskTAS as a siamesed teacher-student architecture to distill knowledge from pre-trained networks, allowing for efficient training of the transformer supernet. To achieve self-supervised transformer architecture search, we further design a novel unsupervised evaluation metric for the evolutionary search algorithm, where each candidate of the student branch is rated by measuring its consistency with the larger teacher network. Extensive experiments demonstrate that the searched architectures can achieve state-of-the-art accuracy on CIFAR-10, CIFAR-100, and ImageNet datasets even without using manual labels. Moreover, the proposed MaskTAS can generalize well to various data domains and tasks by searching specialized transformer architectures in self-supervised manner. Caixia Yan, Xiaojun Chang, Zhihui Li 0001, Lina Yao 0001, Minnan Luo |
ICLR | 1 |
| 2024 | OneActor: Consistent Subject Generation via Cluster-Conditioned GuidanceabstractText-to-image diffusion models benefit artists with high-quality image generation. Yet their stochastic nature hinders artists from creating consistent images of the same subject. Existing methods try to tackle this challenge and generate consistent content in various ways. However, they either depend on external restricted data or require expensive tuning of the diffusion model. For this issue, we propose a novel one-shot tuning paradigm, termed OneActor. It efficiently performs consistent subject generation solely driven by prompts via a learned semantic guidance to bypass the laborious backbone tuning. We lead the way to formalize the objective of consistent subject generation from a clustering perspective, and thus design a cluster-conditioned model. To mitigate the overfitting challenge shared by one-shot tuning pipelines, we augment the tuning with auxiliary samples and devise two inference strategies: semantic interpolation and cluster guidance. These techniques are later verified to significantly improve the generation quality. Comprehensive experiments show that our method outperforms a variety of baselines with satisfactory subject consistency, superior prompt conformity as well as high image quality. Our method is capable of multi-subject generation and compatible with popular diffusion extensions. Besides, we achieve a $4\times$ faster tuning speed than tuning-based baselines and, if desired, avoid increasing the inference time. Furthermore, our method can be naturally utilized to pre-train a consistent subject generation network from scratch, which will implement this research task into more practical applications. (Project page: https://johnneywang.github.io/OneActor-webpage/) Jiahao Wang 0004, Caixia Yan, Haonan Lin, Weizhan Zhang, Mengmeng Wang 0005, Tieliang Gong, Guang Dai, Hao Sun 0015 |
NeurIPS | 2 |
| 2024 | Multiple GRAphs-oriented Random wAlk (MulGRA2) for social link prediction
Tianliang Qi, Weihua Ji, Kuo-Ming Chao, Yan Chen 0031, Caixia Yan, Jun Liu 0002, Mo Xu, Zhihai Suo, Feng Tian 0002 |
Inf. Sci. | 7 |
| 2024 | Adaptive token selection for efficient detection transformer with dual teacher supervision
Muyao Yuan, Weizhan Zhang, Caixia Yan, Tieliang Gong, Yuanhong Zhang, Jiangyong Ying |
Knowl. Based Syst. | 3 |
| 2024 | Towards performance-maximizing neural network pruning via global channel attention
Yingchun Wang 0001, Song Guo 0001, Jingcai Guo, Jie Zhang 0076, Weizhan Zhang, Caixia Yan, Yuanhong Zhang |
Neural Networks | 6 |
| 2024 | Semantics-Guided Contrastive Network for Zero-Shot Object DetectionabstractZero-shot object detection (ZSD), the task that extends conventional detection models to detecting objects from unseen categories, has emerged as a new challenge in computer vision. Most existing approaches tackle the ZSD task with a strict mapping-transfer strategy that may lead to suboptimal ZSD results: 1) the learning process of these models neglects the available semantic information on unseen classes, which can easily bias towards the seen categories; 2) the original visual feature space is not well-structured for the ZSD task due to the lack of discriminative information. To address these issues, we develop a novel Semantics-Guided Contrastive Network for ZSD, named ContrastZSD, a detection framework that first brings contrastive learning mechanism into the realm of zero-shot detection. Particularly, ContrastZSD incorporates two semantics-guided contrastive learning subnets that contrast between region-category and region-region pairs respectively. The pairwise contrastive tasks take advantage of supervision signals derived from both the ground truth label and class similarity information. By performing supervised contrastive learning over those explicit semantic supervision, the model can learn more knowledge about unseen categories to avoid the bias problem to seen concepts, while optimizing the visual data structure to be more discriminative for better visual-semantic alignment. Extensive experiments are conducted on two popular benchmarks for ZSD, i.e., PASCAL VOC and MS COCO. Results show that our method outperforms the previous state-of-the-art on both ZSD and generalized ZSD tasks. Caixia Yan, Xiaojun Chang, Minnan Luo, Huan Liu 0012, Xiaoqin Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Disentangled Generation With Information Bottleneck for Enhanced Few-Shot LearningabstractFew-shot learning (FSL) poses a significant challenge in classifying unseen classes with limited samples, primarily stemming from the scarcity of data. Although numerous generative approaches have been investigated for FSL, their generation process often results in entangled outputs, exacerbating the distribution shift inherent in FSL. Consequently, this considerably hampers the overall quality of the generated samples. Addressing this concern, we present a pioneering framework called DisGenIB, which leverages an Information Bottleneck (IB) approach for Disentangled Generation. Our framework ensures both discrimination and diversity in the generated samples, simultaneously. Specifically, we introduce a groundbreaking Information Theoretic objective that unifies disentangled representation learning and sample generation within a novel framework. In contrast to previous IB-based methods that struggle to leverage priors, our proposed DisGenIB effectively incorporates priors as invariant domain knowledge of sub-features, thereby enhancing disentanglement. This innovative approach enables us to exploit priors to their full potential and facilitates the overall disentanglement process. Moreover, we establish the theoretical foundation that reveals certain prior generative and disentanglement methods as special instances of our DisGenIB, underscoring the versatility of our proposed framework. To solidify our claims, we conduct comprehensive experiments on demanding FSL benchmarks, affirming the remarkable efficacy and superiority of DisGenIB. Furthermore, the validity of our theoretical analyses is substantiated by the experimental results. Our code is available at https://github.com/eric-hang/DisGenIB. Zhuohang Dang, Minnan Luo, Jihong Wang 0003, Chengyou Jia, Caixia Yan, Guang Dai, Xiaojun Chang |
IEEE Trans. Image Process. | 5 |
| 2024 | FHVAC: Feature-Level Hybrid Video Adaptive Configuration for Machine-Centric Live StreamingabstractWith the widespread deployment of edge computing, the focus has shifted to machine-centric live video streaming, where endpoint-collected videos are transmitted over networks to edge servers for analysis. Unlike maximizing user's Quality of Experience (QoE), machine-centric video streaming optimizes the machine's Quality of Inference (QoI) by balancing the inference accuracy, inference delay, and transmission latency with video adaptive configuration. Traditional heuristic configuration adaption methods are reliable but unable to respond to erratic network fluctuations. Reinforcement learning (RL) based algorithms exhibit superior flexibility but suffer from exploration mechanisms, resulting in long-tail effects on upload latency. In this paper, we propose FHVAC, which dynamically selects video encoding parameters for live streaming by coherently fusing rule-based and RL-based agent at the feature level. We initially develop a robust rule-based approach for ensuring the low latency in transmission, and employ imitation learning to convert it into a neural network equivalently. Subsequently, we design a novel module to combine the two approaches and assess various fusion mechanisms. Our evaluation of FHVAC across two vision tasks (pose estimation and semantic segmentation) in two scenarios (trace-driven simulation and testbed-based experiment) shows that FHVAC enhances the average QoI, and reduces 10.61%-65.27% latency tail performance compared to prior work. Yuanhong Zhang, Weizhan Zhang, Haipeng Du, Caixia Yan, Li Liu 0036 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Spatially-Aware Human-Object Interaction Detection with Cross-Modal Enhancement
Gaowen Liu, Huan Liu 0012, Caixia Yan, Rui Li 0073, Sizhe Dang |
ICONIP (5) | 3 |
| 2023 | Tile Classification Based Viewport Prediction with Multi-modal Fusion TransformerabstractViewport prediction is a crucial aspect of tile-based 360° video streaming system. However, existing trajectory based methods lack of robustness, also oversimplify the process of information construction and fusion between different modality inputs, leading to the error accumulation problem. In this paper, we propose a tile classification based viewport prediction method with Multi-modal Fusion Transformer, namely MFTR. Specifically, MFTR utilizes transformer-based networks to extract the long-range dependencies within each modality, then mine intra- and inter-modality relations to capture the combined impact of user historical inputs and video contents on future viewport selection. In addition, MFTR categorizes future tiles into two categories: user interested or not, and selects future viewport as the region that contains most user interested tiles. Comparing with predicting head trajectories, choosing future viewport based on tile's binary classification results exhibits better robustness and interpretability. To evaluate our proposed MFTR, we conduct extensive experiments on two widely used PVS-HM and Xu-Gaze dataset. MFTR shows superior performance over state-of-the-art methods in terms of average prediction accuracy and overlap ratio, also presents competitive computation efficiency. Weizhan Zhang, Caixia Yan, Qi Wang 0180, Wangdu Chen |
ACM Multimedia | 4 |
| 2023 | Social Image-Text Sentiment Classification With Cross-Modal Consistency and Knowledge DistillationabstractSocial media sentiment analysis, which aims to evaluate the attitudes of online users based on their posts, has attracted significant research attention due to its successful application in the field of social media monitoring. It is a beneficial way to utilize multimodal information uploaded by users in order to improve sentiment classification ability. However, existing multimodal fusion-based approaches continue to face difficulties due to the issues of between-modality semantic inconsistency and missing modality. To address these issues, we propose a cross-modal consistency modeling-based knowledge distillation framework for image–text sentiment classification of social media data. Specifically, we design a hybrid curriculum learning strategy to measure the semantic consistency of multimodal data, then gradually train all image–text pairs from easy to hard, which can effectively handle the massive amounts of noise caused by inconsistencies between image and text data on social media. Moreover, in order to alleviate the problem of missing images in unimodal posts, we propose a privileged feature distillation method, in which the teacher model additionally considers images as privileged features, to transfer the visual knowledge to the student model, thereby enhancing the accuracy for text sentiment classification. Extensive experiments conducted over three real-world social media datasets demonstrate the effectiveness and superiority of the proposed multimodal sentiment analysis model. Huan Liu 0012, Jianping Fan 0007, Caixia Yan, Tao Qin 0002 |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | Counterfactual Generation Framework for Few-Shot LearningabstractFew-shot learning (FSL) that aims to recognize novel classes with few labeled samples is troubled by its data scarcity. Though recent works tackle FSL with data augmentation-based methods, these models fail to maintain the discrimination and diversity of the generated samples due to the distribution shift and intra-class bias caused by the data scarcity, therefore greatly undermining the performance. To this end, we use causal mechanisms, which are constant among independent variables across data distribution, to alleviate such effects. In this sense, we decompose the image information into two independent components: sample-specific and class-agnostic information, and further propose a novel Counterfactual Generation Framework (CGF) to learn the underlying causal mechanisms to synthesize faithful samples for FSL. Specifically, based on the counterfactual inference, we design a class-agnostic feature extractor to capture the sample-specific information, together with a counterfactual generation network to simulate the data generation process from a causal perspective. Moreover, to leverage the power of CGF in counterfactual inference, we further develop a novel classifier that classifies samples based on their distributions of counterfactual generations. Extensive experiments demonstrate the effectiveness of CGF on four FSL benchmarks, e.g., 80.12/86.13% accuracy on 5-way 1/5-shot miniImageNet FSL tasks, significantly improving the performance. Our codes and models are available athttps://github.com/eric-hang/CGF. Zhuohang Dang, Minnan Luo, Chengyou Jia, Caixia Yan, Xiaojun Chang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Collaborative Contrastive Refining for Weakly Supervised Person SearchabstractWeakly supervised person search involves training a model with only bounding box annotations, without human-annotated identities. Clustering algorithms are commonly used to assign pseudo-labels to facilitate this task. However, inaccurate pseudo-labels and imbalanced identity distributions can result in severe label and sample noise. In this work, we propose a novel Collaborative Contrastive Refining (CCR) weakly-supervised framework for person search that jointly refines pseudo-labels and the sample-learning process with different contrastive strategies. Specifically, we adopt a hybrid contrastive strategy that leverages both visual and context clues to refine pseudo-labels, and leverage the sample-mining and noise-contrastive strategy to reduce the negative impact of imbalanced distributions by distinguishing positive samples and noise samples. Our method brings two main advantages: 1) it facilitates better clustering results for refining pseudo-labels by exploring the hybrid similarity; 2) it is better at distinguishing query samples and noise samples for refining the sample-learning process. Extensive experiments demonstrate the superiority of our approach over the state-of-the-art weakly supervised methods by a large margin (more than 3% mAP on CUHK-SYSU). Moreover, by leveraging more diverse unlabeled data, our method achieves comparable or even better performance than the state-of-the-art supervised methods. Chengyou Jia, Minnan Luo, Caixia Yan, Linchao Zhu, Xiaojun Chang |
IEEE Trans. Image Process. | 3 |
| 2022 | ZeroNAS: Differentiable Generative Adversarial Networks Search for Zero-Shot LearningabstractIn recent years, remarkable progress in zero-shot learning (ZSL) has been achieved by generative adversarial networks (GAN). To compensate for the lack of training samples in ZSL, a surge of GAN architectures have been developed by human experts through trial-and-error testing. Despite their efficacy, however, there is still no guarantee that these hand-crafted models can consistently achieve good performance across diversified datasets or scenarios. Accordingly, in this paper, we turn to neural architecture search (NAS) and make the first attempt to bring NAS techniques into the ZSL realm. Specifically, we propose a differentiable GAN architecture search method over a specifically designed search space for zero-shot learning, referred to as ZeroNAS. Considering the relevance and balance of the generator and discriminator, ZeroNAS jointly searches their architectures in a min-max player game via adversarial training. Extensive experiments conducted on four widely used benchmark datasets demonstrate that ZeroNAS is capable of discovering desirable architectures that perform favorably against state-of-the-art ZSL and generalized zero-shot learning (GZSL) approaches. Source code is at https://github.com/caixiay/ZeroNAS. Caixia Yan, Xiaojun Chang, Zhihui Li 0001, Weili Guan, ZongYuan Ge, Lei Zhu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Self-weighted Robust LDA for Multiclass Classification with Edge ClassesabstractLinear discriminant analysis (LDA) is a popular technique to learn the most discriminative features for multi-class classification. A vast majority of existing LDA algorithms are prone to be dominated by the class with very large deviation from the others, i.e., edge class, which occurs frequently in multi-class classification. First, the existence of edge classes often makes the total mean biased in the calculation of between-class scatter matrix. Second, the exploitation of ℓ2-norm based between-class distance criterion magnifies the extremely large distance corresponding to edge class. In this regard, a novel self-weighted robust LDA with ℓ2,1-norm based pairwise between-class distance criterion, called SWRLDA, is proposed for multi-class classification especially with edge classes. SWRLDA can automatically avoid the optimal mean calculation and simultaneously learn adaptive weights for each class pair without setting any additional parameter. An efficient re-weighted algorithm is exploited to derive the global optimum of the challenging ℓ2,1-norm maximization problem. The proposed SWRLDA is easy to implement and converges fast in practice. Extensive experiments demonstrate that SWRLDA performs favorably against other compared methods on both synthetic and real-world datasets while presenting superior computational efficiency in comparison with other techniques. Caixia Yan, Xiaojun Chang, Minnan Luo, Xiaoqin Zhang 0002, Zhihui Li 0001, Feiping Nie 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2020 | Cross-Graph Representation Learning for Unsupervised Graph Alignment
Weifan Wang 0004, Minnan Luo, Caixia Yan, Meng Wang 0009, Xiang Zhao 0002 |
DASFAA (2) | 3 |
| 2020 | Unsupervised Hierarchical Feature Selection on Networked Data
Yuzhe Zhang 0003, Chen Chen 0022, Minnan Luo, Jundong Li, Caixia Yan |
DASFAA (3) | 5 |
| 2020 | Memory transformation networks for weakly supervised visual classification
Huan Liu 0012, Minnan Luo, Xiaojun Chang, Caixia Yan, Lina Yao 0001 |
Knowl. Based Syst. | 5 |
| 2020 | Semantics-Preserving Graph Propagation for Zero-Shot Object DetectionabstractMost existing object detection models are restricted to detecting objects from previously seen categories, an approach that tends to become infeasible for rare or novel concepts. Accordingly, in this paper, we explore object detection in the context of zero-shot learning, i.e., Zero-Shot Object Detection (ZSD), to concurrently recognize and localize objects from novel concepts. Existing ZSD algorithms are typically based on a simple mapping-transfer strategy that is susceptible to the domain shift problem. To resolve this problem, we propose a novel Semantics-Preserving Graph Propagation model for ZSD based on Graph Convolutional Networks (GCN). More specifically, we employ a graph construction module to flexibly build category graphs by incorporating diverse correlations between category nodes; this is followed by two semantics preserving modules that enhance both category and region representations through a multi-step graph propagation process. Compared to existing mapping-transfer based methods, both the semantic description and semantic structural knowledge exhibited in prior category graphs can be effectively leveraged to boost the generalization capability of the learned projection function via knowledge transfer, thereby providing a solution to the domain shift problem. Experiments on existing seen/unseen splits of three popular object detection datasets demonstrate that the proposed approach performs favorably against state-of-the-art ZSD methods. Caixia Yan, Xiaojun Chang, Minnan Luo, Chung-Hsing Yeh, Alex Hauptmann 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Discrete Multi-Graph ClusteringabstractSpectral clustering plays a significant role in applications that rely on multi-view data due to its well-defined mathematical framework and excellent performance on arbitrarily-shaped clusters. Unfortunately, directly optimizing the spectral clustering inevitably results in an NP-hard problem due to the discrete constraints on the clustering labels. Hence, conventional approaches intuitively include a relax-and-discretize strategy to approximate the original solution. However, there are no principles in this strategy that prevent the possibility of information loss between each stage of the process. This uncertainty is aggravated when a procedure of heterogeneous features fusion has to be included in multi-view spectral clustering. In this paper, we avoid an NP-hard optimization problem and develop a general framework for multi-view discrete graph clustering by directly learning a consensus partition across multiple views, instead of using the relax-and-discretize strategy. An effective re-weighting optimization algorithm is exploited to solve the proposed challenging problem. Further, we provide a theoretical analysis of the model's convergence properties and computational complexity for the proposed algorithm. Extensive experiments on several benchmark datasets verify the effectiveness and superiority of the proposed algorithm on clustering and image segmentation tasks. Minnan Luo, Caixia Yan, Xiaojun Chang, Ling Chen 0006, Feiping Nie 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Top-k multi-class SVM using multiple features
Caixia Yan, Minnan Luo, Huan Liu 0012, Zhihui Li 0001 |
Inf. Sci. | 1 |
| 2018 | Robust dictionary learning with graph regularization for unsupervised person re-identification
Caixia Yan, Minnan Luo, Wenhe Liu |
Multim. Tools Appl. | 1 |
| 2011 | Design, analysis and experiments of an omni-directional spherical robotabstractThis paper presents a novel design of an omni-directional spherical robot that is mainly composed of a lucent ball-shaped shell and an internal driving unit. Two motors installed on the internal driving unit are used to realize the omni-directional motion of the robot, one motor is used to make the robot move straight and another is used to make it steer. Its motion analysis, kinematics modeling and controllability analysis are presented. Two typical motion simulations show that the unevenness of ground has big influence on the open-loop trajectory tracking of the robot. At last, motion performance of this spherical robot in several typical environments is presented with prototype experiments. Qiang Zhan, Yao Cai, Caixia Yan |
ICRA | 3 |