EDBT 2026 Demo / reviewers in the wild / expert
Lixin Duan
dblp:54/7057
· DBLP profile ↗
108ranked-venue papers
10as first author
70since 2021 · last 2026
0000-0002-0723-4016ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 77 · 5 first-author · 50 since 2021Artificial intelligence and machine learning · 59 · 8 first-author · 33 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CasMoE: A Cascaded Framework for Efficient MoE Inference on Resource-constrained DevicesabstractThe Mixture-of-Experts (MoE) architecture has emerged as a key enabler for scaling large language models (LLMs), empowering increased model capacity with minimal computational overhead through gating-based dynamic expert activation. However, due to the memory demands introduced by expert modules, MoE inference on resource-constrained devices is still challenging. Existing methods such as model compression and parameter offloading provide partial alleviation but often lead to reduced accuracy or increased latency. In this paper, we propose CasMoE, a general and efficient cascaded framework for accelerating MoE inference on resource-constrained devices. CasMoE employs a two-stage offline-online approach to facilitate efficient expert prefetching. In the offline stage, a parameterized Expert Activation Predictor (EAP) is introduced to accurately predict the corresponding expert activation from the incoming prompt. In the online stage, a non-parametric Expert Activation Matcher (EAM) supporting fast expert retrieval is then integrated with the EAP to form a cascade planner that operates independently of the MoE architecture, predicting activated experts for all MoE layers in a single pass prior to decoding. A gating mechanism is also incorporated to dynamically adjust the sensitivity of the EAM and EAP, enabling a flexible trade-off between inference efficiency and quality. Extensive experiments on diverse downstream tasks demonstrate CasMoE’s effectiveness in accelerating inference while preserving high accuracy. Haowen He, Liang Zhao 0004, Xiaoheng Deng, Lixin Duan, Shaohua Wan 0001 |
AAAI | 5 |
| 2026 | Deep contrastive graph clustering with information preservation
Hu Lu, Haotian Hong, Fuhao Shi, Shengli Wu 0001, Lixin Duan, Shaohua Wan 0001 |
Pattern Recognit. | 5 |
| 2026 | Performance Optimization of Split Federated Learning in Heterogeneous Edge Computing EnvironmentsabstractClients in federated learning (FL) may exhibit varying computing capabilities, leading to prolonged training latency when deploying complex deep neural networks. To address this challenge, split federated learning (SFL) presents an approach that offloads the main computational workload from resource-constrained devices to a server, while enabling parallel training. However, there are two significant limitations of existing SFL frameworks: The adoption of a uniform cut layer strategy fails to take into account the heterogeneous among clients; it fails to effectively utilize server-side resources to improve training efficiency. This article presents a framework, i.e., heterogeneous split federated learning, which considers personalized cut layer selection and server resource configuration to accelerate SFL in heterogeneous edge computing environments. By splitting the global model into two components for each client, our framework jointly optimizes both client-side workload, batch size control, and server resource configuration strategy, while considering device heterogeneity. Specifically, we develop an alternating iterative scheduling algorithm to obtain an approximate scheme for the cut layer, batch sizes, and server resource configuration to alleviate the impact of device heterogeneity. The experimental results illustrate that HSFL outperforms the compared methods, achieving performance improvements of up to 3.9%$\sim$32.2% across two datasets under various data distribution scenarios, which demonstrates the effectiveness of the proposed strategies. Junyan Hu, Yuansheng Liang, Yanping Chen 0006, Gang Liu 0038, Weiwei Chen 0004, Lixin Duan |
IEEE Trans. Ind. Informatics | 6 |
| 2026 | ROOT: Region-Word Alignment With Partial Optimal Transport for Open-Vocabulary Object DetectionabstractOpen-vocabulary object detection (OVD) aims to detect novel object concepts by mining region-word correspondences from image-text pairs, yet current methods often produce false correspondences. While some strategies (e.g., one-to-one matching) were proposed to mitigate this issue, they often sacrifice numerous valuable region-word pairs during the matching process. To overcome these challenges, we propose a novel comprehensive alignment method, named Region-word Alignment with Partial Optimal Transport (ROOT) framework, which reframes the region-word matching task as a problem of partial distribution alignment. Unlike traditional optimal transport, which shifts the full mass of the distribution, partial optimal transport enables selective matching, making it more robust to noise in region and word alignment. Specifically, ROOT first employs partial optimal transport to obtain an optimal transport plan for region and word feature alignment. This transport plan is then used to compute a matching reliability score for each region-word pair, which reweights the contrastive alignment loss to enhance accuracy. By enabling more flexible and reliable region-text matches, ROOT significantly reduces misalignment errors while preserving valuable region-word correspondences. Extensive experiments on standard benchmarks OV-COCO and OV-LVIS show that our ROOT outperforms the previous state-of-the-art works, demonstrating the effectiveness of our approach. Jinhong Deng, Yinjie Lei, Wen Li 0001, Lixin Duan |
IEEE Trans. Image Process. | 4 |
| 2026 | Lighted-SAM: Lightening Open-World SAM for Low-Light SegmentationabstractSegment Anything Model (SAM) has achieved impressive segmentation performance in an open-world setting. However, SAM relies heavily on high-quality input images and usually struggles in low-light conditions. This is mainly caused by the pre-training dataset, SA-1B, in which low-light samples constitute a relatively small fraction of the data. This lack of presence leads to a noticeable weakness when SAM is applied in real-world dark environments. With the motivation of improving SAM's performance under low-light conditions while retaining its strong zero-shot capability, this work proposes an alignment stage between the pre-training stage and testing stage. Unlike existing low-light studies that mainly focus on task-specific and close-set settings, our work further emphasizes pursuing the segmentation ability under low-light conditions for open-world models. To this end, we construct DarkSeg58K, a realistic and diverse dataset, which serves as the alignment dataset to support this stage. We further introduce Lighted-SAM as the lightweight repair strategy to fix SAM's performance in low-light conditions. Different from existing methods focusing on introducing spectral adapters into the model design and training this model end-to-end, Lighted-SAM introduces the Spectral Information Resonance (SIR) mechanism to harmoniously integrate the spectral enhancement module into SAM, which is usually kept frozen due to its large-scale parameters. Based on our lightweight repairing strategy, Lighted-SAM can improve SAM's ability in low-light conditions while preserving its zero-shot ability. Experiments on different benchmarks validate the superiority of our approach. Code is available at: https://github.com/Jaaaahan/LightedSAM. Yuhan Jia, Lixin Duan, Wen Li 0001, Fengmao Lv |
IEEE Trans. Image Process. | 2 |
| 2026 | Tuning-Free Adaptive Style Incorporation for Structure-Consistent Text-Driven Style TransferabstractText-driven style transfer methods leveraging diffusion models have shown impressive creativity, yet they still face challenges in maintaining consistent structure and content preservation. Existing methods often directly concatenate the content and style prompts for a prompt-level style injection. However, this coarse-grained style injection strategy inevitably leads to structural deviations in the stylized images. This poses a significant obstacle for professional artists and creators seeking precise artistic editing. In this work, we strive to attain a harmonious balance between content preservation and style transformation. We propose Adaptive Style Incorporation (ASI), to achieve fine-grained feature-level style incorporation. It consists of the Siamese Cross-Attention (SiCA) to decouple the single-track cross-attention to a dual-track structure to obtain separate content and style features, and the Adaptive Content-Style Blending (AdaBlending) module to couple the content and style information from a structure-consistent manner. Experimentally, our method exhibits much better performance in both structure preservation and stylized effects. Yanqi Ge, Jiaqi Liu 0004, Qingnan Fan, Xi Jiang 0009, Shuai Qin, Wen Li 0001, Lixin Duan |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2025 | S-INF: Towards Realistic Indoor Scene Synthesis via Scene Implicit Neural FieldabstractLearning-based methods have become increasingly popular in 3D indoor scene synthesis (ISS), showing superior performance over traditional optimization-based approaches. These learning-based methods typically model distributions on simple yet explicit scene representations using generative models. However, due to the oversimplified explicit representations that overlook detailed information and the lack of guidance from multimodal relationships within the scene, most learning-based methods struggle to generate indoor scenes with realistic object arrangements and styles. In this paper, we introduce a new method, Scene Implicit Neural Field (S-INF), for indoor scene synthesis, aiming to learn meaningful representations of multimodal relationships, to enhance the realism of indoor scene synthesis. S-INF assumes that the scene layout is often related to the object-detailed information. It disentangles the multimodal relationships into scene layout relationships and detailed object relationships, fusing them later through implicit neural fields (INFs). By learning specialized scene layout relationships and projecting them into S-INF, we achieve a realistic generation of scene layout. Additionally, S-INF captures dense and detailed object relationships through differentiable rendering, ensuring stylistic consistency across objects. Through extensive experiments on the benchmark 3D-FRONT dataset, we demonstrate that our method consistently achieves state-of-the-art performance under different types of ISS. Zixi Liang, Haifeng Wu, Wen Li 0001, Lixin Duan |
AAAI | 6 |
| 2025 | LidarGait++: Learning Local Features and Size Awareness from LiDAR Point Clouds for 3D Gait RecognitionabstractPoint clouds have gained growing interest in gait recognition. However, current methods, which typically convert point clouds into 3D voxels, often fail to extract essential gait-specific features. In this paper, we explore gait recognition within 3D point clouds from the perspectives of architectural designs and gait representation modeling. We indicate the significance of local and body size features in 3D gait recognition and introduce LidarGait++, a novel framework combining advanced local representation learning techniques with a novel size-aware learning mechanism. Specifically, LidarGait++ utilizes Set Abstraction (SA) layer and Pyramid Point Pooling (P3) layer for learning locally fine-grained gait representations from 3D point clouds directly. Both the SA and P3layers can be further enhanced with size-aware learning to make the model aware of the actual size of the subjects. In the end, LidarGait++ not only outperforms current state-of-the-art methods, but it also consistently demonstrates robust performance and great generalizability on two benchmarks. Our extensive experiments validate the effectiveness of size and local features in 3D gait recognition. Chuanfu Shen, Lixin Duan, Shiqi Yu 0001 |
CVPR | 3 |
| 2025 | GeoDepth: From Point-to-Depth to Plane-to-Depth Modeling for Self-Supervised Monocular Depth EstimationabstractSelf-supervised monocular depth estimation has long been treated as a point-wise prediction problem, where the depth of each pixel is usually estimated independently. However, artifacts are often observed in the estimated depth map, e.g., depth values for points located in the same region may jump dramatically. To address this issue, we propose a novel self-supervised monocular depth estimation framework called GeoDepth, where we explore the intrinsic geometric representation in 3D scenes for producing accurate and continuous depth maps. In particular, we model the complex 3D scene as a collection of planes with varying sizes, where each plane is characterized by a unique set of parameters, namely planar normal (indicating plane orientation) and planar offset (defining the perpendicular distance from the camera center to the plane). Under this modeling, points in the same plane are enforced to share a unique representation and their depth variations related only to pixel coordinates, thus this geometric relationship can be exploited to regularize the depth variations of these points. To this end, we design a structured plane generation module that introduces spatio-temporal geometric cues and the plane uniqueness principle to recover the correct scene plane representation. In addition, we develop a depth discontinuity module to identify depth discontinuity regions and subsequently optimize them. Our experiments on the KITTI and NYUv2 datasets demonstrate that GeoDepth achieves state-of-the-art performance, with additional tests on Make3D and ScanNet validating its generalization capabilities. Haifeng Wu, Shuhang Gu, Lixin Duan, Wen Li 0001 |
CVPR | 3 |
| 2025 | ResCLIP: Residual Attention for Training-free Dense Vision-language InferenceabstractWhile vision-language models like CLIP have shown remarkable success in open-vocabulary tasks, their application is currently confined to image-level tasks, and they still struggle with dense predictions. Recent works often attribute such deficiency in dense predictions to the self-attention layers in the final block, and have achieved commendable results by modifying the original query-key attention to self-correlation attention, (e.g., query-query and key-key attention). However, these methods overlook the cross-correlation attention (query-key) properties, which capture the rich spatial correspondence. In this paper, we reveal that the cross-correlation of self-attention in non-final layers of CLIP also exhibits localization properties. Therefore, we propose the Residual Cross-correlation Self-attention (RCS) module, which leverages the cross-correlation self-attention from intermediate layers to remold the attention in the final block. The RCS module effectively reorganizes spatial information, unleashing the localization potential within CLIP for dense vision-language inference. Furthermore, to enhance the focus on regions of the same categories and local consistency, we propose the Semantic Feedback Refinement (SFR) module, which utilizes semantic segmentation maps to further adjust the attention scores. By integrating these two strategies, our method, termed ResCLIP, can be easily incorporated into existing approaches as a plug-and-play module, significantly boosting their performance in dense vision-language inference. Extensive experiments across multiple standard benchmarks demonstrate that our method surpasses state-of-the-art training-free methods, validating the effectiveness of the proposed approach. Code is available at https://github.com/yvhangyang/ResCLIP. Jinhong Deng, Wen Li 0001, Lixin Duan |
CVPR | 4 |
| 2025 | Balanced Sharpness-Aware Minimization for Imbalanced Regression
Yahao Liu, Qin Wang 0013, Lixin Duan, Wen Li 0001 |
ICCV | 3 |
| 2025 | The Devil Is in the Spurious Correlations: Boosting Moment Retrieval With Dynamic LearningabstractGiven a textual query along with a corresponding video, the objective of moment retrieval aims to localize the moments relevant to the query within the video. While commendable results have been demonstrated by existing transformer-based approaches, predicting the accurate temporal span of the target moment is still a major challenge. This paper reveals that a crucial reason stems from the spurious correlation between the text query and the moment context. Namely, the model makes predictions by overly associating queries with background frames rather than distinguishing target moments. To address this issue, we propose a dynamic learning approach for moment retrieval, where two strategies are designed to mitigate the spurious correlation. First, we introduce a novel video synthesis approach to construct a dynamic context for the queried moment, enabling the model to attend to the target moment of the corresponding query across dynamic backgrounds. Second, to alleviate the over-association with backgrounds, we enhance representations temporally by incorporating text-dynamics interaction, which encourages the model to align text with target moments through complementary dynamic representations. With the proposed method, our model significantly alleviates the spurious correlation issue in moment retrieval and establishes new state-of-the-art performance on two popular benchmarks, \ie, QVHighlights and Charades-STA. In addition, detailed ablation studies and evaluations across different architectures demonstrate the generalization and effectiveness of the proposed strategies. Our code will be publicly available. Xinyang Zhou, Fanyue Wei, Lixin Duan, Angela Yao, Wen Li 0001 |
ICCV | 3 |
| 2025 | P2WNet: Homography Estimation for Part-To-Whole and Cross-Modality ScenariosabstractDeep learning-based homography estimation has achieved remarkable advances in recent years. However, existing methods face limitations in the Part-To-Whole (P2W) scenario, where the template image corresponds to a small portion of the search image, as their designs are tailored to image pairs with similar content and limited displacement. To address this issue, we propose P2WNet, a novel framework for part-to-whole and cross-modality homography estimation. First, we tailor a pseudo-siamese encoder to handle cross-modal inputs and incorporate a transformer-based cascade for feature enhancement. Furthermore, we design a novel P2W matching module to capture and represent the correspondences between image pairs in the Part-To-Whole scenario. These robust features are fed into a prediction module to estimate the homography matrix. Additionally, we propose a dataset to validate our model in P2W scenario. Experiments on both our dataset and public benchmarks (DroneVehicle, GoogleMap, MSCOCO) demonstrate that P2WNet achieves superior performance in the P2W scenario and performs competitively in conventional scenarios. Code is available at https://github.com/xuanxh1/P2WNET. ShangXuan Xie, Haifeng Wu, Wen Li 0001, Lixin Duan |
ICME | 4 |
| 2025 | EndoDUM: Unsupervised Endoscopic Depth Estimation with Uncertainty MaskabstractUnsupervised depth estimation plays a key role in endoscopic minimally invasive surgery. However, endoscopic images often suffer from abnormal lighting, with some areas being overexposed and others underexposed. This leads to poor performance of current depth estimation methods in these regions. In this paper, we propose Unsupervised Endoscopic Depth Estimation with Uncertainty Mask (EndoDUM), the first unsupervised method for depth estimation that addresses abnormal lighting regions in endoscopic images. EndoDUM addresses the abnormal lighting problem in endoscopic images by learning illumination and applying a soft mask to regions with abnormal lighting. Specifically, it estimates depth information from a single RGB image accurately through three novel components: 1) The illumination calibration network, which calibrates the raw endoscopic image by estimating lighting, 2) The uncertainty mask, which suppresses overexposed or underexposed regions based on the illumination calibration ratio, and 3) A Three-Dimensional dynamic convolution module that captures local fine-grained features by leveraging complementary attention across three dimensions. Experimental results on the SCARED and Hamlyn datasets demonstrate that EndoDUM significantly outperforms existing methods in depth estimation tasks, achieving state-of-the-art (SOTA) performance. Xuanxuan Liu, Minhao Liu, Lixin Duan |
IJCNN | 6 |
| 2025 | Leveraging Diffusion Models for Continual Test-Time Adaptation in Fundus Image Classification
Mingsi Liu, Xiang Li 0115, Mengxiang Guo, Lixin Duan, Huihui Fang, Yanwu Xu 0001 |
MICCAI (5) | 4 |
| 2025 | GL-LCM: Global-Local Latent Consistency Models for Fast High-Resolution Bone Suppression in Chest X-Ray Images
Yifei Sun 0005, Zhanghao Chen, Yuqing Lu, Lixin Duan, Fenglei Fan, Ahmed El-Azab, Changmiao Wang, Ruiquan Ge |
MICCAI (13) | 5 |
| 2025 | Semantic discrete decoder based on adaptive pixel clustering for monocular depth estimation
Xuanxuan Liu, Mingzhi Ye, Tongwei Lu, Lixin Duan |
Neural Networks | 5 |
| 2025 | Adaptive Multi-Scale Language Reinforcement for Multimodal Named Entity RecognitionabstractOver the recent years, multimodal named entity recognition has gained increasing attentions due to its wide applications in social media. The key factor of multimodal named entity recognition is to effectively fuse information of different modalities. Existing works mainly focus on reinforcing textual representations by fusing image features via the cross-modal attention mechanism. However, these works are limited in reinforcing the text modality at the token level. As a named entity usually contains several tokens, modeling token-level inter-modal interactions is suboptimal for the multimodal named entity recognition problem. In this work, we propose a multimodal named entity recognition approach dubbed Adaptive Multi-scale Language Reinforcement (AMLR) to implement entity-level language reinforcement. To this end, our model first expands token-level textual representations into multi-scale textual representations which are composed of language units of different lengths. After that, the visual information reinforces the language modality by modeling the cross-modal attention between images and expanded multi-scale textual representations. Unlike existing token-level language reinforcement methods, the word sequences of named entities can be directly interacted with the visual features as a whole, making the modeled cross-modal correlations more reasonable. Although the underlying entity is not given, the training procedure can encourage the relevant image contents to adaptively attend to the appropriate language units, making our approach not rely on the pipeline design. Comprehensive evaluation results on two public Twitter datasets clearly demonstrate the superiority of our proposed model. Enping Li, Tianrui Li 0001, Huaishao Luo, Jielei Chu, Lixin Duan, Fengmao Lv |
IEEE Trans. Multim. | 5 |
| 2025 | Simultaneous Detection and Interaction Reasoning for Object-Centric Action RecognitionabstractThe interactions between human and objects are important for recognizing object-centric actions. Existing methods usually adopt a two-stage pipeline, where object proposals are first detected using a pretrained detector, and then are fed to an action recognition model for extracting video features and learning the object relations for action recognition. However, since the action prior is unknown in the object detection stage, important objects could be easily overlooked, leading to inferior action recognition performance. In this paper, we propose an end-to-end object-centric action recognition framework that simultaneously performsDetectionAndInteractionReasoning (dubbed DAIR) in one stage. Particularly, after extracting video features using a base network, we design three consecutive modules for simultaneously learning object detection and interaction reasoning. Firstly, we build a Patch-based Object Decoder (PatchDec) to generate object proposals from video patch tokens. Then, we design an Interactive Object Refining and Aggregation (IRA) to identify the interactive objects that are important for action recognition. The IRA module adjusts the interactiveness scores of proposals based on their relative position and appearance, and aggregates the object-level information into global video representation. Finally, we build an Object Relation Modeling (ORM) module to encode the object relations. These three modules together with the video feature extractor can be trained jointly in an end-to-end fashion, thus avoiding the heavy reliance on an off-the-shelf object detector, and reducing the multi-stage training burden. We conduct experiments on two datasets, Something-Else and Ikea-Assembly, to evaluate the performance of our proposed approach on conventional, compositional, and few-shot action recognition tasks. Through in-depth experimental analysis, we show the crucial role ofinteractiveobjects in learning for action recognition, and we can outperform state-of-the-art methods on both datasets. We hope our DAIR can provide a new perspective for object-centric action recognition. Xunsong Li, Pengzhan Sun 0001, Yangcen Liu, Lixin Duan, Wen Li 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Segmenting Anything in the Dark via Depth PerceptionabstractImage segmentation under low-light conditions is essential in real-world applications, such as autonomous driving and video surveillance systems. The recent Segment Anything Model (SAM) exhibits strong segmentation capability in various vision applications. However, its performance could be severely degraded under low-light conditions. On the other hand, multimodal information has been exploited to help models construct more comprehensive understanding of scenes under low-light conditions by providing complementary information (e.g., depth). Therefore, in this work, we present a pioneer attempt that elevates a unimodal vision foundation model (e.g., SAM) to a multimodal one, by efficiently integrating additional depth information under low-light conditions. To achieve that, we propose a novel method called Depth Perception SAM (DPSAM) based on the SAM framework. Specifically, we design a modality encoder to extract the depth information and the Depth Perception Layers (DPLs) for mutual feature refinement between RGB and depth features. The DPLs employ the cross-modal attention mechanism to mutually query effective information from both RGB and depth for the subsequent feature refinement. Thus, DPLs can effectively leverage the complementary information from depth to enrich the RGB representations and obtain comprehensive multimodal visual representations for segmenting anything in the dark. To this end, our DPSAM maximally maintains the instinct expertise of SAM for RGB image segmentation and further leverages on the strength of depth for enhanced segmenting anything capability, especially for cases that are likely to fail with RGB only (e.g., low-light or complex textures). As demonstrated by extensive experiments on four RGBD benchmark datasets, DPSAM clearly improves the performance for the segmenting anything performance in the dark, e.g., +12.90% mIoU and +16.23% mIoU on LLRGBD and DeLiVER, respectively. Our code and model will be made publicly available at:https://github.com/liupeng3425/DPSAM. Peng Liu 0049, Jinhong Deng, Lixin Duan, Wen Li 0001, Fengmao Lv |
IEEE Trans. Multim. | 3 |
| 2025 | Coherence-guided Preference Disentanglement for Cross-domain RecommendationsabstractDiscovering user preferences across different domains is pivotal in cross-domain recommendation systems, particularly when platforms lack comprehensive user-item interactive data. The limited presence of shared users often hampers the effective modeling of common preferences. While leveraging shared items’ attributes, such as category and popularity, can enhance cross-domain recommendation performance, the scarcity of shared items between domains has limited research in this area. To address this, we propose a Coherence-guided Preference Disentanglement (CoPD) method aimed at improving cross-domain recommendation by (i) explicitly extracting shared item attributes to guide the learning of shared user preferences and (ii) disentangling these preferences to identify specific user interests transferred between domains. CoPD introduces coherence constraints on item embeddings of shared and specific domains, aiding in extracting shared attributes. Moreover, it utilizes these attributes to guide the disentanglement of user preferences into separate embeddings for interest and conformity through a popularity-weighted loss. Experiments conducted on real-world datasets demonstrate the superior performance of our proposed CoPD over existing competitive baselines, highlighting its effectiveness in enhancing cross-domain recommendation performance. The code is available at https://github.com/XiangZongyi/CoPD . Zongyi Xiang, Yan Zhang 0036, Lixin Duan, Hongzhi Yin, Ivor W. Tsang |
ACM Trans. Inf. Syst. | 3 |
| 2024 | Beyond Prototypes: Semantic Anchor Regularization for Better Representation LearningabstractOne of the ultimate goals of representation learning is to achieve compactness within a class and well-separability between classes. Many outstanding metric-based and prototype-based methods following the Expectation-Maximization paradigm, have been proposed for this objective. However, they inevitably introduce biases into the learning process, particularly with long-tail distributed training data. In this paper, we reveal that the class prototype is not necessarily to be derived from training features and propose a novel perspective to use pre-defined class anchors serving as feature centroid to unidirectionally guide feature learning. However, the pre-defined anchors may have a large semantic distance from the pixel features, which prevents them from being directly applied. To address this issue and generate feature centroid independent from feature learning, a simple yet effective Semantic Anchor Regularization (SAR) is proposed. SAR ensures the inter-class separability of semantic anchors in the semantic space by employing a classifier-aware auxiliary cross-entropy loss during training via disentanglement learning. By pulling the learned features to these semantic anchors, several advantages can be attained: 1) the intra-class compactness and naturally inter-class separability, 2) induced bias or errors from feature learning can be avoided, and 3) robustness to the long-tailed problem. The proposed SAR can be used in a plug-and-play manner in the existing models. Extensive experiments demonstrate that the SAR performs better than previous sophisticated prototype-based methods. The implementation is available at https://github.com/geyanqi/SAR. Yanqi Ge, Qiang Nie, Yong Liu 0020, Chengjie Wang 0001, Feng Zheng 0001, Wen Li 0001, Lixin Duan |
AAAI | 8 |
| 2024 | Beyond Viewpoint: Robust 3D Object Recognition Under Arbitrary Views Through Joint Multi-part Representation
Linlong Fan, Yanqi Ge, Wen Li 0001, Lixin Duan |
ECCV (52) | 5 |
| 2024 | Powerful and Flexible: Personalized Text-to-Image Generation via Reinforcement Learning
Fanyue Wei, Wei Zeng 0008, Dawei Yin 0001, Lixin Duan, Wen Li 0001 |
ECCV (27) | 5 |
| 2024 | Learning Semantic Latent Directions for Accurate and Controllable Human Motion Prediction
Jiale Tao, Wen Li 0001, Lixin Duan |
ECCV (21) | 4 |
| 2024 | Diffusion-Enhanced Transformation Consistency Learning for Retinal Image Segmentation
Xiang Li 0115, Huihui Fang, Mingsi Liu, Yanwu Xu 0001, Lixin Duan |
MICCAI (11) | 5 |
| 2024 | Cache-Driven Spatial Test-Time Adaptation for Cross-Modality Medical Image Segmentation
Xiang Li 0115, Huihui Fang, Changmiao Wang, Mingsi Liu, Lixin Duan, Yanwu Xu 0001 |
MICCAI (11) | 5 |
| 2024 | Towards Unsupervised Model Selection for Domain Adaptive Object DetectionabstractEvaluating the performance of deep models in new scenarios has drawn increasing attention in recent years due to the wide application of deep learning techniques in various fields. However, while it is possible to collect data from new scenarios, the annotations are not always available. Existing Domain Adaptive Object Detection (DAOD) works usually report their performance by selecting the best model on the validation set or even the test set of the target domain, which is highly impractical in real-world applications. In this paper, we propose a novel unsupervised model selection approach for domain adaptive object detection, which is able to select almost the optimal model for the target domain without using any target labels. Our approach is based on the flat minima principle, i.e., models located in the flat minima region in the parameter space usually exhibit excellent generalization ability. However, traditional methods require labeled data to evaluate how well a model is located in the flat minima region, which is unrealistic for the DAOD task. Therefore, we design a Detection Adaptation Score (DAS) approach to approximately measure the flat minima without using target labels. We show via a generalization bound that the flatness can be deemed as model variance, while the minima depend on the domain distribution distance for the DAOD task. Accordingly, we propose a Flatness Index Score (FIS) to assess the flatness by measuring the classification and localization fluctuation before and after perturbations of model parameters and a Prototypical Distance Ratio (PDR) score to seek the minima by measuring the transferability and discriminability of the models. In this way, the proposed DAS approach can effectively represent the degree of flat minima and evaluate the model generalization ability on the target domain. We have conducted extensive experiments on various DAOD benchmarks and approaches, and the experimental results show that the proposed DAS correlates well with the performance of DAOD models and can be used as an effective tool for model selection after training. The code will be released at https://github.com/HenryYu23/DAS. Hengfu Yu, Jinhong Deng, Wen Li 0001, Lixin Duan |
NeurIPS | 4 |
| 2024 | Deep graph tensor learning for temporal link prediction
Zhen Liu 0006, Wen Li 0001, Lixin Duan |
Inf. Sci. | 4 |
| 2024 | Feature Re-Representation and Reliable Pseudo Label Retraining for Cross-Domain Semantic SegmentationabstractThis paper presents a novel unsupervised domain adaptation method for semantic segmentation. We argue that a good representation of the target-domain data should keep both the knowledge from the source domain and the target-domain-specific information. To obtain the knowledge from the source domain, we first learn a set of bases to characterize the feature distribution of the source domain, then features from both the source and the target domain are re-represented as a weighted summation of the source bases. A discriminator is additionally introduced to make the re-representation responsibilities of both domain features under the same bases indistinguishable. In this way, the domain gap between the source re-representation and target re-representation is minimized, and the re-represented target domain features contain the source domain information. Then we combine the feature re-representation with the original domain-specific feature together for subsequent pixel-wise classification. To further make the re-represented target features semantically meaningful, a Reliable Pseudo Label Retraining (RPLR) strategy is proposed, which utilizes the consistency of the prediction by the networks trained with multi-view source images to select the clean pseudo labels on unlabeled target images for re-training. Extensive experiments demonstrate the competitive performance of our approach for unsupervised domain adaptation on the semantic segmentation benchmarks. Jing Li 0117, Kang Zhou 0001, Shenhan Qian, Wen Li 0001, Lixin Duan, Shenghua Gao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Balanced Teacher for Source-Free Object DetectionabstractWe study a practical domain adaptation task, named source-free object detection (SFOD), which aims to adapt a pre-trained source detector to an unlabeled target domain without access to the original labeled source domain samples. In this paper, we design a new self-training approach for SFOD called Balance Teacher based on the mean teacher model. We target two key issues when using self-training for SFOD: 1) imbalanced label distribution when using pseudo-labels for supervising the model training, and 2) imbalanced image distribution,i.e., significant data variance in the target domain. To address these issues, we first design a Class-balanced Instance Selection (CBIS) module to automatically balance different classes when selecting pseudo-labeled instances during the training process. Then, we propose a Progressive Target Variance Minimization (PTVM) to cope with the imbalanced image distribution in the target domain, where the feature distributions of certainty and uncertainty target samples are progressively aligned to alleviate the data distribution variance. In this way, the teacher model can provide high-quality pseudo-labels and guide the student model to adapt gradually to the target domain. We have conducted extensive experiments on five widely used benchmarks, and the experimental results clearly show the superiority of our method over the state-of-the-art baselines. Jinhong Deng, Wen Li 0001, Lixin Duan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | CARD: Semantic Segmentation With Efficient Class-Aware Regularized DecoderabstractSemantic segmentation has recently achieved notable advances by exploiting “class-level” contextual information during learning, e.g., the Object Contextual Representation (OCR) and Context Prior (CPNet) approaches. However, these approaches simply concatenate class-level information to pixel features to boost pixel representation learning, which cannot fully utilize intra-class and inter-class contextual information. Moreover, these approaches learn soft class centers based on coarse mask prediction, which is prone to error accumulation. To better exploit class-level information, we propose a universal Class-Aware Regularization (CAR) approach to optimize the intra-class variance and inter-class distance during feature learning, motivated by the fact that humans can recognize an object by itself no matter which other objects it appears with. Moreover, we design a dedicated decoder for CAR (named CARD), which consists of a novel spatial token mixer and an upsampling module, to maximize its gain for existing baselines while being highly efficient in terms of computational cost. Specifically, CAR consists of three novel loss functions. The first loss function encourages more compact class representations within each class, the second directly maximizes the distance between different class centers, and the third further pushes the distance between inter-class centers and pixels. Furthermore, the class center in our approach is directly generated from ground truth instead of from the error-prone coarse prediction. CAR can be directly applied to most existing segmentation models during training, including OCR and CPNet, and can largely improve their accuracy at no additional inference overhead. Extensive experiments and ablation studies conducted on multiple benchmark datasets demonstrate that the proposed CAR can boost the accuracy of all baseline models by up to 2.23% mIOU with superior generalization ability. CARD outperforms state-of-the-art approaches on multiple benchmarks with a highly efficient architecture. The code will be available at https://github.com/edwardyehuang/CAR. Liang Chen 0026, Wenjing Jia, Xiangjian He, Lixin Duan, Xuefei Zhe, Linchao Bao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | High-Level Feature Guided Decoding for Semantic SegmentationabstractExisting pyramid-based upsamplers (e.g. SemanticFPN), although efficient, usually produce less accurate results compared to dilation-based models when using the same backbone. This is partially caused by thecontaminatedhigh-level features since they are fused and fine-tuned with noisy low-level features on limited data. To address this issue, we propose to use powerful pre-trainedhigh-levelfeatures asguidance (HFG) so that the upsampler can produce robust results. Specifically,onlythe high-level features from the backbone are used to train the class tokens, which are then reused by the upsampler for classification, guiding the upsampler features to more discriminative backbone features. One crucial design of the HFG is to protect the high-level features from being contaminated by using proper stop-gradient operations so that the backbone does not update according to the noisy gradient from the upsampler. To push the upper limit of HFG, we introduce acontextaugmentationencoder (CAE) that can efficiently and effectively operate on the low-resolution high-level feature, resulting in improved representation and thus better guidance. We named our complete solution as the High-Level Features Guided Decoder (HFGD). We evaluate the proposed HFGD on three benchmarks: Pascal Context, COCOStuff164k, and Cityscapes. HFGD achieves state-of-the-art results among methods that do not use extra training data, demonstrating its effectiveness and generalization ability. Shenghua Gao, Wen Li 0001, Lixin Duan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | CAFA: Cross-Modal Attentive Feature Alignment for Cross-Domain Urban Scene SegmentationabstractAutonomous driving systems rely heavily on semantic segmentation models for accurate and safe decision-making. High segmentation performance in real-world urban scenes is crucial for autonomous vehicles, while substantial pixel-level labels are required during model training. Unsupervised domain adaptation (UDA) techniques are widely used to adapt the segmentation model trained on the synthetic data (i.e., source domain) to the real-world data (i.e., target domain) since obtaining pixel-level annotations is fairly easy in the synthetic environment. Recently, increasing UDA approaches promote cross-domain semantic segmentation (CDSS) by fusing the depth information into the RGB features. However, feature fusion does not necessarily eliminate the domain-specific components in the RGB features, which can result in the features still being influenced by domain-specific information. To address this, we propose a novel cross-modal attentive feature alignment (CAFA) framework for CDSS, which provides an explicit perspective of using depth information to align the main backbone RGB features of both domains in a nonadversarial manner. In particular, considering that the depth modality is less affected by the domain gap, we employ depth as an intermediate modality and align the RGB features by attending RGB features to the depth modality through constructing an auxiliary multimodal segmentation task. The state-of-the-art performance of our CAFA can be achieved on benchmark tasks, such as Synthia$\to$Cityscapes and grand theft auto (GTA)$\to$Cityscapes. Peng Liu 0049, Yanqi Ge, Lixin Duan, Wen Li 0001, Fengmao Lv |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Transferring Multi-Modal Domain Knowledge to Uni-Modal Domain for Urban Scene SegmentationabstractSynthetic data (i.e., source domain) have been widely adopted to improve the semantic segmentation performance for real-world images (i.e., target domain), since obtaining pixel-level annotations is fairly easy in the synthetic environment. Traditional domain adaptation methods normally focus on learning in the RGB modality only. We notice that the synthetic environment can generate depth information of semantic objects at almost no cost, while it is nontrivial to collect such information in the real-world scenario. In this case, we employ the depth information of synthetic data in this work to further boost the segmentation performance, and then transform the uni-modal problem into a multi-modal one. In this work, we focus on urban scene understanding and make a pioneer attempt on learning uni-modal feature representations for real-world images by mining from multi-modal knowledge of synthetic images with additional depth information. To this end, we propose a novel method called Multi-modal Domain Knowledge Transfer (MDKT), which transfers the multi-modal knowledge of the source domain to the uni-modal target domain through domain adaptation. In MDKT, we first employ the Cross-Modal Correlation (CMC) module to enhance the source features by fusing the RGB and depth information. Then, the uni-modal target domain feature and multi-modal source domain feature are aligned through the Modal-Imbalanced Adversarial Training (MIAT) strategy, which transfers the multi-modal knowledge to the uni-modal network in the target domain. We conduct extensive experiments on several benchmark settings for urban scene understanding. The promising results clearly show the effectiveness of our proposed MDKT approach. Peng Liu 0049, Yanqi Ge, Lixin Duan, Wen Li 0001, Haonan Luo 0002, Fengmao Lv |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Calibrate the Inter-Observer Segmentation Uncertainty via Diagnosis-First PrincipleabstractMany of the tissues/lesions in the medical images may be ambiguous. Therefore, medical segmentation is typically annotated by a group of clinical experts to mitigate personal bias. A common solution to fuse different annotations is the majority vote, e.g., taking the average of multiple labels. However, such a strategy ignores the difference between the grader expertness. Inspired by the observation that medical image segmentation is usually used to assist the disease diagnosis in clinical practice, we propose the diagnosis-first principle, which is to take disease diagnosis as the criterion to calibrate the inter-observer segmentation uncertainty. Following this idea, a framework named Diagnosis-First segmentation Framework (DiFF) is proposed. Specifically, DiFF will first learn to fuse the multi-rater segmentation labels to a single ground-truth which could maximize the disease diagnosis performance. We dubbed the fused ground-truth as Diagnosis-First Ground-truth (DF-GT). Then, the Take and Give Model (T&G Model) to segment DF-GT from the raw image is proposed. With the T&G Model, DiFF can learn the segmentation with the calibrated uncertainty that facilitate the disease diagnosis. We verify the effectiveness of DiFF on three different medical segmentation tasks: optic-disc/optic-cup (OD/OC) segmentation on fundus images, thyroid nodule segmentation on ultrasound images, and skin lesion segmentation on dermoscopic images. Experimental results show that the proposed DiFF can effectively calibrate the segmentation uncertainty, and thus significantly facilitate the corresponding disease diagnosis, which outperforms previous state-of-the-art multi-rater learning methods. Yu Zhang 0091, Huihui Fang, Lixin Duan, Mingkui Tan, Weihua Yang, Yueming Jin, Yanwu Xu 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Cross-Domain Detection Transformer Based on Spatial-Aware and Semantic-Aware Token AlignmentabstractDetection transformers such as DETR [1] have recently exhibited promising performance for many object detection tasks, but the generalization ability of those methods is still quite limited for cross-domain adaptation scenarios. To address the cross-domain issue, a straightforward method is to perform token alignment with adversarial training in transformers. However, its performance is often unsatisfactory because the tokens in detection transformers are quite diverse and represent different spatial and semantic information. In this paper, we propose a new method for cross-domain detection transformers called spatial-aware and semantic-aware token alignment (SSTA). Specifically, we take advantage of the characteristics of cross-attention as used in the detection transformer and propose spatial-aware token alignment (SpaTA) and semantic-aware token alignment (SemTA) strategies to guide the token alignment across domains. For spatial-aware token alignment, we extract the information from the cross-attention map (CAM) to align the distribution of tokens according to their attention to object queries. For semantic-aware token alignment, we inject the category information into the cross-attention map and construct domain embedding to guide the learning of a multi-class discriminator to model the category relationship and achieve category-level token alignment during the entire adaptation process. We conduct extensive experiments on several widely-used benchmarks, and the results clearly show the effectiveness of our proposed approach over existing state-of-the-art methods. Jinhong Deng, Wen Li 0001, Lixin Duan, Dong Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Harmonious Teacher for Cross-Domain Object DetectionabstractSelf-training approaches recently achieved promising results in cross-domain object detection, where people iteratively generate pseudo labels for unlabeled target domain samples with a model, and select high-confidence samples to refine the model. In this work, we reveal that the consistency of classification and localization predictions are crucial to measure the quality of pseudo labels, and propose a new Harmonious Teacher approach to improve the self-training for cross-domain object detection. In particular, we first propose to enhance the quality of pseudo labels by regularizing the consistency of the classification and localization scores when training the detection model. The consistency losses are defined for both labeled source samples and the unlabeled target samples. Then, we further remold the traditional sample selection method by a sample reweighing strategy based on the consistency of classification and localization scores to improve the ranking of predictions. This allows us to fully exploit all instance predictions from the target domain without abandoning valuable hard examples. Without bells and whistles, our method shows superior performance in various cross-domain scenarios compared with the state-of-the-art baselines, which validates the effectiveness of our Harmonious Teacher. Our codes will be available at https://github.com/kinredon/Harmonious-Teacher. Jinhong Deng, Dongli Xu, Wen Li 0001, Lixin Duan |
CVPR | 4 |
| 2023 | Minimizing Maximum Model Discrepancy for Transferable Black-box Targeted AttacksabstractIn this work, we study the black-box targeted attack problem from the model discrepancy perspective. On the theoretical side, we present a generalization error bound for black-box targeted attacks, which gives a rigorous theoretical analysis for guaranteeing the success of the attack. We reveal that the attack error on a target model mainly depends on empirical attack error on the substitute model and the maximum model discrepancy among substitute models. On the algorithmic side, we derive a new algorithm for black-box targeted attacks based on our theoretical analysis, in which we additionally minimize the maximum model discrepancy (M3D) of the substitute models when training the generator to generate adversarial examples. In this way, our model is capable of crafting highly transferable adversarial examples that are robust to the model variation, thus improving the success rate for attacking the black-box model. We conduct extensive experiments on the ImageNet dataset with different classification models, and our proposed approach outperforms existing state-of-the-art methods by a significant margin. The code will be available at https://github.com/Asteriajojo/M3D. Tong Chu, Yahao Liu, Wen Li 0001, Jingjing Li 0001, Lixin Duan |
CVPR | 6 |
| 2023 | Multi-View Token Clustering and Fusion for 3D Object Recognition and Retrievalabstract3D object recognition has received extensive attention in recent years. Many existing methods tackle the task by rendering 3D objects from multiple views. However, most multi-view recognition methods do not utilize fine-grained information from different views, which is found to be crucial for improving 3D object representation in the multi-view setting. In this paper, we propose a transformer-based method, referred to as MVCFormer, for multi-view feature clustering and fusion. MVCFormer clusters semantically similar tokens at the same stages and selects representative fine-grained features, which helps to eliminate feature redundancy and remove cluttered backgrounds and make the selected features more diverse. On the other hand, our model also integrates selected features from all stages to obtain a discriminative 3D object representation by a cross-attention fusion method. Extensive experiments on benchmark datasets (e.g., ModelNet40, ModelNet10, ShapeNetCore55, and RGBD) clearly demonstrate the effectiveness of our proposed MVCFormer over existing baselines. Linlong Fan, Yanqi Ge, Wen Li 0001, Lixin Duan |
ICME | 4 |
| 2023 | Learning continuous piecewise non-linear activation functions for deep neural networksabstractActivation functions provide the non-linearity to deep neural networks, which are crucial for the optimization and performance improvement. In this paper, we propose a learnable continuous piece-wise nonlinear activation function (or CPN in short), which improves the widely used ReLU from three directions, i.e., finer pieces, non-linear terms and learnable parameterization. CPN is a continuous activation function with multiple pieces and incorporates non-linear terms in every interval. We give a general formulation of CPN and provide different implementations according to three key factors: whether the activation space is divided uniformly or not, whether the non-linear terms exist or not, and whether the activation function is continuous or not. We demonstrate the effectiveness of our method on image classification and single image super-resolution tasks by simply changing the activation function. For example, CPN improves 4.78% / 4.52% top-1 accuracy over ReLU on MobileNetV2_0.25 / MobileNetV2_0.35 for ImageNet classification and achieves better PSNR on several benchmarks for super-resolution. Our implementation is available at https://github.com/xc-G/CPN. Xinchen Gao, Yawei Li 0001, Wen Li 0001, Lixin Duan, Luc Van Gool, Luca Benini, Michele Magno |
ICME | 4 |
| 2023 | Double-Fine-Tuning Multi-Objective Vision-and-Language Transformer for Social Media Popularity PredictionabstractSocial media popularity prediction aims to predict future interaction or attractiveness of new posts. However, in most existing works, there is a notable deficiency in the effective treatment of numerical features. Despite their significant potential to provide ample information, these features are often inadequately processed, leading to insufficiency of information acquirement. In this paper, we introduce a method, named Double-Fine-Tuning Multi-Objective Vision-and-Language Transformer (DFT-MOVLT). To supplement the information in vision-and-language pre-training (VLP), we propose compound text, which is concatenated by numerical data and text. Furthermore, during VLP, a transformer is trained using 3 objectives to ensure thorough feature extraction. Finally, for more generalized prediction, we fine-tune 2 models using different training ways and ensemble them. To evaluate the effectiveness of each mechanism adopted in the proposed method, we conduct an array of ablation experiments. Our team achieve the 3rd place in Social Media Prediction (SMP) Challenge 2023. Xiaolu Chen, Weilong Chen, Zhongjian Zhang, Lixin Duan, Yanru Zhang |
ACM Multimedia | 5 |
| 2023 | Learning Motion Refinement for Unsupervised Face AnimationabstractUnsupervised face animation aims to generate a human face video based on the
appearance of a source image, mimicking the motion from a driving video. Existing
methods typically adopted a prior-based motion model (e.g., the local affine motion
model or the local thin-plate-spline motion model). While it is able to capture
the coarse facial motion, artifacts can often be observed around the tiny motion
in local areas (e.g., lips and eyes), due to the limited ability of these methods
to model the finer facial motions. In this work, we design a new unsupervised
face animation approach to learn simultaneously the coarse and finer motions. In
particular, while exploiting the local affine motion model to learn the global coarse
facial motion, we design a novel motion refinement module to compensate for
the local affine motion model for modeling finer face motions in local areas. The
motion refinement is learned from the dense correlation between the source and
driving images. Specifically, we first construct a structure correlation volume based
on the keypoint features of the source and driving images. Then, we train a model
to generate the tiny facial motions iteratively from low to high resolution. The
learned motion refinements are combined with the coarse motion to generate the
new image. Extensive experiments on widely used benchmarks demonstrate that
our method achieves the best results among state-of-the-art baselines. Jiale Tao, Shuhang Gu, Wen Li 0001, Lixin Duan |
NeurIPS | 4 |
| 2023 | Unsupervised Domain Adaptation for Optical Flow Estimation
Jianpeng Ding, Jinhong Deng, Yanru Zhang, Shaohua Wan 0001, Lixin Duan |
PRCV (3) | 5 |
| 2023 | Multi-modal Instance Refinement for Cross-Domain Action Recognition
Yuan Qing, Naixing Wu, Shaohua Wan 0001, Lixin Duan |
PRCV (1) | 4 |
| 2023 | MoVAE: A Variational AutoEncoder for Molecular Graph GenerationabstractMolecule generation plays an important role in accelerating drug discovery. In recent years, many molecule generation methods have been proposed based on variational autoencoders (VAEs), due to its advantages in latent manifold representation learning and training stability. However, most of the existing VAE-based models require tedious graph matching operations during training, and tend to generate invalid molecules. To overcome these limitations, in this paper, we propose a novel molecular graph variational autoencoder (MoVAE). Firstly, to avoid complicated graph matching, the proposed MoVAE only encodes and decodes all the nodes and edges individually. Secondly, to improve the generation validity, it adversarially trains the model by treating the encoder and decoder as the discriminator and generator. In addition, to generate molecules with various target conditions, the MoVAE also introduces drug property constraints and valence histogram constraints. Experiment results on two real datasets show that our model outperforms almost all the state-of-the-art algorithms. Zerun Lin, Lixin Duan, Le Ou-Yang, Peilin Zhao |
SDM | 3 |
| 2023 | Noisy Label Learning With Provable Consistency for a Wider Family of LossesabstractDeep models have achieved state-of-the-art performance on a broad range of visual recognition tasks. Nevertheless, the generalization ability of deep models is seriously affected by noisy labels. Though deep learning packages have different losses, this is not transparent for users to choose consistent losses. This paper addresses the problem of how to use abundant loss functions designed for the traditional classification problem in the presence of label noise. We present a dynamic label learning (DLL) algorithm for noisy label learning and then prove that any surrogate loss function can be used for classification with noisy labels by using our proposed algorithm, with a consistency guarantee that the label noise does not ultimately hinder the search for the optimal classifier of the noise-free sample. In addition, we provide a depth theoretical analysis of our algorithm to verify the justifies' correctness and explain the powerful robustness. Finally, experimental results on synthetic and real datasets confirm the efficiency of our algorithm and the correctness of our justifies and show that our proposed algorithm significantly outperforms or is comparable to current state-of-the-art counterparts. Defu Liu 0001, Wen Li 0001, Lixin Duan, Ivor W. Tsang, Guowu Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Learning cross-domain semantic-visual relationships for transductive zero-shot learning
Fengmao Lv, Jianyang Zhang, Guowu Yang, Lei Feng 0006, Lixin Duan |
Pattern Recognit. | 6 |
| 2023 | MetaCAR: Cross-Domain Meta-Augmentation for Content-Aware RecommendationabstractCold-start has become critical for recommendations, especially for sparse user-item interactions. Recent approaches based on meta-learning succeed in alleviating the issue, owing to the fact that these methods have strong generalization, so they can fast adapt to new tasks under cold-start settings. However, these meta-learning-based recommendation models learned with single and spase ratings are easily falling into the meta-overfitting, since the one and only rating$r_{ui}$to a specific item$i$cannot reflect a user's diverse interests under various circumstances(e.g., time, mood, age, etc), i.e. if$r_{ui}$equals to 1 in the historical dataset, but$r_{ui}$could be 0 in some circumstance. In meta-learning, tasks with these single ratings are called Non-Mutually-Exclusive(Non-ME) tasks, and tasks with diverse ratings are called Mutually-Exclusive(ME) tasks. Fortunately, a meta-augmentation technique is proposed to relief the meta-overfitting for meta-learning methods by transferring Non-ME tasks into ME tasks by adding noises to labels without changing inputs. Motivated by the meta-augmentation method, in this paper, we propose a cross-domain meta-augmentation technique for content-aware recommendation systems (MetaCAR) to construct ME tasks in the recommendation scenario. Our proposed method consists of two stages: meta-augmentation and meta-learning. In the meta-augmentation stage, we first conduct domain adaptation by a dual conditional variational autoencoder (CVAE) with a multi-view information bottleneck constraint, and then apply the learned CVAE to generate ratings for users in the target domain. In the meta-learning stage, we introduce both the true and generated ratings to construct ME tasks that enables the meta-learning recommendations to avoid meta-overfitting. Experiments evaluated in real-world datasets show the significant superiority of MetaCAR for coping with the cold-start user issue over competing baselines including cross-domain, content-aware, and meta-learning-based recommendations. Changyu Li, Yan Zhang 0036, Lixin Duan, Ivor W. Tsang, Jie Shao 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Deep Cross-Attention Network for Crowdfunding Success PredictionabstractCrowdfunding creates opportunities for entrepre- neurs. It allows startup companies to reach a large audience for fundraising and bring their creative ideas to life. In this work, we are concerned with crowdfunding project success prediction problem,i.e., to predict whether a project will successfully reach its funding goal by using its project profiles. This is important for startup companies to refine their project profiles and achieve their goals. Crowdfunding project success prediction is a typical classification problem but with a few critical challenges. On the one hand, with only coarse-grained project status as weak supervision, it is hard for a deep learning network to learn the relationship between project profiles and explain why it makes this prediction. On the other hand, on the project homepage, there are various modalities of description, including metadata, textual description, images, and videos. Among those, videos play an important role in the success of a crowdfunding project, however, were ignored in previous works, due to the difficulty in extracting useful semantic and authentic information from videos, especially for the crowdfunding project where information in different modalities are unaligned. To this end, we propose a novel framework called Deep Cross-Attention Network to learn and fuse information from introduction videos and textual descriptions of project profiles. More specifically, we develop a cross-attention block to align and represent mismatched textual description and untrimmed introduction videos and fuse the information from these two modalities, which effectively remedies the lack of supervised information caused by project status as weak supervision. More importantly, with our cross-attention mechanism, the model is able to interpret how it makes such predictions and show which keywords and keyframes it depends on. We conduct extensive experiments on two crowdfunding datasets (collected from Kickstarter and Indiegogo) and show that our method achieves superior performance over existing state-of-the-art baselines. Yi Yang 0042, Wen Li 0001, Defu Lian, Lixin Duan |
IEEE Trans. Multim. | 5 |
| 2023 | Open Set Domain Adaptation With Soft Unknown-Class RejectionabstractThe goal of domain adaptation (DA) is to train a good model for a target domain, with a large amount of labeled data in a source domain but only limited labeled data in the target domain. Conventional closed set domain adaptation (CSDA) assumes source and target label spaces are the same. However, this is not quite practical in real-world applications. In this work, we study the problem of open set domain adaptation (OSDA), which only requires the target label space to partially overlap with the source label space. Consequently, the solution to OSDA requires unknown classes detection and separation, which is normally achieved by introducing a threshold for the prediction of target unknown classes; however, the performance can be quite sensitive to that threshold. In this article, we tackle the above issues by proposing a novel OSDA method to perform soft rejection of unknown target classes and simultaneously match the source and target domains. Extensive experiments on three standard datasets validate the effectiveness of the proposed method over the state-of-the-art competitors. Yiming Xu 0006, Lin Chen 0021, Lixin Duan, Ivor W. Tsang, Jiebo Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Denoised Maximum Classifier Discrepancy for Source-Free Unsupervised Domain AdaptationabstractSource-Free Unsupervised Domain Adaptation(SFUDA) aims to adapt a pre-trained source model to an unlabeled target domain without access to the original labeled source domain samples. Many existing SFUDA approaches apply the self-training strategy, which involves iteratively selecting confidently predicted target samples as pseudo-labeled samples used to train the model to fit the target domain. However, the self-training strategy may also suffer from sample selection bias and be impacted by the label noise of the pseudo-labeled samples. In this work, we provide a rigorous theoretical analysis on how these two issues affect the model generalization ability when applying the self-training strategy for the SFUDA problem. Based on this theoretical analysis, we then propose a new Denoised Maximum Classifier Discrepancy (D-MCD) method for SFUDA to effectively address these two issues. In particular, we first minimize the distribution mismatch between the selected pseudo-labeled samples and the remaining target domain samples to alleviate the sample selection bias. Moreover, we design a strong-weak self-training paradigm to denoise the selected pseudo-labeled samples, where the strong network is used to select pseudo-labeled samples while the weak network helps the strong network to filter out hard samples to avoid incorrect labels. In this way, we are able to ensure both the quality of the pseudo-labels and the generalization ability of the trained model on the target domain. We achieve state-of-the-art results on three domain adaptation benchmark datasets, which clearly validates the effectiveness of our proposed approach. Full code is available at https://github.com/kkkkkkon/D-MCD. Tong Chu, Yahao Liu, Jinhong Deng, Wen Li 0001, Lixin Duan |
AAAI | 5 |
| 2022 | Undoing the Damage of Label Shift for Cross-domain Semantic SegmentationabstractExisting works typically treat cross-domain semantic segmentation (CDSS) as a data distribution mismatch prob-lem and focus on aligning the marginal distribution or con-ditional distribution. However, the label shift issue is un-fortunately overlooked, which actually commonly exists in the CDSS task, and often causes a classifier bias in the learnt model. In this paper, we give an in-depth analysis and show that the damage of label shift can be overcome by aligning the data conditional distribution and correcting the posterior probability. To this end, we propose a novel approach to undo the damage of the label shift problem in CDSS. In implementation, we adopt class-level feature alignment for conditional distribution alignment, as well as two simple yet effective methods to rectify the classifier bias from source to target by remolding the classifier predictions. We conduct extensive experiments on the benchmark datasets of urban scenes, including GTA5 to Cityscapes and SYNTHIA to Cityscapes, where our proposed approach outperforms previous methods by a large margin. For instance, our model equipped with a self-training strat-egy reaches 59.3% mIoU on GTA5 to Cityscapes, pushing to a new state-of-the-art. The code will be available at https://github.com/manmanjun/Undoing_UDA. Yahao Liu, Jinhong Deng, Jiale Tao, Tong Chu, Lixin Duan, Wen Li 0001 |
CVPR | 5 |
| 2022 | Structure-Aware Motion Transfer with Deformable Anchor ModelabstractGiven a source image and a driving video depicting the same object type, the motion transfer task aims to generate a video by learning the motion from the driving video while preserving the appearance from the source image. In this paper, we propose a novel structure-aware motion modeling approach, the deformable anchor model (DAM), which can automatically discover the motion structure of arbitrary objects without leveraging their prior structure information. Specifically, inspired by the known deformable part model (DPM), our DAM introduces two types of anchors or key-points: i) a number of motion anchors that capture both appearance and motion information from the source image and driving video; ii) a latent root anchor, which is linked to the motion anchors to facilitate better learning of the representations of the object structure information. More-over, DAM can be further extended to a hierarchical version through the introduction of additional latent anchors to model more complicated structures. By regularizing motion anchors with latent anchor(s), DAM enforces the corre-spondences between them to ensure the structural information is well captured and preserved. Moreover, DAM can be learned effectively in an unsupervised manner. We validate our proposed DAM for motion transfer on different bench-mark datasets. Extensive experiments clearly demonstrate that DAM achieves superior performance relative to existing state-of-the-art methods. Jiale Tao, Borun Xu, Tiezheng Ge, Yuning Jiang 0001, Wen Li 0001, Lixin Duan |
CVPR | 7 |
| 2022 | Learning Pixel-Level Distinctions for Video Highlight DetectionabstractThe goal of video highlight detection is to select the most attractive segments from a long video to depict the most interesting parts of the video. Existing methods typically focus on modeling relationship between different video segments in order to learning a model that can assign highlight scores to these segments; however, these approaches do not explicitly consider the contextual dependency within individual segments. To this end, we propose to learn pixel-level distinctions to improve the video highlight detection. This pixel-level distinction indicates whether or not each pixel in one video belongs to an interesting section. The advantages of modeling such fine-level distinctions are two-fold. First, it allows us to exploit the temporal and spatial relations of the content in one video, since the distinction of a pixel in one frame is highly dependent on both the content before this frame and the content around this pixel in this frame. Second, learning the pixel-level distinction also gives a good explanation to the video highlight task regarding what contents in a highlight segment will be attractive to people. We design an encoder-decoder network to estimate the pixel-level distinction, in which we leverage the 3D convolutional neural networks to exploit the temporal context information, and further take advantage of the visual saliency to model the spatial distinction. State-of-the-art performance on three public benchmarks clearly validates the effectiveness of our framework for video highlight detection. Fanyue Wei, Tiezheng Ge, Yuning Jiang 0001, Wen Li 0001, Lixin Duan |
CVPR | 6 |
| 2022 | Motion Transformer for Unsupervised Image Animation
Jiale Tao, Tiezheng Ge, Yuning Jiang 0001, Wen Li 0001, Lixin Duan |
ECCV (16) | 6 |
| 2022 | Motion and Appearance Adaptation for Cross-domain Motion Transfer
Borun Xu, Jinhong Deng, Jiale Tao, Tiezheng Ge, Yuning Jiang 0001, Wen Li 0001, Lixin Duan |
ECCV (16) | 8 |
| 2022 | Diverse Preference Augmentation with Multiple Domains for Cold-start RecommendationsabstractCold-start issues have been more and more challenging for providing accurate recommendations with the fast increase of users and items. Most existing approaches attempt to solve the intractable problems via content-aware recommendations based on auxiliary information and/or cross-domain recommendations with transfer learning. Their performances are often constrained by the extremely sparse user-item interactions, unavailable side information, or very limited domain-shared users. Recently, meta-learners with meta-augmentation by adding noises to labels have been proven to be effective to avoid overfitting and shown good performance on new tasks. Motivated by the idea of meta-augmentation, in this paper, by treating a user's preference over items as a task, we propose a so-called Diverse Preference Augmentation framework with multiple source domains based on meta-learning (referred to as MetaDPA) to i) generate diverse ratings in a new domain of interest (known as target domain) to handle overfitting on the case of sparse interactions, and to ii) learn a preference model in the target domain via a meta-learning scheme to alleviate cold-start issues. Specifically, we first conduct multi-source domain adaptation by dual conditional variational autoencoders and impose a Multi-domain InfoMax (MDI) constraint on the latent representations to learn domain-shared and domain-specific preference properties. To avoid overfitting, we add a Mutually-Exclusive (ME) constraint on the output of decoders to generate diverse ratings given content data. Finally, these generated diverse ratings and the original ratings are introduced into the meta-training procedure to learn a preference meta-learner, which produces good generalization ability on cold-start recommendation tasks. Experiments on real-world datasets show our proposed MetaDPA clearly outperforms the current state-of-the-art baselines. Yan Zhang 0036, Changyu Li, Ivor W. Tsang, Lixin Duan, Hongzhi Yin, Wen Li 0001, Jie Shao 0001 |
ICDE | 5 |
| 2022 | Domain Adversarial Reinforcement Learning for Partial Domain AdaptationabstractPartial domain adaptation aims to transfer knowledge from a label-rich source domain to a label-scarce target domain (i.e., the target categories are a subset of the source ones), which relaxes the common assumption in traditional domain adaptation that the label space is fully shared across different domains. In this more general and practical scenario on partial domain adaptation, a major challenge is how to select source instances from the shared categories to ensure positive transfer for the target domain. To address this problem, we propose a domain adversarial reinforcement learning (DARL) framework to progressively select source instances to learn transferable features between domains by reducing the domain shift. Specifically, we employ a deep Q-learning to learn policies for an agent to make selection decisions by approximating the action-value function. Moreover, domain adversarial learning is introduced to learn a common feature subspace for the selected source instances and the target instances, and also to contribute to the reward calculation for the agent that is based on the relevance of the selected source instances with respect to the target domain. Extensive experiments on several benchmark data sets clearly demonstrate the superior performance of our proposed DARL over existing state-of-the-art methods for partial domain adaptation. Jin Chen 0009, Xinxiao Wu, Lixin Duan, Shenghua Gao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Unbiased Mean Teacher for Cross-Domain Object DetectionabstractCross-domain object detection is challenging, because object detection model is often vulnerable to data variance, especially to the considerable domain shift between two distinctive domains. In this paper, we propose a new Unbiased Mean Teacher (UMT) model for cross-domain object detection. We reveal that there often exists a considerable model bias for the simple mean teacher (MT) model in cross-domain scenarios, and eliminate the model bias with several simple yet highly effective strategies. In particular, for the teacher model, we propose a cross-domain distillation method for MT to maximally exploit the expertise of the teacher model. Moreover, for the student model, we alleviate its bias by augmenting training samples with pixel-level adaptation. Finally, for the teaching process, we employ an out-of-distribution estimation strategy to select samples that most fit the current model to further enhance the cross-domain distillation process. By tackling the model bias issue with these strategies, our UMT model achieves mAPs of 44.1%, 58.1%, 41.7%, and 43.1% on benchmark datasets Clipart1k, Watercolor2k, Foggy Cityscapes, and Cityscapes, respectively, which outperforms the existing state-of-the-art results in notable margins. Our implementation is available at https://github.com/kinredon/umt. Jinhong Deng, Wen Li 0001, Lixin Duan |
CVPR | 4 |
| 2021 | Progressive Modality Reinforcement for Human Multimodal Emotion Recognition From Unaligned Multimodal SequencesabstractHuman multimodal emotion recognition involves time-series data of different modalities, such as natural language, visual motions, and acoustic behaviors. Due to the variable sampling rates for sequences from different modalities, the collected multimodal streams are usually unaligned. The asynchrony across modalities increases the difficulty on conducting efficient multimodal fusion. Hence, this work mainly focuses on multimodal fusion from unaligned multimodal sequences. To this end, we propose the Progressive Modality Reinforcement (PMR) approach based on the recent advances of crossmodal transformer. Our approach introduces a message hub to exchange information with each modality. The message hub sends common messages to each modality and reinforces their features via crossmodal attention. In turn, it also collects the reinforced features from each modality and uses them to generate a reinforced common message. By repeating the cycle process, the common message and the modalities’ features can progressively complement each other. Finally, the reinforced features are used to make predictions for human emotion. Comprehensive experiments on different human multimodal emotion recognition benchmarks clearly demonstrate the superiority of our approach. Fengmao Lv, Yanyong Huang, Lixin Duan, Guosheng Lin |
CVPR | 4 |
| 2021 | BAPA-Net: Boundary Adaptation and Prototype Alignment for Cross-domain Semantic SegmentationabstractExisting cross-domain semantic segmentation methods usually focus on the overall segmentation results of whole objects but neglect the importance of object boundaries. In this work, we find that the segmentation performance can be considerably boosted if we treat object boundaries properly. For that, we propose a novel method called BAPA-Net, which is based on a convolutional neural network via Boundary Adaptation and Prototype Alignment, under the unsupervised domain adaptation setting. Specifically, we first construct additional images by pasting objects from source images to target images, and we develop a so-called boundary adaptation module to weigh each pixel based on its distance to the nearest boundary pixel of those pasted source objects. Moreover, we propose another prototype alignment module to reduce the domain mismatch by minimizing distances between the class prototypes of the source and target domains, where boundaries are removed to avoid domain confusion during prototype calculation. By integrating the boundary adaptation and prototype alignment, we are able to train a discriminative and domain-invariant model for cross-domain semantic segmentation. We conduct extensive experiments on the benchmark datasets of urban scenes (i.e., GTA5→Cityscapes and SYNTHIA→Cityscapes). And the promising results clearly show the effectiveness of our BAPA-Net method over existing state-of-the-art for cross-domain semantic segmentation. Our implementation is available at https://github.com/manmanjun/BAPA-Net. Yahao Liu, Jinhong Deng, Xinchen Gao, Wen Li 0001, Lixin Duan |
ICCV | 5 |
| 2021 | Multi-Scale Enhanced Active Learning for Skeleton-Based Action RecognitionabstractSkeleton-based models have been widely used, because of their robustness to complex backgrounds and high computational efficiency. However, annotating skeleton sequences is labor-intensive. It is appealing to reduce the cost of acquiring data with accurate labels for skeleton-based models. This paper presents an active learning method for the skeleton-based action recognition model, which boosts the performance of the model with less labeled data by instructing humans to annotate the most valuable samples. The key issue in active learning is to train a model that precisely predicts the values of samples. To achieve this, we propose to enhance the ability of our model to evaluate samples by modeling actions from different granularities from multi-scale representations of skeletons. The multi-scale method is simple and easy to use, which can be treated as a plug-and-play extension to strengthen the skeleton-based models. We conduct experiments on the SHREC, NTU-60, and Kinetics-Skeleton. Extensive experimental results demonstrate the effectiveness of the proposed method. Wen Li 0001, Lixin Duan |
ICME | 4 |
| 2021 | Counterfactual Debiasing Inference for Compositional Action RecognitionabstractCompositional action recognition is a novel challenge in the computer vision community and focuses on revealing the different combinations of verbs and nouns instead of treating subject-object interactions in videos as individual instances only. Existing methods tackle this challenging task by simply ignoring appearance information or fusing object appearances with dynamic instance tracklets. However, those strategies usually do not perform well for unseen action instances. For that, in this work we propose a novel learning framework called Counterfactual Debiasing Network (CDN) to improve the model generalization ability by removing the interference introduced by visual appearances of objects/subjects. It explicitly learns the appearance information in action representations and later removes the effect of such information in a causal inference manner. Specifically, we use tracklets and video content to model the factual inference by considering both appearance information and structure information. In contrast, only video content with appearance information is leveraged in the counterfactual inference. With the two inferences, we conduct a causal graph which captures and removes the bias introduced by the appearance information by subtracting the result of the counterfactual inference from that of the factual inference. By doing that, our proposed CDN method can better recognize unseen action instances by debiasing the effect of appearances. Extensive experiments on the Something-Else dataset clearly show the effectiveness of our proposed CDN over existing state-of-the-art methods. Pengzhan Sun 0001, Bo Wu 0018, Xunsong Li, Wen Li 0001, Lixin Duan, Chuang Gan 0001 |
ACM Multimedia | 5 |
| 2021 | Move As You Like: Image Animation in E-Commerce ScenarioabstractCreative image animations are attractive in e-commerce applications, where motion transfer is one of the import ways to generate animations from static images. However, existing methods rarely transfer motion to objects other than human body or human face, and even fewer apply motion transfer in practical scenarios. In this work, we apply motion transfer on the Taobao product images in real e-commerce scenario to generate creative animations, which are more attractive than static images and they will bring more benefits. We animate the Taobao products of dolls, copper running horses and toy dinosaurs based on motion transfer method for demonstration. Borun Xu, Jiale Tao, Tiezheng Ge, Yuning Jiang 0001, Wen Li 0001, Lixin Duan |
ACM Multimedia | 7 |
| 2021 | STST: Spatial-Temporal Specialized Transformer for Skeleton-based Action RecognitionabstractSkeleton-based action recognition has been widely investigated considering their strong adaptability to dynamic circumstances and complicated backgrounds. To recognize different actions from skeleton sequences, it is essential and crucial to model the posture of the human represented by the skeleton and its changes in the temporal dimension. However, most of the existing works treat skeleton sequences in the temporal and spatial dimension in the same way, ignoring the difference between the temporal and spatial dimension in skeleton data which is not an optimal way to model skeleton sequences. The posture represented by the skeleton in each frame is proposed to be modeled individually. Meanwhile, capturing the movement of the entire skeleton in the temporal dimension is needed. So, we designed Spatial Transformer Block and Directional Temporal Transformer Block for modeling skeleton sequences in spatial and temporal dimensions respectively. Due to occlusion/sensor/raw video, etc., there are noises on both temporal and spatial dimensions in the extracted skeleton data reducing the recognition capabilities of models. To adapt to this imperfect information condition, we propose a multi-task self-supervised learning method by providing confusing samples in different situations to improve the robustness of our model. Combining the above design, we propose our Spatial-Temporal Specialized Transformer~(STST) and conduct experiments with our model on the SHREC, NTU-RGB+D, and Kinetics-Skeleton. Extensive experimental results demonstrate the improved performances and analysis of the proposed method. Bo Wu 0018, Wen Li 0001, Lixin Duan, Chuang Gan 0001 |
ACM Multimedia | 4 |
| 2021 | Video Anomaly Detection with Sparse Coding Inspired Deep Neural NetworksabstractThis paper presents an anomaly detection method that is based on a sparse coding inspired Deep Neural Networks (DNN). Specifically, in light of the success of sparse coding based anomaly detection, we propose a Temporally-coherent Sparse Coding (TSC), where a temporally-coherent term is used to preserve the similarity between two similar frames. The optimization of sparse coefficients in TSC with the Sequential Iterative Soft-Thresholding Algorithm (SIATA) is equivalent to a special stacked Recurrent Neural Networks (sRNN) architecture. Further, to reduce the computational cost in alternatively updating the dictionary and sparse coefficients in TSC optimization and to alleviate hyperparameters selection in TSC, we stack one more layer on top of the TSC-inspired sRNN to reconstruct the inputs, and arrive at an sRNN-AE. We further improve sRNN-AE in the following aspects: i) rather than using a predefined similarity measurement between two frames, we propose to learn a data-dependent similarity measurement between neighboring frames in sRNN-AE to make it more suitable for anomaly detection; ii) to reduce computational costs in the inference stage, we reduce the depth of the sRNN in sRNN-AE and, consequently, our framework achieves real-time anomaly detection; iii) to improve computational efficiency, we conduct temporal pooling over the appearance features of several consecutive frames for summarizing information temporally, then we feed appearance features and temporally summarized features into a separate sRNN-AE for more robust anomaly detection. To facilitate anomaly detection evaluation, we also build a large-scale anomaly detection dataset which is even larger than the summation of all existing datasets for anomaly detection in terms of both the volume of data and the diversity of scenes. Extensive experiments on both a toy dataset under controlled settings and real datasets demonstrate that our method significantly outperforms existing methods, which validates the effectiveness of our sRNN-AE method for anomaly detection. Codes and data have been released at https://github.com/StevenLiuWen/sRNN_TSC_Anomaly_Detection. Weixin Luo, Wen Liu 0003, Dongze Lian, Jinhui Tang 0001, Lixin Duan, Xi Peng 0001, Shenghua Gao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Weakly-Supervised Cross-Domain Road Scene Segmentation via Multi-Level Curriculum AdaptationabstractSemantic segmentation, which aims to acquire pixel-level understanding about images, is among the key components in computer vision. To train a good segmentation model for real-world images, it usually requires a huge amount of time and labor effort to obtain sufficient pixel-level annotations of real-world images beforehand. To get rid of such a nontrivial burden, one can use simulators to automatically generate synthetic images that inherently contain full pixel-level annotations and use them to train a segmentation model for the real-world images. However, training with synthetic images usually cannot lead to good performance due to the domain difference between the synthetic images (i.e., source domain) and the real-world images (i.e., target domain). To deal with this issue, a number of unsupervised domain adaptation (UDA) approaches have been proposed, where no labeled real-world images are available. Different from those methods, in this work, we conduct a pioneer attempt by using easy-to-collect image-level annotations for target images to improve the performance of cross-domain segmentation. Specifically, we leverage those image-level annotations to construct curriculums for the domain adaptation problem. The curriculums describe multi-level properties of the target domain, including label distributions over full images, local regions and single pixels. Since image annotations are “weak” labels compared to pixel annotations for segmentation, we coin this new problem as weakly-supervised cross-domain segmentation. Comprehensive experiments on the GTA5→ Cityscapes and SYNTHIA→ Cityscapes settings demonstrate the effectiveness of our method over the existing state-of-the-art baselines. Fengmao Lv, Guosheng Lin, Peng Liu 0049, Guowu Yang, Sinno Jialin Pan, Lixin Duan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Sequential Instance Refinement for Cross-Domain Object Detection in ImagesabstractCross-domain object detection in images has attracted increasing attention in the past few years, which aims at adapting the detection model learned from existing labeled images (source domain) to newly collected unlabeled ones (target domain). Existing methods usually deal with the cross-domain object detection problem through direct feature alignment between the source and target domains at the image level, the instance level (i.e., region proposals) or both. However, we have observed that directly aligning features of all object instances from the two domains often results in the problem of negative transfer, due to the existence of (1) outlier target instances that contain confusing objects not belonging to any category of the source domain and thus are hard to be captured by detectors and (2) low-relevance source instances that are considerably statistically different from target instances although their contained objects are from the same category. With this in mind, we propose a reinforcement learning based method, coined as sequential instance refinement, where two agents are learned to progressively refine both source and target instances by taking sequential actions to remove both outlier target instances and low-relevance source instances step by step. Extensive experiments on several benchmark datasets demonstrate the superior performance of our method over existing state-of-the-art baselines for cross-domain object detection. Jin Chen 0009, Xinxiao Wu, Lixin Duan, Lin Chen 0021 |
IEEE Trans. Image Process. | 3 |
| 2021 | Adversarial Multimodal Network for Movie Story Question AnsweringabstractVisual question answering by using information from multiple modalities has attracted more and more attention in recent years. However, it is a very challenging task, as the visual content and natural language have quite different statistical properties. In this work, we present a method called Adversarial Multimodal Network (AMN) to better understand video stories for question answering. In AMN, we propose to learn multimodal feature representations by finding a more coherent subspace for video clips and the corresponding texts (e.g., subtitles and questions) based on generative adversarial networks. Moreover, a self-attention mechanism is developed to enforce our newly introduced consistency constraint in order to preserve the self-correlation between the visual cues of the original video clips in the learned multimodal representations. Extensive experiments on the benchmark MovieQA and TVQA datasets show the effectiveness of our proposed AMN over other published state-of-the-art methods. Zhaoquan Yuan, Lixin Duan, Xiao Wu 0001, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2020 | Dynamic and Static Context-Aware LSTM for Multi-agent Motion Prediction
Chaofan Tao, Qinhong Jiang, Lixin Duan, Ping Luo 0002 |
ECCV (21) | 3 |
| 2020 | Temporal Action Proposal Generation via Multi-Task Feature LearningabstractTemporal action proposal generation is an active topic in computer vision and image/video processing communities, which aims to predict a set of temporal proposals to discover all action instances with high recall and intersection over union (IoU) in real-world untrimmed videos. Many previous approaches rely on the two-stream network which intends to simultaneously extract spatial (appearance) and temporal features for representing each video. However, few work considers to capture the correlations between the two kinds of features, leaving a large room for improving the model. In this paper, we present a new method for generating temporal action proposals based on multi-task feature learning. Specifically, we aim to learn shared representation between the spatial and temporal features in a multi-task learning framework, so as to acquire a compact and precise feature representation. Moreover, we devise a correlation loss to address the `weak-correlation' problem with high IoUs but low confidences cores. Finally, we take an ensemble learning strategy in order to inherit the advantages of existing works. Extensive experimental results on the ActivityNet-1.3 challenge dataset show that the proposed method achieves the best performance, compared with the state-of-the-arts reported in the literature and the official leaderboard. Our code will be released soon. Handong Ma, Lixin Duan |
VCIP | 2 |
| 2020 | Reconstruction Regularized Deep Metric Learning for Multi-Label Image ClassificationabstractIn this paper, we present a novel deep metric learning method to tackle the multi-label image classification problem. In order to better learn the correlations among images features, as well as labels, we attempt to explore a latent space, where images and labels are embedded via two unique deep neural networks, respectively. To capture the relationships between image features and labels, we aim to learn a two-way deep distance metric over the embedding space from two different views, i.e., the distance between one image and its labels is not only smaller than those distances between the image and its labels' nearest neighbors but also smaller than the distances between the labels and other images corresponding to the labels' nearest neighbors. Moreover, a reconstruction module for recovering correct labels is incorporated into the whole framework as a regularization term, such that the label embedding space is more representative. Our model can be trained in an end-to-end manner. Experimental results on publicly available image data sets corroborate the efficacy of our method compared with the state of the arts. Chong Liu 0003, Lixin Duan, Peng Gao 0015, Kai Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Learning Transferable Self-Attentive Representations for Action Recognition in Untrimmed Videos with Weak SupervisionabstractAction recognition in videos has attracted a lot of attention in the past decade. In order to learn robust models, previous methods usually assume videos are trimmed as short sequences and require ground-truth annotations of each video frame/sequence, which is quite costly and time-consuming. In this paper, given only video-level annotations, we propose a novel weakly supervised framework to simultaneously locate action frames as well as recognize actions in untrimmed videos. Our proposed framework consists of two major components. First, for action frame localization, we take advantage of the self-attention mechanism to weight each frame, such that the influence of background frames can be effectively eliminated. Second, considering that there are trimmed videos publicly available and also they contain useful information to leverage, we present an additional module to transfer the knowledge from trimmed videos for improving the classification performance in untrimmed ones. Extensive experiments are conducted on two benchmark datasets (i.e., THUMOS14 and ActivityNet1.3), and experimental results clearly corroborate the efficacy of our method. Xiaoyu Zhang 0002, Haichao Shi, Kai Zheng 0001, Xiaobin Zhu 0001, Lixin Duan |
AAAI | 6 |
| 2019 | Constructing Self-Motivated Pyramid Curriculums for Cross-Domain Semantic Segmentation: A Non-Adversarial ApproachabstractWe propose a new approach, called self-motivated pyramid curriculum domain adaptation (PyCDA), to facilitate the adaptation of semantic segmentation neural networks from synthetic source domains to real target domains. Our approach draws on an insight connecting two existing works: curriculum domain adaptation and self-training. Inspired by the former, PyCDA constructs a pyramid curriculum which contains various properties about the target domain. Those properties are mainly about the desired label distributions over the target domain images, image regions, and pixels. By enforcing the segmentation neural network to observe those properties, we can improve the network's generalization capability to the target domain. Motivated by the self-training, we infer this pyramid of properties by resorting to the semantic segmentation network itself. Unlike prior work, we do not need to maintain any additional models (e.g., logistic regression or discriminator networks) or to solve minmax problems which are often difficult to optimize. We report state-of-the-art results for the adaptation from both GTAV and SYNTHIA to Cityscapes, two popular settings in unsupervised domain adaptation for semantic segmentation. Qing Lian, Lixin Duan, Fengmao Lv, Boqing Gong |
ICCV | 2 |
| 2019 | TarGAN: Generating target data with class labels for unsupervised domain adaptation
Fengmao Lv, Guowu Yang, Lixin Duan |
Knowl. Based Syst. | 4 |
| 2019 | Exploiting Images for Video Recognition: Heterogeneous Feature Augmentation via Symmetric Adversarial LearningabstractTraining deep models of video recognition usually requires sufficient labeled videos in order to achieve good performance without over-fitting. However, it is quite labor-intensive and time-consuming to collect and annotate a large amount of videos. Moreover, training deep neural networks on large-scale video datasets always demands huge computational resources which further hold back many researchers and practitioners. To resolve that, collecting and training on annotated images are much easier. However, thoughtlessly applying images to help recognize videos may result in noticeable performance degeneration due to the well-known domain shift and feature heterogeneity. This proposes a novel symmetric adversarial learning approach for heterogeneous image-to-video adaptation, which augments deep image and video features by learning domain-invariant representations of source images and target videos. Primarily focusing on an unsupervised scenario where the labeled source images are accompanied by unlabeled target videos in the training phrase, we present a data-driven approach to respectively learn the augmented features of images and videos with superior transformability and distinguishability. Starting with learning a common feature space (called image-frame feature space) between images and video frames, we then build new symmetric generative adversarial networks (Sym-GANs) where one GAN maps image-frame features to video features and the other maps video features to image-frame features. Using the Sym-GANs, the source image feature is augmented with the generated video-specific representation to capture the motion dynamics while the target video feature is augmented with the image-specific representation to take the static appearance information. Finally, the augmented features from the source domain are fed into a network with fully connected layers for classification. Thanks to an end-to-end training procedure of the Sym-GANs and the classification network, our approach achieves better results than other state-of-the-arts, which is clearly validated by experiments on two video datasets, i.e., the UCF101 and HMDB51 datasets. Feiwu Yu, Xinxiao Wu, Lixin Duan |
IEEE Trans. Image Process. | 4 |
| 2019 | Multiview Multitask Gaze Estimation With Deep Convolutional Neural NetworksabstractGaze estimation, which aims to predict gaze points with given eye images, is an important task in computer vision because of its applications in human visual attention understanding. Many existing methods are based on a single camera, and most of them only focus on either the gaze point estimation or gaze direction estimation. In this paper, we propose a novel multitask method for the gaze point estimation using multiview cameras. Specifically, we analyze the close relationship between the gaze point estimation and gaze direction estimation, and we use a partially shared convolutional neural networks architecture to simultaneously estimate the gaze direction and gaze point. Furthermore, we also introduce a new multiview gaze tracking data set that consists of multiview eye images of different subjects. As far as we know, it is the largest multiview gaze tracking data set. Comprehensive experiments on our multiview gaze tracking data set and existing data sets demonstrate that our multiview multitask gaze point estimation solution consistently outperforms existing methods. Dongze Lian, Lina Hu, Weixin Luo, Yanyu Xu 0001, Lixin Duan, Jingyi Yu 0001, Shenghua Gao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | Exploiting Images for Video Recognition with Hierarchical Generative Adversarial NetworksabstractExisting deep learning methods of video recognition usually require a large number of labeled videos for training. But for a new task, videos are often unlabeled and it is also time-consuming and labor-intensive to annotate them. Instead of human annotation, we try to make use of existing fully labeled images to help recognize those videos. However, due to the problem of domain shifts and heterogeneous feature representations, the performance of classifiers trained on images may be dramatically degraded for video recognition tasks. In this paper, we propose a novel method, called Hierarchical Generative Adversarial Networks (HiGAN), to enhance recognition in videos (i.e., target domain) by transferring knowledge from images (i.e., source domain). The HiGAN model consists of a \emph{low-level} conditional GAN and a \emph{high-level} conditional GAN. By taking advantage of these two-level adversarial learning, our method is capable of learning a domain-invariant feature representation of source images and target videos. Comprehensive experiments on two challenging video recognition datasets (i.e. UCF101 and HMDB51) demonstrate the effectiveness of the proposed method when compared with the existing state-of-the-art domain adaptation methods. Feiwu Yu, Xinxiao Wu, Yuchao Sun, Lixin Duan |
IJCAI | 4 |
| 2017 | Learning User Dependencies for RecommendationabstractSocial recommender systems exploit users' social relationships to improve recommendation accuracy. Intuitively, a user tends to trust different people regarding with different scenarios. Therefore, one main challenge of social recommendation is to exploit the most appropriate dependencies between users for a given recommendation task. Previous social recommendation methods are usually developed based on pre-defined user dependencies. Thus, they may not be optimal for a specific recommendation task. In this paper, we propose a novel recommendation method, named probabilistic relational matrix factorization (PRMF), which can automatically learn the dependencies between users to improve recommendation accuracy. In PRMF, users' latent features are assumed to follow a matrix variate normal (MVN) distribution. Both positive and negative user dependencies can be modeled by the row precision matrix of the MVN distribution. Moreover, we also propose an alternating optimization algorithm to solve the optimization problem of PRMF. Extensive experiments on four real datasets have been performed to demonstrate the effectiveness of the proposed PRMF model. Yong Liu 0020, Peilin Zhao, Xin Liu 0027, Min Wu 0008, Lixin Duan, Xiaoli Li 0001 |
IJCAI | 5 |
| 2017 | Action and Event Recognition in Videos by Learning From Heterogeneous Web SourcesabstractIn this paper, we propose new approaches for action and event recognition by leveraging a large number of freely available Web videos (e.g., from Flickr video search engine) and Web images (e.g., from Bing and Google image search engines). We address this problem by formulating it as a new multi-domain adaptation problem, in which heterogeneous Web sources are provided. Specifically, we are given different types of visual features (e.g., the DeCAF features from Bing/Google images and the trajectory-based features from Flickr videos) from heterogeneous source domains and all types of visual features from the target domain. Considering the target domain is more relevant to some source domains, we propose a new approach named multi-domain adaptation with heterogeneous sources (MDA-HS) to effectively make use of the heterogeneous sources. In MDA-HS, we simultaneously seek for the optimal weights of multiple source domains, infer the labels of target domain samples, and learn an optimal target classifier. Moreover, as textual descriptions are often available for both Web videos and images, we propose a novel approach called MDA-HS using privileged information (MDA-HS+) to effectively incorporate the valuable textual information into our MDA-HS method, based on the recent learning using privileged information paradigm. MDA-HS+ can be further extended by using a new elastic-net-like regularization. We solve our MDA-HS and MDA-HS+ methods by using the cutting-plane algorithm, in which a multiple kernel learning problem is derived and solved. Extensive experiments on three benchmark data sets demonstrate that our proposed approaches are effective for action and event recognition without requiring any labeled samples from the target domain. Li Niu 0002, Xinxing Xu, Lin Chen 0021, Lixin Duan, Dong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2016 | Webly-Supervised Video Recognition by Mutually Voting for Relevant Web Images and Web Video Frames
Chuang Gan 0001, Chen Sun 0002, Lixin Duan, Boqing Gong |
ECCV (3) | 3 |
| 2016 | Axial Alignment for Anterior Segment Swept Source Optical Coherence Tomography via Robust Low-Rank Tensor RecoveryabstractWe present a one-step approach based on low-rank tensor recovery for axial alignment in 360-degree anterior chamber optical coherence tomography. Achieving translational alignment and rotation correction of cross-sections simultaneously, this technique obtains a better anterior segment topographical representation and improves quantitative measurement accuracy and reproducibility of disease related parameters. Through its use of global information, the proposed method is more robust compared to using only individual or paired slices, and less sensitive to noise and motion artifacts. In angle closure analysis on 30 patient eyes, the preliminary results indicate that the proposed axial alignment method can not only facilitate manual qualitative analysis with more distinct landmark representation and much less human labor, but also can improve the accuracy of automatic quantitative assessment by 2.9 %, which demonstrates that the proposed approach is promising for a wide range of clinical applications. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Yanwu Xu 0001, Lixin Duan, Huazhu Fu, Xiaoqin Zhang 0002, Damon Wing Kee Wong, Mani Baskaran, Tin Aung, Jiang Liu 0001 |
MICCAI (3) | 2 |
| 2016 | Semantic Reconstruction-Based Nuclear Cataract Grading from Slit-Lamp Lens ImagesabstractCataracts are the leading cause of visual impairment and blindness worldwide. Cataract grading, i.e. assessing the presence and severity of cataracts, is essential for diagnosis and progression monitoring. We present in this work an automatic method for predicting cataract grades from slit-lamp lens images. Different from existing techniques which normally formulate cataract grading as a regression problem, we solve it through reconstruction-based classification, which has been shown to yield higher performance when the available training data is densely distributed within the feature space. To heighten the effectiveness of this reconstruction-based approach, we introduce a new semantic feature representation that facilitates alignment of test and reference images, and include locality constraints on the linear reconstruction to reduce the influence of less relevant reference samples. In experiments on the large ACHIKO-NC database comprised of 5378 images, our system outperforms the state-of-the-art regression methods over a range of evaluation metrics. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Yanwu Xu 0001, Lixin Duan, Damon Wing Kee Wong, Tien Yin Wong, Jiang Liu 0001 |
MICCAI (3) | 2 |
| 2016 | DEFEATnet - A Deep Conventional Image Representation for Image ClassificationabstractTo study underlying possibilities for the successes of conventional image representation and deep neural networks (DNNs) in image representation, we propose a deep feature extraction, encoding, and pooling network (DEFEATnet) architecture, which is a marriage between conventional image representation approaches and DNNs. In particular, in DEFEATnet, each layer consists of three components: feature extraction, feature encoding, and pooling. The primary advantage of DEFEATnet is twofold. First, it consolidates the prior knowledge (e.g., translation invariance) from extracting, encoding, and pooling handcrafted features, as in the conventional feature representation approaches. Second, it represents the object parts at different granularities by gradually increasing the local receptive fields in different layers, as in DNNs. Moreover, DEFEATnet is a generalized framework that can readily incorporate all types of local features as well as all kinds of well-designed feature encoding and pooling methods. Since prior knowledge is preserved in DEFEATnet, it is especially useful for image representation on small/medium size data sets, where DNNs usually fail due to the lack of sufficient training data. Promising experimental results clearly show that DEFEATnets outperform shallow conventional image representation approaches by a large margin when the same type of features, feature encoding and pooling are used. The extensive experiments also demonstrate the effectiveness of the deep architecture of our DEFEATnet in improving the robustness for image presentation. Shenghua Gao, Lixin Duan, Ivor W. Tsang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Multiple Ocular Diseases Classification with Graph Regularized Probabilistic Multi-label Learning
Yanwu Xu 0001, Lixin Duan, Shuicheng Yan, Zhuo Zhang 0001, Damon Wing Kee Wong, Jiang Liu 0001 |
ACCV (4) | 3 |
| 2014 | Speckle Reduction in Optical Coherence Tomography by Image Registration and Matrix Completion
Jun Cheng 0003, Lixin Duan, Damon Wing Kee Wong, Dacheng Tao, Masahiro Akiba, Jiang Liu 0001 |
MICCAI (1) | 2 |
| 2014 | Incorporating Privileged Genetic Information for Fundus Image Based Glaucoma Detection
Lixin Duan, Yanwu Xu 0001, Wen Li 0001, Lin Chen 0021, Damon Wing Kee Wong, Tien Yin Wong, Jiang Liu 0001 |
MICCAI (2) | 1 |
| 2014 | Optic Cup Segmentation for Glaucoma Detection Using Low-Rank Superpixel Representation
Yanwu Xu 0001, Lixin Duan, Stephen Lin 0001, Damon Wing Kee Wong, Tien Yin Wong, Jiang Liu 0001 |
MICCAI (1) | 2 |
| 2014 | Learning With Augmented Features for Supervised and Semi-Supervised Heterogeneous Domain AdaptationabstractIn this paper, we study the heterogeneous domain adaptation (HDA) problem, in which the data from the source domain and the target domain are represented by heterogeneous features with different dimensions. By introducing two different projection matrices, we first transform the data from two domains into a common subspace such that the similarity between samples across different domains can be measured. We then propose a new feature mapping function for each domain, which augments the transformed samples with their original features and zeros. Existing supervised learning methods (e.g., SVM and SVR) can be readily employed by incorporating our newly proposed augmented feature representations for supervised HDA. As a showcase, we propose a novel method called Heterogeneous Feature Augmentation (HFA) based on SVM. We show that the proposed formulation can be equivalently derived as a standard Multiple Kernel Learning (MKL) problem, which is convex and thus the global solution can be guaranteed. To additionally utilize the unlabeled data in the target domain, we further propose the semi-supervised HFA (SHFA) which can simultaneously learn the target classifier as well as infer the labels of unlabeled target samples. Comprehensive experiments on three different applications clearly demonstrate that our SHFA and HFA outperform the existing HDA methods. Wen Li 0001, Lixin Duan, Dong Xu 0001, Ivor W. Tsang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Hybrid constraint SVR for facial age estimation
Lixin Duan, Yuehu Liu |
Signal Process. | 3 |
| 2013 | Event Recognition in Videos by Learning from Heterogeneous Web SourcesabstractIn this work, we propose to leverage a large number of loosely labeled web videos (e.g., from YouTube) and web images (e.g., from Google/Bing image search) for visual event recognition in consumer videos without requiring any labeled consumer videos. We formulate this task as a new multi-domain adaptation problem with heterogeneous sources, in which the samples from different source domains can be represented by different types of features with different dimensions (e.g., the SIFT features from web images and space-time (ST) features from web videos) while the target domain samples have all types of features. To effectively cope with the heterogeneous sources where some source domains are more relevant to the target domain, we propose a new method called Multi-domain Adaptation with Heterogeneous Sources (MDA-HS) to learn an optimal target classifier, in which we simultaneously seek the optimal weights for different source domains with different types of features as well as infer the labels of unlabeled target domain data based on multiple types of features. We solve our optimization problem by using the cutting-plane algorithm based on group based multiple kernel learning. Comprehensive experiments on two datasets demonstrate the effectiveness of MDA-HS for event recognition in consumer videos. Lin Chen 0021, Lixin Duan, Dong Xu 0001 |
CVPR | 2 |
| 2013 | Action Recognition Using Multilevel Features and Latent Structural SVMabstractWe first propose a new low-level visual feature, called spatio-temporal context distribution feature of interest points, to describe human actions. Each action video is expressed as a set of relative XYT coordinates between pairwise interest points in a local region. We learn a global Gaussian mixture model (GMM) (referred to as a universal background model) using the relative coordinate features from all the training videos, and then we represent each video as the normalized parameters of a video-specific GMM adapted from the global GMM. In order to capture the spatio-temporal relationships at different levels, multiple GMMs are utilized to describe the context distributions of interest points over multiscale local regions. Motivated by the observation that some actions share similar motion patterns, we additionally propose a novel mid-level class correlation feature to capture the semantic correlations between different action classes. Each input action video is represented by a set of decision values obtained from the pre-learned classifiers of all the action classes, with each decision value measuring the likelihood that the input video belongs to the corresponding action class. Moreover, human actions are often associated with some specific natural environments and also exhibit high correlation with particular scene classes. It is therefore beneficial to utilize the contextual scene information for action recognition. In this paper, we build the high-level co-occurrence relationship between action classes and scene classes to discover the mutual contextual constraints between action and scene. By treating the scene class label as a latent variable, we propose to use the latent structural SVM (LSSVM) model to jointly capture the compatibility between multilevel action features (e.g., low-level visual context distribution feature and the corresponding mid-level class correlation feature) and action classes, the compatibility between multilevel scene features (i.e., SIFT feature and the corresponding class correlation feature) and scene classes, and the contextual relationship between action classes and scene classes. Extensive experiments on UCF Sports, YouTube and UCF50 datasets demonstrate the effectiveness of the proposed multilevel features and action-scene interaction based LSSVM model for human action recognition. Moreover, our method generally achieves higher recognition accuracy than other state-of-the-art methods on these datasets. Xinxiao Wu, Dong Xu 0001, Lixin Duan, Jiebo Luo 0001, Yunde Jia |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Efficient Discriminative Learning of Class Hierarchy for Many Class Prediction
Lin Chen 0021, Lixin Duan, Ivor W. Tsang, Dong Xu 0001 |
ACCV (1) | 2 |
| 2012 | Exploiting web images for event recognition in consumer videos: A multiple source domain adaptation approachabstractRecent work has demonstrated the effectiveness of domain adaptation methods for computer vision applications. In this work, we propose a new multiple source domain adaptation method called Domain Selection Machine (DSM) for event recognition in consumer videos by leveraging a large number of loosely labeled web images from different sources (e.g., Flickr.com and Photosig.com), in which there are no labeled consumer videos. Specifically, we first train a set of SVM classifiers (referred to as source classifiers) by using the SIFT features of web images from different source domains. We propose a new parametric target decision function to effectively integrate the static SIFT features from web images/video keyframes and the spacetime (ST) features from consumer videos. In order to select the most relevant source domains, we further introduce a new data-dependent regularizer into the objective of Support Vector Regression (SVR) using the ∊-insensitive loss, which enforces the target classifier shares similar decision values on the unlabeled consumer videos with the selected source classifiers. Moreover, we develop an alternating optimization algorithm to iteratively solve the target decision function and a domain selection vector which indicates the most relevant source domains. Extensive experiments on three real-world datasets demonstrate the effectiveness of our proposed method DSM over the state-of-the-art by a performance gain up to 46.41%. Lixin Duan, Dong Xu 0001, Shih-Fu Chang |
CVPR | 1 |
| 2012 | Batch mode Adaptive Multiple Instance Learning for computer vision tasksabstractMultiple Instance Learning (MIL) has been widely exploited in many computer vision tasks, such as image retrieval, object tracking and so on. To handle ambiguity of instance labels in positive bags, the training process of traditional MIL methods is usually computationally expensive, which limits the applications of MIL in more computer vision tasks. In this paper, we propose a novel batch mode framework, namely Batch mode Adaptive Multiple Instance Learning (BAMIL), to accelerate the instance-level MIL methods. Specifically, instead of using all training bags at once, we divide the training bags into several sets of bags (i.e., batches). At each time, we use one batch of training bags to train a new classifier which is adapted from the latest pre-learned classifier. Such batch mode framework significantly accelerates the traditional MIL methods for large scale applications and can be also used in dynamic environments such as object tracking. The experimental results show that our BAMIL is much faster than the recently developed MIL with constrained positive bags while achieves comparable performance for text-based web image retrieval. In dynamic settings, BAMIL also achieves the better overall performance for object tracking when compared with other online MIL methods. Wen Li 0001, Lixin Duan, Ivor W. Tsang, Dong Xu 0001 |
CVPR | 2 |
| 2012 | Co-labeling: A New Multi-view Learning Approach for Ambiguous ProblemsabstractWe propose a multi-view learning approach called co-labeling which is applicable for several machine learning problems where the labels of training samples are uncertain, including semi-supervised learning (SSL), multi-instance learning (MIL) and max-margin clustering (MMC). Particularly, we first unify those problems into a general ambiguous problem in which we simultaneously learn a robust classifier as well as find the optimal training labels from a finite label candidate set. To effectively utilize multiple views of data, we then develop our co-labeling approach for the general multi-view ambiguous problem. In our work, classifiers trained on different views can teach each other by iteratively passing the predictions of training samples from one classifier to the others. The predictions from one classifier are considered as label candidates for the other classifiers. To train a classifier with a label candidate set for each view, we adopt the Multiple Kernel Learning (MKL) technique by constructing the base kernel through associating the input kernel calculated from input features with one label candidate. Compared with the traditional co-training method which was specifically designed for SSL, the advantages of our co-labeling are two-fold: 1) it can be applied to other ambiguous problems such as MIL and MMC, 2) it is more robust by using the MKL method to integrate multiple labeling candidates obtained from different iterations and biases. Promising results on several real-world multi-view data sets clearly demonstrate the effectiveness of our proposed co-labeling for both MIL and SSL. Wen Li 0001, Lixin Duan, Ivor W. Tsang, Dong Xu 0001 |
ICDM | 2 |
| 2012 | Learning with Augmented Features for Heterogeneous Domain Adaptation
Lixin Duan, Dong Xu 0001, Ivor W. Tsang |
ICML | 1 |
| 2012 | Domain Transfer Multiple Kernel LearningabstractCross-domain learning methods have shown promising results by leveraging labeled patterns from the auxiliary domain to learn a robust classifier for the target domain which has only a limited number of labeled samples. To cope with the considerable change between feature distributions of different domains, we propose a new cross-domain kernel learning framework into which many existing kernel methods can be readily incorporated. Our framework, referred to as Domain Transfer Multiple Kernel Learning (DTMKL), simultaneously learns a kernel function and a robust classifier by minimizing both the structural risk functional and the distribution mismatch between the labeled and unlabeled samples from the auxiliary and target domains. Under the DTMKL framework, we also propose two novel methods by using SVM and prelearned classifiers, respectively. Comprehensive experiments on three domain adaptation data sets (i.e., TRECVID, 20 Newsgroups, and email spam data sets) demonstrate that DTMKL-based methods outperform existing cross-domain learning and multiple kernel learning methods. Lixin Duan, Ivor W. Tsang, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Visual Event Recognition in Videos by Learning from Web DataabstractWe propose a visual event recognition framework for consumer videos by leveraging a large amount of loosely labeled web videos (e.g., from YouTube). Observing that consumer videos generally contain large intraclass variations within the same type of events, we first propose a new method, called Aligned Space-Time Pyramid Matching (ASTPM), to measure the distance between any two video clips. Second, we propose a new transfer learning method, referred to as Adaptive Multiple Kernel Learning (A-MKL), in order to 1) fuse the information from multiple pyramid levels and features (i.e., space-time features and static SIFT features) and 2) cope with the considerable variation in feature distributions between videos from two domains (i.e., web video domain and consumer video domain). For each pyramid level and each type of local features, we first train a set of SVM classifiers based on the combined training set from two domains by using multiple base kernels from different kernel types and parameters, which are then fused with equal weights to obtain a prelearned average classifier. In A-MKL, for each event class we learn an adapted target classifier based on multiple base kernels and the prelearned average classifiers from this event class or all the event classes by minimizing both the structural risk functional and the mismatch between data distributions of two domains. Extensive experiments demonstrate the effectiveness of our proposed framework that requires only a small number of labeled consumer videos by leveraging web data. We also conduct an in-depth investigation on various aspects of the proposed method A-MKL, such as the analysis on the combination coefficients on the prelearned classifiers, the convergence of the learning algorithm, and the performance variation by using different proportions of labeled consumer videos. Moreover, we show that A-MKL using the prelearned classifiers from all the event classes leads to better performance when compared with A-MK- using the prelearned classifiers only from each individual event class. Lixin Duan, Dong Xu 0001, Ivor W. Tsang, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Domain Adaptation From Multiple Sources: A Domain-Dependent Regularization ApproachabstractIn this paper, we propose a new framework called domain adaptation machine (DAM) for the multiple source domain adaption problem. Under this framework, we learn a robust decision function (referred to as target classifier) for label prediction of instances from the target domain by leveraging a set of base classifiers which are prelearned by using labeled instances either from the source domains or from the source domains and the target domain. With the base classifiers, we propose a new domain-dependent regularizer based on smoothness assumption, which enforces that the target classifier shares similar decision values with the relevant base classifiers on the unlabeled instances from the target domain. This newly proposed regularizer can be readily incorporated into many kernel methods (e.g., support vector machines (SVM), support vector regression, and least-squares SVM (LS-SVM)). For domain adaptation, we also develop two new domain adaptation methods referred to as FastDAM and UniverDAM. In FastDAM, we introduce our proposed domain-dependent regularizer into LS-SVM as well as employ a sparsity regularizer to learn a sparse target classifier with the support vectors only from the target domain, which thus makes the label prediction on any test instance very fast. In UniverDAM, we additionally make use of the instances from the source domains as Universum to further enhance the generalization ability of the target classifier. We evaluate our two methods on the challenging TRECIVD 2005 dataset for the large-scale video concept detection task as well as on the 20 newsgroups and email spam datasets for document retrieval. Comprehensive experiments demonstrate that FastDAM and UniverDAM outperform the existing multiple source domain adaptation methods for the two applications. Lixin Duan, Dong Xu 0001, Ivor W. Tsang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2011 | Action recognition using context and appearance distribution featuresabstractWe first propose a new spatio-temporal context distribution feature of interest points for human action recognition. Each action video is expressed as a set of relative XYT coordinates between pairwise interest points in a local region. We learn a global GMM (referred to as Universal Background Model, UBM) using the relative coordinate features from all the training videos, and then represent each video as the normalized parameters of a video-specific GMM adapted from the global GMM. In order to capture the spatio-temporal relationships at different levels, multiple GMMs are utilized to describe the context distributions of interest points over multi-scale local regions. To describe the appearance information of an action video, we also propose to use GMM to characterize the distribution of local appearance features from the cuboids centered around the interest points. Accordingly, an action video can be represented by two types of distribution features: 1) multiple GMM distributions of spatio-temporal context; 2) GMM distribution of local video appearance. To effectively fuse these two types of heterogeneous and complementary distribution features, we additionally propose a new learning algorithm, called Multiple Kernel Learning with Augmented Features (AFMKL), to learn an adapted classifier based on multiple kernels and the pre-learned classifiers of other action classes. Extensive experiments on KTH, multi-view IXMAS and complex UCF sports datasets demonstrate that our method generally achieves higher recognition accuracy than other state-of-the-art methods. Xinxiao Wu, Dong Xu 0001, Lixin Duan, Jiebo Luo 0001 |
CVPR | 3 |
| 2011 | Text-based image retrieval using progressive multi-instance learningabstractRelevant and irrelevant images collected from the Web (e.g., Flickr.com) have been employed as loosely labeled training data for image categorization and retrieval. In this work, we propose a new approach to learn a robust classifier for text-based image retrieval (TBIR) using relevant and irrelevant training web images, in which we explicitly handle noise in the loose labels of training images. Specifically, we first partition the relevant and irrelevant training web images into clusters. By treating each cluster as a "bag" and the images in each bag as "instances", we formulate this task as a multi-instance learning problem with constrained positive bags, in which each positive bag contains at least a portion of positive instances. We present a new algorithm called MIL-CPB to effectively exploit such constraints on positive bags and predict the labels of test instances (images). Observing that the constraints on positive bags may not always be satisfied in our application, we additionally propose a progressive scheme (referred to as Progressive MIL-CPB, or PMIL-CPB) to further improve the retrieval performance, in which we iteratively partition the top-ranked training web images from the current MIL-CPB classifier to construct more confident positive "bags "and then add these new "bags" as training data to learn the subsequent MIL-CPB classifiers. Comprehensive experiments on two challenging real-world web image data sets demonstrate the effectiveness of our approach. © 2011 IEEE. Wen Li 0001, Lixin Duan, Dong Xu 0001, Ivor W. Tsang |
ICCV | 2 |
| 2011 | Improving Web Image Search by Bag-Based RerankingabstractGiven a textual query in traditional text-based image retrieval (TBIR), relevant images are to be reranked using visual features after the initial text-based search. In this paper, we propose a new bag-based reranking framework for large-scale TBIR. Specifically, we first cluster relevant images using both textual and visual features. By treating each cluster as a "bag" and the images in the bag as "instances," we formulate this problem as a multi-instance (MI) learning problem. MI learning methods such as mi-SVM can be readily incorporated into our bag-based reranking framework. Observing that at least a certain portion of a positive bag is of positive instances while a negative bag might also contain positive instances, we further use a more suitable generalized MI (GMI) setting for this application. To address the ambiguities on the instance labels in the positive and negative bags under this GMI setting, we develop a new method referred to as GMI-SVM to enhance retrieval performance by propagating the labels from the bag level to the instance level. To acquire bag annotations for (G)MI learning, we propose a bag ranking method to rank all the bags according to the defined bag ranking score. The top ranked bags are used as pseudopositive training bags, while pseudonegative training bags can be obtained by randomly sampling a few irrelevant images that are not associated with the textual query. Comprehensive experiments on the challenging real-world data set NUS-WIDE demonstrate our framework with automatic bag annotation can achieve the best performances compared with existing image reranking methods. Our experiments also demonstrate that GMI-SVM can achieve better performances when using the manually labeled training bags obtained from relevance feedback. Lixin Duan, Wen Li 0001, Ivor W. Tsang, Dong Xu 0001 |
IEEE Trans. Image Process. | 1 |
| 2010 | Visual event recognition in videos by learning from web dataabstractWe propose a visual event recognition framework for consumer domain videos by leveraging a large amount of loosely labeled web videos (e.g., from YouTube). First, we propose a new aligned space-time pyramid matching method to measure the distances between two video clips, where each video clip is divided into space-time volumes over multiple levels. We calculate the pair-wise distances between any two volumes and further integrate the information from different volumes with Integer-flow Earth Mover's Distance (EMD) to explicitly align the volumes. Second, we propose a new cross-domain learning method in order to 1) fuse the information from multiple pyramid levels and features (i.e., space-time feature and static SIFT feature) and 2) cope with the considerable variation in feature distributions between videos from two domains (i.e., web domain and consumer domain). For each pyramid level and each type of local features, we train a set of SVM classifiers based on the combined training set from two domains using multiple base kernels of different kernel types and parameters, which are fused with equal weights to obtain an average classifier. Finally, we propose a cross-domain learning method, referred to as Adaptive Multiple Kernel Learning (A-MKL), to learn an adapted classifier based on multiple base kernels and the prelearned average classifiers by minimizing both the structural risk functional and the mismatch between data distributions from two domains. Extensive experiments demonstrate the effectiveness of our proposed framework that requires only a small number of labeled consumer videos by leveraging web data. Lixin Duan, Dong Xu 0001, Ivor W. Tsang, Jiebo Luo 0001 |
CVPR | 1 |
| 2010 | Near Duplicate Identification With Spatially Aligned Pyramid MatchingabstractA new framework, termed spatially aligned pyramid matching, is proposed for near duplicate image identification. The proposed method robustly handles spatial shifts as well as scale changes, and is extensible for video data. Images are divided into both overlapped and non-overlapped blocks over multiple levels. In the first matching stage, pairwise distances between blocks from the examined image pair are computed using earth mover's distance (EMD) or the visual word with$\chi^{2}$distance based method with scale-invariant feature transform (SIFT) features. In the second stage, multiple alignment hypotheses that consider piecewise spatial shifts and scale variation are postulated and resolved using integer-flow EMD. Moreover, to compute the distances between two videos, we conduct the third step matching (i.e., temporal matching) after spatial matching. Two application scenarios are addressed—near duplicate retrieval (NDR) and near duplicate detection (NDD). For retrieval ranking, a pyramid-based scheme is constructed to fuse matching results from different partition levels. For NDD, we also propose a dual-sample approach by using the multilevel distances as features and support vector machine for binary classification. The proposed methods are shown to clearly outperform existing methods through extensive testing on the Columbia Near Duplicate Image Database and two new datasets. In addition, we also discuss in depth our framework in terms of the extension for video NDR and NDD, the sensitivity to parameters, the utilization of multiscale dense SIFT descriptors, and the test of scalability in image NDD. Dong Xu 0001, Tat-Jen Cham, Shuicheng Yan, Lixin Duan, Shih-Fu Chang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | Domain Transfer SVM for video concept detectionabstractCross-domain learning methods have shown promising results by leveraging labeled patterns from auxiliary domains to learn a robust classifier for target domain, which has a limited number of labeled samples. To cope with the tremendous change of feature distribution between different domains in video concept detection, we propose a new cross-domain kernel learning method. Our method, referred to as Domain Transfer SVM (DTSVM), simultaneously learns a kernel function and a robust SVM classifier by minimizing both the structural risk functional of SVM and the distribution mismatch of labeled and unlabeled samples between the auxiliary and target domains. Comprehensive experiments on the challenging TRECVID corpus demonstrate that DTSVM outperforms existing cross-domain learning and multiple kernel learning methods. Lixin Duan, Ivor W. Tsang, Dong Xu 0001, Stephen J. Maybank |
CVPR | 1 |
| 2009 | Domain adaptation from multiple sources via auxiliary classifiersabstractWe propose a multiple source domain adaptation method, referred to as Domain Adaptation Machine (DAM), to learn a robust decision function (referred to as target classifier) for label prediction of patterns from the target domain by leveraging a set of pre-computed classifiers (referred to as auxiliary/source classifiers) independently learned with the labeled patterns from multiple source domains. We introduce a new data-dependent regularizer based on smoothness assumption into Least-Squares SVM (LS-SVM), which enforces that the target classifier shares similar decision values with the auxiliary classifiers from relevant source domains on the unlabeled patterns of the target domain. In addition, we employ a sparsity regularizer to learn a sparse target classifier. Comprehensive experiments on the challenging TRECVID 2005 corpus demonstrate that DAM outperforms the existing multiple source domain adaptation methods for video concept detection in terms of effectiveness and efficiency. Lixin Duan, Ivor W. Tsang, Dong Xu 0001, Tat-Seng Chua |
ICML | 1 |