EDBT 2026 Demo / reviewers in the wild / expert
Wen Li 0001
dblp:06/721-1
· DBLP profile ↗
116ranked-venue papers
9as first author
72since 2021 · last 2026
0000-0002-5559-8594ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 83 · 8 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 83 · 4 first-author · 54 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unsupervised Robust Domain Adaptation: Paradigm, Theory and Algorithm
Fuxiang Huang, Xiaowei Fu, Shiyu Ye, Wen Li 0001, Xinbo Gao 0001, David Zhang 0001, Lei Zhang 0038 |
Int. J. Comput. Vis. | 5 |
| 2026 | PMDAv2: Multi-scale prototype matching for domain adaptive semantic segmentation
Weiwei Li 0005, Yuchen Zheng 0001, Yuanyuan Ren, Junzhuo Liu 0002, Yahao Liu, Wen Li 0001 |
Pattern Recognit. | 6 |
| 2026 | ROOT: Region-Word Alignment With Partial Optimal Transport for Open-Vocabulary Object DetectionabstractOpen-vocabulary object detection (OVD) aims to detect novel object concepts by mining region-word correspondences from image-text pairs, yet current methods often produce false correspondences. While some strategies (e.g., one-to-one matching) were proposed to mitigate this issue, they often sacrifice numerous valuable region-word pairs during the matching process. To overcome these challenges, we propose a novel comprehensive alignment method, named Region-word Alignment with Partial Optimal Transport (ROOT) framework, which reframes the region-word matching task as a problem of partial distribution alignment. Unlike traditional optimal transport, which shifts the full mass of the distribution, partial optimal transport enables selective matching, making it more robust to noise in region and word alignment. Specifically, ROOT first employs partial optimal transport to obtain an optimal transport plan for region and word feature alignment. This transport plan is then used to compute a matching reliability score for each region-word pair, which reweights the contrastive alignment loss to enhance accuracy. By enabling more flexible and reliable region-text matches, ROOT significantly reduces misalignment errors while preserving valuable region-word correspondences. Extensive experiments on standard benchmarks OV-COCO and OV-LVIS show that our ROOT outperforms the previous state-of-the-art works, demonstrating the effectiveness of our approach. Jinhong Deng, Yinjie Lei, Wen Li 0001, Lixin Duan |
IEEE Trans. Image Process. | 3 |
| 2026 | Lighted-SAM: Lightening Open-World SAM for Low-Light SegmentationabstractSegment Anything Model (SAM) has achieved impressive segmentation performance in an open-world setting. However, SAM relies heavily on high-quality input images and usually struggles in low-light conditions. This is mainly caused by the pre-training dataset, SA-1B, in which low-light samples constitute a relatively small fraction of the data. This lack of presence leads to a noticeable weakness when SAM is applied in real-world dark environments. With the motivation of improving SAM's performance under low-light conditions while retaining its strong zero-shot capability, this work proposes an alignment stage between the pre-training stage and testing stage. Unlike existing low-light studies that mainly focus on task-specific and close-set settings, our work further emphasizes pursuing the segmentation ability under low-light conditions for open-world models. To this end, we construct DarkSeg58K, a realistic and diverse dataset, which serves as the alignment dataset to support this stage. We further introduce Lighted-SAM as the lightweight repair strategy to fix SAM's performance in low-light conditions. Different from existing methods focusing on introducing spectral adapters into the model design and training this model end-to-end, Lighted-SAM introduces the Spectral Information Resonance (SIR) mechanism to harmoniously integrate the spectral enhancement module into SAM, which is usually kept frozen due to its large-scale parameters. Based on our lightweight repairing strategy, Lighted-SAM can improve SAM's ability in low-light conditions while preserving its zero-shot ability. Experiments on different benchmarks validate the superiority of our approach. Code is available at: https://github.com/Jaaaahan/LightedSAM. Yuhan Jia, Lixin Duan, Wen Li 0001, Fengmao Lv |
IEEE Trans. Image Process. | 3 |
| 2026 | Tuning-Free Adaptive Style Incorporation for Structure-Consistent Text-Driven Style TransferabstractText-driven style transfer methods leveraging diffusion models have shown impressive creativity, yet they still face challenges in maintaining consistent structure and content preservation. Existing methods often directly concatenate the content and style prompts for a prompt-level style injection. However, this coarse-grained style injection strategy inevitably leads to structural deviations in the stylized images. This poses a significant obstacle for professional artists and creators seeking precise artistic editing. In this work, we strive to attain a harmonious balance between content preservation and style transformation. We propose Adaptive Style Incorporation (ASI), to achieve fine-grained feature-level style incorporation. It consists of the Siamese Cross-Attention (SiCA) to decouple the single-track cross-attention to a dual-track structure to obtain separate content and style features, and the Adaptive Content-Style Blending (AdaBlending) module to couple the content and style information from a structure-consistent manner. Experimentally, our method exhibits much better performance in both structure preservation and stylized effects. Yanqi Ge, Jiaqi Liu 0004, Qingnan Fan, Xi Jiang 0009, Shuai Qin, Wen Li 0001, Lixin Duan |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2025 | S-INF: Towards Realistic Indoor Scene Synthesis via Scene Implicit Neural FieldabstractLearning-based methods have become increasingly popular in 3D indoor scene synthesis (ISS), showing superior performance over traditional optimization-based approaches. These learning-based methods typically model distributions on simple yet explicit scene representations using generative models. However, due to the oversimplified explicit representations that overlook detailed information and the lack of guidance from multimodal relationships within the scene, most learning-based methods struggle to generate indoor scenes with realistic object arrangements and styles. In this paper, we introduce a new method, Scene Implicit Neural Field (S-INF), for indoor scene synthesis, aiming to learn meaningful representations of multimodal relationships, to enhance the realism of indoor scene synthesis. S-INF assumes that the scene layout is often related to the object-detailed information. It disentangles the multimodal relationships into scene layout relationships and detailed object relationships, fusing them later through implicit neural fields (INFs). By learning specialized scene layout relationships and projecting them into S-INF, we achieve a realistic generation of scene layout. Additionally, S-INF captures dense and detailed object relationships through differentiable rendering, ensuring stylistic consistency across objects. Through extensive experiments on the benchmark 3D-FRONT dataset, we demonstrate that our method consistently achieves state-of-the-art performance under different types of ISS. Zixi Liang, Haifeng Wu, Wen Li 0001, Lixin Duan |
AAAI | 5 |
| 2025 | Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of SamplesabstractDeep learning models are known to often learn features that spuriously correlate with the class label during training but are irrelevant to the prediction task. Existing methods typically address this issue by annotating potential spurious attributes, or filtering spurious features based on some empirical assumptions (e.g., simplicity of bias). However, these methods may yield unsatisfactory performance due to the intricate and elusive nature of spurious correlations in real-world data. In this paper, we propose a data-oriented approach1to mitigate the spurious correlation in deep learning models. We observe that samples that are influenced by spurious features tend to exhibit a dispersed distribution in the learned feature space. This allows us to identify the presence of spurious features. Subsequently, we obtain a bias-invariant representation by neutralizing the spurious features based on a simple grouping strategy. Then, we learn a feature transformation to eliminate the spurious features by aligning with this bias-invariant representation. Finally, we update the classifier by incorporating the learned feature transformation and obtain an unbiased model. By integrating the aforementioned identifying, neutralizing, eliminating and updating procedures, we build an effective pipeline for mitigating spurious correlation. Experiments on image and NLP debiasing benchmarks show an improvement in worst group accuracy of more than 20% compared to standard empirical risk minimization (ERM). Weiwei Li 0005, Junzhuo Liu 0002, Yuanyuan Ren, Yuchen Zheng 0001, Yahao Liu, Wen Li 0001 |
CVPR | 6 |
| 2025 | Learned Image Compression with Dictionary-based Entropy ModelabstractLearned image compression methods have attracted great research interest and exhibited superior rate-distortion performance to the best classical image compression standards of the present. The entropy model plays a key role in learned image compression, which estimates the probability distribution of the latent representation for further entropy coding. Most existing methods employed hyper-prior and auto-regressive architectures to form their entropy models. However, they only aimed to explore the internal dependencies of latent representation while neglecting the importance of extracting prior from training data. In this work, we propose a novel entropy model named Dictionary-Based Cross Attention Entropy model, which introduces a learnable dictionary to summarize the typical structures occurring in the training dataset to enhance the entropy model. Extensive experimental results have demonstrated that the proposed model strikes a better balance between performance and latency, achieving state-of-the-art results on various benchmark datasets. Jingbo Lu, Leheng Zhang, Mu Li 0005, Wen Li 0001, Shuhang Gu |
CVPR | 5 |
| 2025 | GeoDepth: From Point-to-Depth to Plane-to-Depth Modeling for Self-Supervised Monocular Depth EstimationabstractSelf-supervised monocular depth estimation has long been treated as a point-wise prediction problem, where the depth of each pixel is usually estimated independently. However, artifacts are often observed in the estimated depth map, e.g., depth values for points located in the same region may jump dramatically. To address this issue, we propose a novel self-supervised monocular depth estimation framework called GeoDepth, where we explore the intrinsic geometric representation in 3D scenes for producing accurate and continuous depth maps. In particular, we model the complex 3D scene as a collection of planes with varying sizes, where each plane is characterized by a unique set of parameters, namely planar normal (indicating plane orientation) and planar offset (defining the perpendicular distance from the camera center to the plane). Under this modeling, points in the same plane are enforced to share a unique representation and their depth variations related only to pixel coordinates, thus this geometric relationship can be exploited to regularize the depth variations of these points. To this end, we design a structured plane generation module that introduces spatio-temporal geometric cues and the plane uniqueness principle to recover the correct scene plane representation. In addition, we develop a depth discontinuity module to identify depth discontinuity regions and subsequently optimize them. Our experiments on the KITTI and NYUv2 datasets demonstrate that GeoDepth achieves state-of-the-art performance, with additional tests on Make3D and ScanNet validating its generalization capabilities. Haifeng Wu, Shuhang Gu, Lixin Duan, Wen Li 0001 |
CVPR | 4 |
| 2025 | ResCLIP: Residual Attention for Training-free Dense Vision-language InferenceabstractWhile vision-language models like CLIP have shown remarkable success in open-vocabulary tasks, their application is currently confined to image-level tasks, and they still struggle with dense predictions. Recent works often attribute such deficiency in dense predictions to the self-attention layers in the final block, and have achieved commendable results by modifying the original query-key attention to self-correlation attention, (e.g., query-query and key-key attention). However, these methods overlook the cross-correlation attention (query-key) properties, which capture the rich spatial correspondence. In this paper, we reveal that the cross-correlation of self-attention in non-final layers of CLIP also exhibits localization properties. Therefore, we propose the Residual Cross-correlation Self-attention (RCS) module, which leverages the cross-correlation self-attention from intermediate layers to remold the attention in the final block. The RCS module effectively reorganizes spatial information, unleashing the localization potential within CLIP for dense vision-language inference. Furthermore, to enhance the focus on regions of the same categories and local consistency, we propose the Semantic Feedback Refinement (SFR) module, which utilizes semantic segmentation maps to further adjust the attention scores. By integrating these two strategies, our method, termed ResCLIP, can be easily incorporated into existing approaches as a plug-and-play module, significantly boosting their performance in dense vision-language inference. Extensive experiments across multiple standard benchmarks demonstrate that our method surpasses state-of-the-art training-free methods, validating the effectiveness of the proposed approach. Code is available at https://github.com/yvhangyang/ResCLIP. Jinhong Deng, Wen Li 0001, Lixin Duan |
CVPR | 3 |
| 2025 | Balanced Sharpness-Aware Minimization for Imbalanced Regression
Yahao Liu, Qin Wang 0013, Lixin Duan, Wen Li 0001 |
ICCV | 4 |
| 2025 | PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Model
Hualian Sheng, Sijia Cai, Bing Deng, Qiao Liang 0002, Wen Li 0001, Jieping Ye, Shuhang Gu |
ICCV | 6 |
| 2025 | The Devil Is in the Spurious Correlations: Boosting Moment Retrieval With Dynamic LearningabstractGiven a textual query along with a corresponding video, the objective of moment retrieval aims to localize the moments relevant to the query within the video. While commendable results have been demonstrated by existing transformer-based approaches, predicting the accurate temporal span of the target moment is still a major challenge. This paper reveals that a crucial reason stems from the spurious correlation between the text query and the moment context. Namely, the model makes predictions by overly associating queries with background frames rather than distinguishing target moments. To address this issue, we propose a dynamic learning approach for moment retrieval, where two strategies are designed to mitigate the spurious correlation. First, we introduce a novel video synthesis approach to construct a dynamic context for the queried moment, enabling the model to attend to the target moment of the corresponding query across dynamic backgrounds. Second, to alleviate the over-association with backgrounds, we enhance representations temporally by incorporating text-dynamics interaction, which encourages the model to align text with target moments through complementary dynamic representations. With the proposed method, our model significantly alleviates the spurious correlation issue in moment retrieval and establishes new state-of-the-art performance on two popular benchmarks, \ie, QVHighlights and Charades-STA. In addition, detailed ablation studies and evaluations across different architectures demonstrate the generalization and effectiveness of the proposed strategies. Our code will be publicly available. Xinyang Zhou, Fanyue Wei, Lixin Duan, Angela Yao, Wen Li 0001 |
ICCV | 5 |
| 2025 | P2WNet: Homography Estimation for Part-To-Whole and Cross-Modality ScenariosabstractDeep learning-based homography estimation has achieved remarkable advances in recent years. However, existing methods face limitations in the Part-To-Whole (P2W) scenario, where the template image corresponds to a small portion of the search image, as their designs are tailored to image pairs with similar content and limited displacement. To address this issue, we propose P2WNet, a novel framework for part-to-whole and cross-modality homography estimation. First, we tailor a pseudo-siamese encoder to handle cross-modal inputs and incorporate a transformer-based cascade for feature enhancement. Furthermore, we design a novel P2W matching module to capture and represent the correspondences between image pairs in the Part-To-Whole scenario. These robust features are fed into a prediction module to estimate the homography matrix. Additionally, we propose a dataset to validate our model in P2W scenario. Experiments on both our dataset and public benchmarks (DroneVehicle, GoogleMap, MSCOCO) demonstrate that P2WNet achieves superior performance in the P2W scenario and performs competitively in conventional scenarios. Code is available at https://github.com/xuanxh1/P2WNET. ShangXuan Xie, Haifeng Wu, Wen Li 0001, Lixin Duan |
ICME | 3 |
| 2025 | SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMsabstractMultimodal Large Language Models (MLLMs) typically process a large number of visual tokens, leading to considerable computational overhead, even though many of these tokens are redundant. Existing visual token pruning methods primarily focus on selecting the most salient tokens based on attention scores, resulting in the semantic incompleteness of the selected tokens. In this paper, we propose a novel visual token pruning strategy, called **S**aliency-**C**overage **O**riented token **P**runing for **E**fficient MLLMs (SCOPE), to jointly model both the saliency and coverage of the selected visual tokens to better preserve semantic completeness. Specifically, we introduce a set-coverage for a given set of selected tokens, computed based on the token relationships. We then define a token-coverage gain for each unselected token, quantifying how much additional coverage would be obtained by including it. By integrating the saliency score into the token-coverage gain, we propose our SCOPE score and iteratively select the token with the highest SCOPE score. We conduct extensive experiments on multiple vision-language understanding benchmarks using the LLaVA-1.5 and LLaVA-Next models. Experimental results demonstrate that our method consistently outperforms prior approaches. Jinhong Deng, Wen Li 0001, Joey Tianyi Zhou, Yang He 0002 |
NeurIPS | 2 |
| 2025 | Domain Adaptive Detection of MAVs: A Benchmark and Noise Suppression NetworkabstractVisual detection of Micro Air Vehicles (MAVs) has attracted increasing attention in recent years due to its important application in various tasks. The existing methods for MAV detection assume that the training set and testing set have the same distribution. As a result, when deployed in new domains, the detectors would have a significant performance degradation due to domain discrepancy. In this paper, we study the problem of cross-domain MAV detection. The contributions of this paper are threefold. 1) We propose a Multi-MAV-Multi-Domain (M3D) dataset consisting of both simulation and realistic images. Compared to other existing datasets, the proposed one is more comprehensive in the sense that it covers rich scenes, diverse MAV types, and various viewing angles. A new benchmark for cross-domain MAV detection is proposed based on the proposed dataset. 2) We propose a Noise Suppression Network (NSN) based on the framework of pseudo-labeling and a large-to-small training procedure. To reduce the challenging pseudo-label noises, two novel modules are designed in this network. The first is a prior-based curriculum learning module for allocating adaptive thresholds for pseudo labels with different difficulties. The second is a masked copy-paste augmentation module for pasting truly-labeled MAVs on unlabeled target images and thus decreasing pseudo-label noises. 3) Extensive experimental results verify the superior performance of the proposed method compared to the state-of-the-art ones. In particular, it achieves mAP of 46.9%(+5.8%), 50.5%(+3.7%), and 61.5%(+11.3%) on the tasks of simulation-to-real adaptation, cross-scene adaptation, and cross-camera adaptation, respectively.Note to Practitioners— To study the cross-domain MAV detection problem, this paper establishes a novel benchmark that consists of three domain adaptation tasks: simulation-to-real adaptation, cross-scene adaptation, and cross-camera adaptation, respectively. The benchmark is based on a novel MAV dataset called Multi-MAV-Multi-Domain (M3D), which is available at: https://github.com/WestlakeAerialRobotics/M3D. To reduce the noises caused by pseudo labels, a noise suppression network is proposed to overcome the error accumulation. Extensive experiments are conducted to prove the effectiveness of the proposed approach. Jinhong Deng, Peidong Liu 0001, Wen Li 0001, Shiyu Zhao 0002 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Simultaneous Detection and Interaction Reasoning for Object-Centric Action RecognitionabstractThe interactions between human and objects are important for recognizing object-centric actions. Existing methods usually adopt a two-stage pipeline, where object proposals are first detected using a pretrained detector, and then are fed to an action recognition model for extracting video features and learning the object relations for action recognition. However, since the action prior is unknown in the object detection stage, important objects could be easily overlooked, leading to inferior action recognition performance. In this paper, we propose an end-to-end object-centric action recognition framework that simultaneously performsDetectionAndInteractionReasoning (dubbed DAIR) in one stage. Particularly, after extracting video features using a base network, we design three consecutive modules for simultaneously learning object detection and interaction reasoning. Firstly, we build a Patch-based Object Decoder (PatchDec) to generate object proposals from video patch tokens. Then, we design an Interactive Object Refining and Aggregation (IRA) to identify the interactive objects that are important for action recognition. The IRA module adjusts the interactiveness scores of proposals based on their relative position and appearance, and aggregates the object-level information into global video representation. Finally, we build an Object Relation Modeling (ORM) module to encode the object relations. These three modules together with the video feature extractor can be trained jointly in an end-to-end fashion, thus avoiding the heavy reliance on an off-the-shelf object detector, and reducing the multi-stage training burden. We conduct experiments on two datasets, Something-Else and Ikea-Assembly, to evaluate the performance of our proposed approach on conventional, compositional, and few-shot action recognition tasks. Through in-depth experimental analysis, we show the crucial role ofinteractiveobjects in learning for action recognition, and we can outperform state-of-the-art methods on both datasets. We hope our DAIR can provide a new perspective for object-centric action recognition. Xunsong Li, Pengzhan Sun 0001, Yangcen Liu, Lixin Duan, Wen Li 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Segmenting Anything in the Dark via Depth PerceptionabstractImage segmentation under low-light conditions is essential in real-world applications, such as autonomous driving and video surveillance systems. The recent Segment Anything Model (SAM) exhibits strong segmentation capability in various vision applications. However, its performance could be severely degraded under low-light conditions. On the other hand, multimodal information has been exploited to help models construct more comprehensive understanding of scenes under low-light conditions by providing complementary information (e.g., depth). Therefore, in this work, we present a pioneer attempt that elevates a unimodal vision foundation model (e.g., SAM) to a multimodal one, by efficiently integrating additional depth information under low-light conditions. To achieve that, we propose a novel method called Depth Perception SAM (DPSAM) based on the SAM framework. Specifically, we design a modality encoder to extract the depth information and the Depth Perception Layers (DPLs) for mutual feature refinement between RGB and depth features. The DPLs employ the cross-modal attention mechanism to mutually query effective information from both RGB and depth for the subsequent feature refinement. Thus, DPLs can effectively leverage the complementary information from depth to enrich the RGB representations and obtain comprehensive multimodal visual representations for segmenting anything in the dark. To this end, our DPSAM maximally maintains the instinct expertise of SAM for RGB image segmentation and further leverages on the strength of depth for enhanced segmenting anything capability, especially for cases that are likely to fail with RGB only (e.g., low-light or complex textures). As demonstrated by extensive experiments on four RGBD benchmark datasets, DPSAM clearly improves the performance for the segmenting anything performance in the dark, e.g., +12.90% mIoU and +16.23% mIoU on LLRGBD and DeLiVER, respectively. Our code and model will be made publicly available at:https://github.com/liupeng3425/DPSAM. Peng Liu 0049, Jinhong Deng, Lixin Duan, Wen Li 0001, Fengmao Lv |
IEEE Trans. Multim. | 4 |
| 2025 | Evidential Deep Learning for Open-Set Active Domain AdaptationabstractOpen-set domain adaptation (OSDA) seeks to transfer knowledge from a labeled source domain to an unlabeled target domain containing novel classes. Traditional OSDA methods rarely account for the uncertainty in predictions and typically require additional training overhead. Evidential deep learning (EDL) transforms the model's predictions from point estimates to distributions over the probability simplex by replacing the standard softmax output of classification neural networks with Dirichlet distributions. Considering the presence of out-of-distribution novel classes in OSDA and the additional overhead of existing methods, we propose EDL for open-set active domain adaptation (EOSADA). Leveraging EDL, we construct an open-set classifier and employ a two-round selection strategy guided by the data uncertainty of target domain samples and semantic similarity scores with known classes. This strategy balances the selection of samples from known and novel classes while identifying informative samples, thereby maximizing the performance of the model in OSDA scenarios without modifying the model structure and utilizing a limited annotation budget. Extensive experiments demonstrate the superiority of our approach. Qing Tian 0001, Jiangsen Yu, Wen Li 0001, Zhen Lei 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Beyond Prototypes: Semantic Anchor Regularization for Better Representation LearningabstractOne of the ultimate goals of representation learning is to achieve compactness within a class and well-separability between classes. Many outstanding metric-based and prototype-based methods following the Expectation-Maximization paradigm, have been proposed for this objective. However, they inevitably introduce biases into the learning process, particularly with long-tail distributed training data. In this paper, we reveal that the class prototype is not necessarily to be derived from training features and propose a novel perspective to use pre-defined class anchors serving as feature centroid to unidirectionally guide feature learning. However, the pre-defined anchors may have a large semantic distance from the pixel features, which prevents them from being directly applied. To address this issue and generate feature centroid independent from feature learning, a simple yet effective Semantic Anchor Regularization (SAR) is proposed. SAR ensures the inter-class separability of semantic anchors in the semantic space by employing a classifier-aware auxiliary cross-entropy loss during training via disentanglement learning. By pulling the learned features to these semantic anchors, several advantages can be attained: 1) the intra-class compactness and naturally inter-class separability, 2) induced bias or errors from feature learning can be avoided, and 3) robustness to the long-tailed problem. The proposed SAR can be used in a plug-and-play manner in the existing models. Extensive experiments demonstrate that the SAR performs better than previous sophisticated prototype-based methods. The implementation is available at https://github.com/geyanqi/SAR. Yanqi Ge, Qiang Nie, Yong Liu 0020, Chengjie Wang 0001, Feng Zheng 0001, Wen Li 0001, Lixin Duan |
AAAI | 7 |
| 2024 | Beyond Viewpoint: Robust 3D Object Recognition Under Arbitrary Views Through Joint Multi-part Representation
Linlong Fan, Yanqi Ge, Wen Li 0001, Lixin Duan |
ECCV (52) | 4 |
| 2024 | Powerful and Flexible: Personalized Text-to-Image Generation via Reinforcement Learning
Fanyue Wei, Wei Zeng 0008, Dawei Yin 0001, Lixin Duan, Wen Li 0001 |
ECCV (27) | 6 |
| 2024 | Learning Semantic Latent Directions for Accurate and Controllable Human Motion Prediction
Jiale Tao, Wen Li 0001, Lixin Duan |
ECCV (21) | 3 |
| 2024 | Towards Unsupervised Model Selection for Domain Adaptive Object DetectionabstractEvaluating the performance of deep models in new scenarios has drawn increasing attention in recent years due to the wide application of deep learning techniques in various fields. However, while it is possible to collect data from new scenarios, the annotations are not always available. Existing Domain Adaptive Object Detection (DAOD) works usually report their performance by selecting the best model on the validation set or even the test set of the target domain, which is highly impractical in real-world applications. In this paper, we propose a novel unsupervised model selection approach for domain adaptive object detection, which is able to select almost the optimal model for the target domain without using any target labels. Our approach is based on the flat minima principle, i.e., models located in the flat minima region in the parameter space usually exhibit excellent generalization ability. However, traditional methods require labeled data to evaluate how well a model is located in the flat minima region, which is unrealistic for the DAOD task. Therefore, we design a Detection Adaptation Score (DAS) approach to approximately measure the flat minima without using target labels. We show via a generalization bound that the flatness can be deemed as model variance, while the minima depend on the domain distribution distance for the DAOD task. Accordingly, we propose a Flatness Index Score (FIS) to assess the flatness by measuring the classification and localization fluctuation before and after perturbations of model parameters and a Prototypical Distance Ratio (PDR) score to seek the minima by measuring the transferability and discriminability of the models. In this way, the proposed DAS approach can effectively represent the degree of flat minima and evaluate the model generalization ability on the target domain. We have conducted extensive experiments on various DAOD benchmarks and approaches, and the experimental results show that the proposed DAS correlates well with the performance of DAOD models and can be used as an effective tool for model selection after training. The code will be released at https://github.com/HenryYu23/DAS. Hengfu Yu, Jinhong Deng, Wen Li 0001, Lixin Duan |
NeurIPS | 3 |
| 2024 | Adversarial Neon Beam: A light-based physical attack to DNNs
Chengyin Hu, Weiwen Shi, Ling Tian, Wen Li 0001 |
Comput. Vis. Image Underst. | 4 |
| 2024 | Deep graph tensor learning for temporal link prediction
Zhen Liu 0006, Wen Li 0001, Lixin Duan |
Inf. Sci. | 3 |
| 2024 | Feature Re-Representation and Reliable Pseudo Label Retraining for Cross-Domain Semantic SegmentationabstractThis paper presents a novel unsupervised domain adaptation method for semantic segmentation. We argue that a good representation of the target-domain data should keep both the knowledge from the source domain and the target-domain-specific information. To obtain the knowledge from the source domain, we first learn a set of bases to characterize the feature distribution of the source domain, then features from both the source and the target domain are re-represented as a weighted summation of the source bases. A discriminator is additionally introduced to make the re-representation responsibilities of both domain features under the same bases indistinguishable. In this way, the domain gap between the source re-representation and target re-representation is minimized, and the re-represented target domain features contain the source domain information. Then we combine the feature re-representation with the original domain-specific feature together for subsequent pixel-wise classification. To further make the re-represented target features semantically meaningful, a Reliable Pseudo Label Retraining (RPLR) strategy is proposed, which utilizes the consistency of the prediction by the networks trained with multi-view source images to select the clean pseudo labels on unlabeled target images for re-training. Extensive experiments demonstrate the competitive performance of our approach for unsupervised domain adaptation on the semantic segmentation benchmarks. Jing Li 0117, Kang Zhou 0001, Shenhan Qian, Wen Li 0001, Lixin Duan, Shenghua Gao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Balanced Teacher for Source-Free Object DetectionabstractWe study a practical domain adaptation task, named source-free object detection (SFOD), which aims to adapt a pre-trained source detector to an unlabeled target domain without access to the original labeled source domain samples. In this paper, we design a new self-training approach for SFOD called Balance Teacher based on the mean teacher model. We target two key issues when using self-training for SFOD: 1) imbalanced label distribution when using pseudo-labels for supervising the model training, and 2) imbalanced image distribution,i.e., significant data variance in the target domain. To address these issues, we first design a Class-balanced Instance Selection (CBIS) module to automatically balance different classes when selecting pseudo-labeled instances during the training process. Then, we propose a Progressive Target Variance Minimization (PTVM) to cope with the imbalanced image distribution in the target domain, where the feature distributions of certainty and uncertainty target samples are progressively aligned to alleviate the data distribution variance. In this way, the teacher model can provide high-quality pseudo-labels and guide the student model to adapt gradually to the target domain. We have conducted extensive experiments on five widely used benchmarks, and the experimental results clearly show the superiority of our method over the state-of-the-art baselines. Jinhong Deng, Wen Li 0001, Lixin Duan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | High-Level Feature Guided Decoding for Semantic SegmentationabstractExisting pyramid-based upsamplers (e.g. SemanticFPN), although efficient, usually produce less accurate results compared to dilation-based models when using the same backbone. This is partially caused by thecontaminatedhigh-level features since they are fused and fine-tuned with noisy low-level features on limited data. To address this issue, we propose to use powerful pre-trainedhigh-levelfeatures asguidance (HFG) so that the upsampler can produce robust results. Specifically,onlythe high-level features from the backbone are used to train the class tokens, which are then reused by the upsampler for classification, guiding the upsampler features to more discriminative backbone features. One crucial design of the HFG is to protect the high-level features from being contaminated by using proper stop-gradient operations so that the backbone does not update according to the noisy gradient from the upsampler. To push the upper limit of HFG, we introduce acontextaugmentationencoder (CAE) that can efficiently and effectively operate on the low-resolution high-level feature, resulting in improved representation and thus better guidance. We named our complete solution as the High-Level Features Guided Decoder (HFGD). We evaluate the proposed HFGD on three benchmarks: Pascal Context, COCOStuff164k, and Cityscapes. HFGD achieves state-of-the-art results among methods that do not use extra training data, demonstrating its effectiveness and generalization ability. Shenghua Gao, Wen Li 0001, Lixin Duan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | CAFA: Cross-Modal Attentive Feature Alignment for Cross-Domain Urban Scene SegmentationabstractAutonomous driving systems rely heavily on semantic segmentation models for accurate and safe decision-making. High segmentation performance in real-world urban scenes is crucial for autonomous vehicles, while substantial pixel-level labels are required during model training. Unsupervised domain adaptation (UDA) techniques are widely used to adapt the segmentation model trained on the synthetic data (i.e., source domain) to the real-world data (i.e., target domain) since obtaining pixel-level annotations is fairly easy in the synthetic environment. Recently, increasing UDA approaches promote cross-domain semantic segmentation (CDSS) by fusing the depth information into the RGB features. However, feature fusion does not necessarily eliminate the domain-specific components in the RGB features, which can result in the features still being influenced by domain-specific information. To address this, we propose a novel cross-modal attentive feature alignment (CAFA) framework for CDSS, which provides an explicit perspective of using depth information to align the main backbone RGB features of both domains in a nonadversarial manner. In particular, considering that the depth modality is less affected by the domain gap, we employ depth as an intermediate modality and align the RGB features by attending RGB features to the depth modality through constructing an auxiliary multimodal segmentation task. The state-of-the-art performance of our CAFA can be achieved on benchmark tasks, such as Synthia$\to$Cityscapes and grand theft auto (GTA)$\to$Cityscapes. Peng Liu 0049, Yanqi Ge, Lixin Duan, Wen Li 0001, Fengmao Lv |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Transferring Multi-Modal Domain Knowledge to Uni-Modal Domain for Urban Scene SegmentationabstractSynthetic data (i.e., source domain) have been widely adopted to improve the semantic segmentation performance for real-world images (i.e., target domain), since obtaining pixel-level annotations is fairly easy in the synthetic environment. Traditional domain adaptation methods normally focus on learning in the RGB modality only. We notice that the synthetic environment can generate depth information of semantic objects at almost no cost, while it is nontrivial to collect such information in the real-world scenario. In this case, we employ the depth information of synthetic data in this work to further boost the segmentation performance, and then transform the uni-modal problem into a multi-modal one. In this work, we focus on urban scene understanding and make a pioneer attempt on learning uni-modal feature representations for real-world images by mining from multi-modal knowledge of synthetic images with additional depth information. To this end, we propose a novel method called Multi-modal Domain Knowledge Transfer (MDKT), which transfers the multi-modal knowledge of the source domain to the uni-modal target domain through domain adaptation. In MDKT, we first employ the Cross-Modal Correlation (CMC) module to enhance the source features by fusing the RGB and depth information. Then, the uni-modal target domain feature and multi-modal source domain feature are aligned through the Modal-Imbalanced Adversarial Training (MIAT) strategy, which transfers the multi-modal knowledge to the uni-modal network in the target domain. We conduct extensive experiments on several benchmark settings for urban scene understanding. The promising results clearly show the effectiveness of our proposed MDKT approach. Peng Liu 0049, Yanqi Ge, Lixin Duan, Wen Li 0001, Haonan Luo 0002, Fengmao Lv |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Cross-Domain Detection Transformer Based on Spatial-Aware and Semantic-Aware Token AlignmentabstractDetection transformers such as DETR [1] have recently exhibited promising performance for many object detection tasks, but the generalization ability of those methods is still quite limited for cross-domain adaptation scenarios. To address the cross-domain issue, a straightforward method is to perform token alignment with adversarial training in transformers. However, its performance is often unsatisfactory because the tokens in detection transformers are quite diverse and represent different spatial and semantic information. In this paper, we propose a new method for cross-domain detection transformers called spatial-aware and semantic-aware token alignment (SSTA). Specifically, we take advantage of the characteristics of cross-attention as used in the detection transformer and propose spatial-aware token alignment (SpaTA) and semantic-aware token alignment (SemTA) strategies to guide the token alignment across domains. For spatial-aware token alignment, we extract the information from the cross-attention map (CAM) to align the distribution of tokens according to their attention to object queries. For semantic-aware token alignment, we inject the category information into the cross-attention map and construct domain embedding to guide the learning of a multi-class discriminator to model the category relationship and achieve category-level token alignment during the entire adaptation process. We conduct extensive experiments on several widely-used benchmarks, and the results clearly show the effectiveness of our proposed approach over existing state-of-the-art methods. Jinhong Deng, Wen Li 0001, Lixin Duan, Dong Xu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Harmonious Teacher for Cross-Domain Object DetectionabstractSelf-training approaches recently achieved promising results in cross-domain object detection, where people iteratively generate pseudo labels for unlabeled target domain samples with a model, and select high-confidence samples to refine the model. In this work, we reveal that the consistency of classification and localization predictions are crucial to measure the quality of pseudo labels, and propose a new Harmonious Teacher approach to improve the self-training for cross-domain object detection. In particular, we first propose to enhance the quality of pseudo labels by regularizing the consistency of the classification and localization scores when training the detection model. The consistency losses are defined for both labeled source samples and the unlabeled target samples. Then, we further remold the traditional sample selection method by a sample reweighing strategy based on the consistency of classification and localization scores to improve the ranking of predictions. This allows us to fully exploit all instance predictions from the target domain without abandoning valuable hard examples. Without bells and whistles, our method shows superior performance in various cross-domain scenarios compared with the state-of-the-art baselines, which validates the effectiveness of our Harmonious Teacher. Our codes will be available at https://github.com/kinredon/Harmonious-Teacher. Jinhong Deng, Dongli Xu, Wen Li 0001, Lixin Duan |
CVPR | 3 |
| 2023 | Minimizing Maximum Model Discrepancy for Transferable Black-box Targeted AttacksabstractIn this work, we study the black-box targeted attack problem from the model discrepancy perspective. On the theoretical side, we present a generalization error bound for black-box targeted attacks, which gives a rigorous theoretical analysis for guaranteeing the success of the attack. We reveal that the attack error on a target model mainly depends on empirical attack error on the substitute model and the maximum model discrepancy among substitute models. On the algorithmic side, we derive a new algorithm for black-box targeted attacks based on our theoretical analysis, in which we additionally minimize the maximum model discrepancy (M3D) of the substitute models when training the generator to generate adversarial examples. In this way, our model is capable of crafting highly transferable adversarial examples that are robust to the model variation, thus improving the success rate for attacking the black-box model. We conduct extensive experiments on the ImageNet dataset with different classification models, and our proposed approach outperforms existing state-of-the-art methods by a significant margin. The code will be available at https://github.com/Asteriajojo/M3D. Tong Chu, Yahao Liu, Wen Li 0001, Jingjing Li 0001, Lixin Duan |
CVPR | 4 |
| 2023 | Multi-View Token Clustering and Fusion for 3D Object Recognition and Retrievalabstract3D object recognition has received extensive attention in recent years. Many existing methods tackle the task by rendering 3D objects from multiple views. However, most multi-view recognition methods do not utilize fine-grained information from different views, which is found to be crucial for improving 3D object representation in the multi-view setting. In this paper, we propose a transformer-based method, referred to as MVCFormer, for multi-view feature clustering and fusion. MVCFormer clusters semantically similar tokens at the same stages and selects representative fine-grained features, which helps to eliminate feature redundancy and remove cluttered backgrounds and make the selected features more diverse. On the other hand, our model also integrates selected features from all stages to obtain a discriminative 3D object representation by a cross-attention fusion method. Extensive experiments on benchmark datasets (e.g., ModelNet40, ModelNet10, ShapeNetCore55, and RGBD) clearly demonstrate the effectiveness of our proposed MVCFormer over existing baselines. Linlong Fan, Yanqi Ge, Wen Li 0001, Lixin Duan |
ICME | 3 |
| 2023 | Learning continuous piecewise non-linear activation functions for deep neural networksabstractActivation functions provide the non-linearity to deep neural networks, which are crucial for the optimization and performance improvement. In this paper, we propose a learnable continuous piece-wise nonlinear activation function (or CPN in short), which improves the widely used ReLU from three directions, i.e., finer pieces, non-linear terms and learnable parameterization. CPN is a continuous activation function with multiple pieces and incorporates non-linear terms in every interval. We give a general formulation of CPN and provide different implementations according to three key factors: whether the activation space is divided uniformly or not, whether the non-linear terms exist or not, and whether the activation function is continuous or not. We demonstrate the effectiveness of our method on image classification and single image super-resolution tasks by simply changing the activation function. For example, CPN improves 4.78% / 4.52% top-1 accuracy over ReLU on MobileNetV2_0.25 / MobileNetV2_0.35 for ImageNet classification and achieves better PSNR on several benchmarks for super-resolution. Our implementation is available at https://github.com/xc-G/CPN. Xinchen Gao, Yawei Li 0001, Wen Li 0001, Lixin Duan, Luc Van Gool, Luca Benini, Michele Magno |
ICME | 3 |
| 2023 | Learning Motion Refinement for Unsupervised Face AnimationabstractUnsupervised face animation aims to generate a human face video based on the
appearance of a source image, mimicking the motion from a driving video. Existing
methods typically adopted a prior-based motion model (e.g., the local affine motion
model or the local thin-plate-spline motion model). While it is able to capture
the coarse facial motion, artifacts can often be observed around the tiny motion
in local areas (e.g., lips and eyes), due to the limited ability of these methods
to model the finer facial motions. In this work, we design a new unsupervised
face animation approach to learn simultaneously the coarse and finer motions. In
particular, while exploiting the local affine motion model to learn the global coarse
facial motion, we design a novel motion refinement module to compensate for
the local affine motion model for modeling finer face motions in local areas. The
motion refinement is learned from the dense correlation between the source and
driving images. Specifically, we first construct a structure correlation volume based
on the keypoint features of the source and driving images. Then, we train a model
to generate the tiny facial motions iteratively from low to high resolution. The
learned motion refinements are combined with the coarse motion to generate the
new image. Extensive experiments on widely used benchmarks demonstrate that
our method achieves the best results among state-of-the-art baselines. Jiale Tao, Shuhang Gu, Wen Li 0001, Lixin Duan |
NeurIPS | 3 |
| 2023 | Noisy Label Learning With Provable Consistency for a Wider Family of LossesabstractDeep models have achieved state-of-the-art performance on a broad range of visual recognition tasks. Nevertheless, the generalization ability of deep models is seriously affected by noisy labels. Though deep learning packages have different losses, this is not transparent for users to choose consistent losses. This paper addresses the problem of how to use abundant loss functions designed for the traditional classification problem in the presence of label noise. We present a dynamic label learning (DLL) algorithm for noisy label learning and then prove that any surrogate loss function can be used for classification with noisy labels by using our proposed algorithm, with a consistency guarantee that the label noise does not ultimately hinder the search for the optimal classifier of the noise-free sample. In addition, we provide a depth theoretical analysis of our algorithm to verify the justifies' correctness and explain the powerful robustness. Finally, experimental results on synthetic and real datasets confirm the efficiency of our algorithm and the correctness of our justifies and show that our proposed algorithm significantly outperforms or is comparable to current state-of-the-art counterparts. Defu Liu 0001, Wen Li 0001, Lixin Duan, Ivor W. Tsang, Guowu Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Co-MDA: Federated Multisource Domain Adaptation on Black-Box ModelsabstractFederated domain adaptation (FDA) is an effective method for performing learning tasks over distributed networks, which well improves data privacy and portability in unsupervised multi-source domain adaptation (UMDA) tasks. Despite the impressive gains achieved, two common limitations exist in current FDA works. First, most previous studies require access to the model parameters or gradient details of each source party. However, the raw source data can be reconstructed from the model gradients or parameters, which may leak individual information. Second, these works assume that different parties share an identical network architecture, which is impractical and not desirable for low- or high-resource target users. To address these issues, in this work, we propose a more practical UMDA setting, called Federated Multi-source Domain Adaptation on Black-box Models (B2FDA), where all data are stored locally and only the input-output interface of the source model is available. To tackle B2FDA, we propose an effective method, termed Co2-Learning with Multi-Domain Attention (Co-MDA). Experiments on multiple benchmark datasets demonstrate the effectiveness of our proposed method. Notably, Co-MDA performs comparably with traditional UMDA methods where the source data or the trained model are fully available. Wei Xi 0003, Wen Li 0001, Dong Xu 0001, Gairui Bai, Jizhong Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Domain Adaptive Sampling for Cross-Domain Point Cloud RecognitionabstractPoint cloud recognition has recently gained increasing research interest due to the huge potential in real-world applications such as autonomous driving, robotics, etc. However, the point clouds of similar objects often exhibit notable geometric variations due to the difference in capturing devices or environmental changes. This leads to significant performance degradation when the learned point cloud recognition model is applied to a new scenario, which is also known as the domain adaptation issue. In this work, we propose a new unsupervised domain adaptation approach for point cloud recognition via domain adaptive sampling (DAS). In particular, we propose a two-level sampling strategy of point level and instance level to improve the cross-domain recognition ability of the model. First, we propose a domain adaptive point sampling (DAPS) strategy to enhance the domain-invariant representation of point clouds by progressively focusing on representative points in each point cloud based on geometric consistency. Then, we further propose an instance-level domain adaptive cloud sampling (DACS) strategy to learn target-specific information based on a self-paced learning paradigm, where we select a set of pseudo-labeled target point clouds to train our designed light-weighted adapters without modifying the learned domain-invariant representation. We validate our domain adaptive sampling approach on the benchmark datasets PointDA-10 and GraspNetPC-10, where our method achieves new state-of-the-art performance. Zicheng Wang 0012, Wen Li 0001, Dong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Deep Cross-Attention Network for Crowdfunding Success PredictionabstractCrowdfunding creates opportunities for entrepre- neurs. It allows startup companies to reach a large audience for fundraising and bring their creative ideas to life. In this work, we are concerned with crowdfunding project success prediction problem,i.e., to predict whether a project will successfully reach its funding goal by using its project profiles. This is important for startup companies to refine their project profiles and achieve their goals. Crowdfunding project success prediction is a typical classification problem but with a few critical challenges. On the one hand, with only coarse-grained project status as weak supervision, it is hard for a deep learning network to learn the relationship between project profiles and explain why it makes this prediction. On the other hand, on the project homepage, there are various modalities of description, including metadata, textual description, images, and videos. Among those, videos play an important role in the success of a crowdfunding project, however, were ignored in previous works, due to the difficulty in extracting useful semantic and authentic information from videos, especially for the crowdfunding project where information in different modalities are unaligned. To this end, we propose a novel framework called Deep Cross-Attention Network to learn and fuse information from introduction videos and textual descriptions of project profiles. More specifically, we develop a cross-attention block to align and represent mismatched textual description and untrimmed introduction videos and fuse the information from these two modalities, which effectively remedies the lack of supervised information caused by project status as weak supervision. More importantly, with our cross-attention mechanism, the model is able to interpret how it makes such predictions and show which keywords and keyframes it depends on. We conduct extensive experiments on two crowdfunding datasets (collected from Kickstarter and Indiegogo) and show that our method achieves superior performance over existing state-of-the-art baselines. Yi Yang 0042, Wen Li 0001, Defu Lian, Lixin Duan |
IEEE Trans. Multim. | 3 |
| 2022 | Denoised Maximum Classifier Discrepancy for Source-Free Unsupervised Domain AdaptationabstractSource-Free Unsupervised Domain Adaptation(SFUDA) aims to adapt a pre-trained source model to an unlabeled target domain without access to the original labeled source domain samples. Many existing SFUDA approaches apply the self-training strategy, which involves iteratively selecting confidently predicted target samples as pseudo-labeled samples used to train the model to fit the target domain. However, the self-training strategy may also suffer from sample selection bias and be impacted by the label noise of the pseudo-labeled samples. In this work, we provide a rigorous theoretical analysis on how these two issues affect the model generalization ability when applying the self-training strategy for the SFUDA problem. Based on this theoretical analysis, we then propose a new Denoised Maximum Classifier Discrepancy (D-MCD) method for SFUDA to effectively address these two issues. In particular, we first minimize the distribution mismatch between the selected pseudo-labeled samples and the remaining target domain samples to alleviate the sample selection bias. Moreover, we design a strong-weak self-training paradigm to denoise the selected pseudo-labeled samples, where the strong network is used to select pseudo-labeled samples while the weak network helps the strong network to filter out hard samples to avoid incorrect labels. In this way, we are able to ensure both the quality of the pseudo-labels and the generalization ability of the trained model on the target domain. We achieve state-of-the-art results on three domain adaptation benchmark datasets, which clearly validates the effectiveness of our proposed approach. Full code is available at https://github.com/kkkkkkon/D-MCD. Tong Chu, Yahao Liu, Jinhong Deng, Wen Li 0001, Lixin Duan |
AAAI | 4 |
| 2022 | Revisiting Random Channel Pruning for Neural Network CompressionabstractChannel (or 3D filter) pruning serves as an effective way to accelerate the inference of neural networks. There has been a flurry of algorithms that try to solve this practical problem, each being claimed effective in some ways. Yet, a benchmark to compare those algorithms directly is lacking, mainly due to the complexity of the algorithms and some custom settings such as the particular network configuration or training procedure. A fair benchmark is important for the further development of channel pruning. Meanwhile, recent investigations reveal that the channel configurations discovered by pruning algorithms are at least as important as the pre-trained weights. This gives channel pruning a new role, namely searching the optimal channel configuration. In this paper, we try to determine the channel configuration of the pruned models by random search. The proposed approach provides a new way to compare different methods, namely how well they behave compared with random pruning. We show that this simple strategy works quite well compared with other channel pruning methods. We also show that under this setting, there are surprisingly no clear winners among different channel importance evaluation methods, which then may tilt the research efforts into advanced channel configuration searching methods. Code will be released at https://github.com/ofsoundof/random_channel_pruning. Yawei Li 0001, Kamil Adamczewski, Wen Li 0001, Shuhang Gu, Radu Timofte, Luc Van Gool |
CVPR | 3 |
| 2022 | Undoing the Damage of Label Shift for Cross-domain Semantic SegmentationabstractExisting works typically treat cross-domain semantic segmentation (CDSS) as a data distribution mismatch prob-lem and focus on aligning the marginal distribution or con-ditional distribution. However, the label shift issue is un-fortunately overlooked, which actually commonly exists in the CDSS task, and often causes a classifier bias in the learnt model. In this paper, we give an in-depth analysis and show that the damage of label shift can be overcome by aligning the data conditional distribution and correcting the posterior probability. To this end, we propose a novel approach to undo the damage of the label shift problem in CDSS. In implementation, we adopt class-level feature alignment for conditional distribution alignment, as well as two simple yet effective methods to rectify the classifier bias from source to target by remolding the classifier predictions. We conduct extensive experiments on the benchmark datasets of urban scenes, including GTA5 to Cityscapes and SYNTHIA to Cityscapes, where our proposed approach outperforms previous methods by a large margin. For instance, our model equipped with a self-training strat-egy reaches 59.3% mIoU on GTA5 to Cityscapes, pushing to a new state-of-the-art. The code will be available at https://github.com/manmanjun/Undoing_UDA. Yahao Liu, Jinhong Deng, Jiale Tao, Tong Chu, Lixin Duan, Wen Li 0001 |
CVPR | 6 |
| 2022 | Meta Distribution Alignment for Generalizable Person Re-IdentificationabstractDomain Generalizable (DG) person ReID is a challenging task which trains a model on source domains yet generalizes well on target domains. Existing methods use source domains to learn domain-invariant features, and assume those features are also irrelevant with target domains. However, they do not consider the target domain information which is unavailable in the training phrase of DG. To address this issue, we propose a novel Meta Distribution Alignment (MDA) method to enable them to share similar distribution in a test-time-training fashion. Specifically, since high-dimensional features are difficult to constrain with a known simple distribution, we first introduce an intermediate latent space constrained to a known prior distribution. The source domain data is mapped to this latent space and then reconstructed back. A meta-learning strategy is introduced to facilitate generalization and support fast adaption. To reduce their discrepancy, we further propose a test-time adaptive updating strategy based on the latent space which efficiently adapts model to unseen domains with a few samples. Extensive experimental results show that our model outperforms the state-of-the-art methods by up to 5.2% R-1 on average on the large-scale and 4.7% R-1 on the single-source domain generalization ReID benchmark. Source code is publicly available at https://github.com/haoni0812/MDA.git. Hao Ni 0002, Jingkuan Song, Xiaopeng Luo, Feng Zheng 0001, Wen Li 0001, Heng Tao Shen |
CVPR | 5 |
| 2022 | Semantic-Aware Domain Generalized SegmentationabstractDeep models trained on source domain lack generalization when evaluated on unseen target domains with different data distributions. The problem becomes even more pro-nounced when we have no access to target domain samples for adaptation. In this paper, we address domain generalized semantic segmentation, where a segmentation model is trained to be domain-invariant without using any target domain data. Existing approaches to tackle this problem standardize data into a unified distribution. We argue that while such a standardization promotes global normalization, the resulting features are not discriminative enough to get clear segmentation boundaries. To enhance separation between categories while simultaneously promoting domain invariance, we propose a framework including two novel modules: Semantic-Aware Normalization (SAN) and Semantic-Aware Whitening (SAW). Specifically, SAN focuses on category-level center alignment between features from different image styles, while SAW enforces distributed alignment for the already center-aligned features. With the help of SAN and SAW, we encourage both intra-category compactness and inter-category separability. We validate our approach through extensive experiments on widely-used datasets (i.e. GTAV, SYNTHIA, Cityscapes, Mapillary and BDDS). Our approach shows significant improvements over existing state-of-the-art on various backbone networks. Code is available at https://github.com/leolyj/SAN-SAW Duo Peng, Yinjie Lei, Munawar Hayat, Yulan Guo, Wen Li 0001 |
CVPR | 5 |
| 2022 | Structure-Aware Motion Transfer with Deformable Anchor ModelabstractGiven a source image and a driving video depicting the same object type, the motion transfer task aims to generate a video by learning the motion from the driving video while preserving the appearance from the source image. In this paper, we propose a novel structure-aware motion modeling approach, the deformable anchor model (DAM), which can automatically discover the motion structure of arbitrary objects without leveraging their prior structure information. Specifically, inspired by the known deformable part model (DPM), our DAM introduces two types of anchors or key-points: i) a number of motion anchors that capture both appearance and motion information from the source image and driving video; ii) a latent root anchor, which is linked to the motion anchors to facilitate better learning of the representations of the object structure information. More-over, DAM can be further extended to a hierarchical version through the introduction of additional latent anchors to model more complicated structures. By regularizing motion anchors with latent anchor(s), DAM enforces the corre-spondences between them to ensure the structural information is well captured and preserved. Moreover, DAM can be learned effectively in an unsupervised manner. We validate our proposed DAM for motion transfer on different bench-mark datasets. Extensive experiments clearly demonstrate that DAM achieves superior performance relative to existing state-of-the-art methods. Jiale Tao, Borun Xu, Tiezheng Ge, Yuning Jiang 0001, Wen Li 0001, Lixin Duan |
CVPR | 6 |
| 2022 | Learning Pixel-Level Distinctions for Video Highlight DetectionabstractThe goal of video highlight detection is to select the most attractive segments from a long video to depict the most interesting parts of the video. Existing methods typically focus on modeling relationship between different video segments in order to learning a model that can assign highlight scores to these segments; however, these approaches do not explicitly consider the contextual dependency within individual segments. To this end, we propose to learn pixel-level distinctions to improve the video highlight detection. This pixel-level distinction indicates whether or not each pixel in one video belongs to an interesting section. The advantages of modeling such fine-level distinctions are two-fold. First, it allows us to exploit the temporal and spatial relations of the content in one video, since the distinction of a pixel in one frame is highly dependent on both the content before this frame and the content around this pixel in this frame. Second, learning the pixel-level distinction also gives a good explanation to the video highlight task regarding what contents in a highlight segment will be attractive to people. We design an encoder-decoder network to estimate the pixel-level distinction, in which we leverage the 3D convolutional neural networks to exploit the temporal context information, and further take advantage of the visual saliency to model the spatial distinction. State-of-the-art performance on three public benchmarks clearly validates the effectiveness of our framework for video highlight detection. Fanyue Wei, Tiezheng Ge, Yuning Jiang 0001, Wen Li 0001, Lixin Duan |
CVPR | 5 |
| 2022 | Revisiting AP Loss for Dense Object Detection: Adaptive Ranking Pair SelectionabstractAverage precision (AP) loss has recently shown promising performance on the dense object detection task. However, a deep understanding of how AP loss affects the detector from a pairwise ranking perspective has not yet been developed. In this work, we revisit the average precision (AP) loss and reveal that the crucial element is that of selecting the ranking pairs between positive and negative samples. Based on this observation, we propose two strategies to improve the AP loss. The first of these is a novel Adaptive Pairwise Error (APE) loss that focusing on ranking pairs in both positive and negative samples. Moreover, we select more accurate ranking pairs by exploiting the normalized ranking scores and localization scores with a clustering algorithm. Experiments conducted on the MSCOCO dataset support our analysis and demonstrate the superiority of our proposed method compared with current classification and ranking loss. The code is available at https://github.com/Xudangliatiger/APE-Loss. Dongli Xu, Jinhong Deng, Wen Li 0001 |
CVPR | 3 |
| 2022 | Interpretable Open-Set Domain Adaptation via Angular Margin Separation
Xinhao Li 0002, Jingjing Li 0001, Zhekai Du, Lei Zhu 0002, Wen Li 0001 |
ECCV (34) | 5 |
| 2022 | Motion Transformer for Unsupervised Image Animation
Jiale Tao, Tiezheng Ge, Yuning Jiang 0001, Wen Li 0001, Lixin Duan |
ECCV (16) | 5 |
| 2022 | Motion and Appearance Adaptation for Cross-domain Motion Transfer
Borun Xu, Jinhong Deng, Jiale Tao, Tiezheng Ge, Yuning Jiang 0001, Wen Li 0001, Lixin Duan |
ECCV (16) | 7 |
| 2022 | Diverse Preference Augmentation with Multiple Domains for Cold-start RecommendationsabstractCold-start issues have been more and more challenging for providing accurate recommendations with the fast increase of users and items. Most existing approaches attempt to solve the intractable problems via content-aware recommendations based on auxiliary information and/or cross-domain recommendations with transfer learning. Their performances are often constrained by the extremely sparse user-item interactions, unavailable side information, or very limited domain-shared users. Recently, meta-learners with meta-augmentation by adding noises to labels have been proven to be effective to avoid overfitting and shown good performance on new tasks. Motivated by the idea of meta-augmentation, in this paper, by treating a user's preference over items as a task, we propose a so-called Diverse Preference Augmentation framework with multiple source domains based on meta-learning (referred to as MetaDPA) to i) generate diverse ratings in a new domain of interest (known as target domain) to handle overfitting on the case of sparse interactions, and to ii) learn a preference model in the target domain via a meta-learning scheme to alleviate cold-start issues. Specifically, we first conduct multi-source domain adaptation by dual conditional variational autoencoders and impose a Multi-domain InfoMax (MDI) constraint on the latent representations to learn domain-shared and domain-specific preference properties. To avoid overfitting, we add a Mutually-Exclusive (ME) constraint on the output of decoders to generate diverse ratings given content data. Finally, these generated diverse ratings and the original ratings are introduced into the meta-training procedure to learn a preference meta-learner, which produces good generalization ability on cold-start recommendation tasks. Experiments on real-world datasets show our proposed MetaDPA clearly outperforms the current state-of-the-art baselines. Yan Zhang 0036, Changyu Li, Ivor W. Tsang, Lixin Duan, Hongzhi Yin, Wen Li 0001, Jie Shao 0001 |
ICDE | 7 |
| 2022 | Partial Label Learning with Semantic Label RepresentationsabstractPartial-label learning (PLL) solves the problem where each training instance is assigned a candidate label set, among which only one is the ground-truth label. The core of PLL is to learn efficient feature representations to facilitate label disambiguation. However, existing PLL methods only learn plain representations by coarse supervision, which is incapable of capturing sufficiently distinguishable representations, especially when confronted with the knotty label ambiguity, i.e., certain candidate labels share similar visual patterns. In this paper, we propose a novel framework partial label learning with semantic label representations dubbed ParSE, which consists of two synergistic processes, including visual-semantic representation learning and powerful label disambiguation. In the former process, we propose a novel weighted calibration rank loss that has two implications. First, it implies a progressive calibration strategy that utilizes the disambiguated label confidence to weight the similarity between each image feature embedding and its corresponding semantic label representations of all candidates. Second, it also considers the ranking relationship between candidate and non-candidate ones. Based on learned visual-semantic representations, subsequent label disambiguation is desirably endowed with more powerful abilities. Experiments on benchmarks show that ParSE outperforms state-of-the-art counterparts. Shuo He 0001, Lei Feng 0006, Fengmao Lv, Wen Li 0001, Guowu Yang |
KDD | 4 |
| 2022 | Text Enhancement Network for Cross-Domain Scene Text DetectionabstractConventional scene text detection approaches essentially assume that training and test data are drawn from the same distribution and have achieved compelling results. However, scene text detectors often suffer from performance degradation in real-world applications, since the feature distribution of training images is different from that of test images obtained from a new scene. To address the above problems, we propose a novel method called Text Enhancement Network (TEN) based on adversarial learning for cross-domain scene text detection. Specifically, we first design a Multi-adversarial Feature Alignment (MFA) module to maximally align features of the source and target data from low-level texture to high-level semantics. Second, we develop the Text Attention Enhancement (TAE) module to re-weigh the importance of text regions and accordingly enhance the corresponding features, in order to improve the robustness against noisy background. Additionally, we design a self-training strategy to further boost the performance of our TEN. We conduct extensive experiments on five benchmarks, and the experimental results demonstrate the effectiveness of our TEN. Jinhong Deng, Xiulian Luo, Jiawen Zheng, Wanli Dang, Wen Li 0001 |
IEEE Signal Process. Lett. | 5 |
| 2022 | VDM-DA: Virtual Domain Modeling for Source Data-Free Domain AdaptationabstractDomain adaptation aims to leverage a label-rich domain (the source domain) to help model learning in a label-scarce domain (the target domain). Most domain adaptation methods require the co-existence of source and target domain samples to reduce the distribution mismatch. However, access to the source domain samples may not always be feasible in real-world applications due to different problems (e.g., storage, transmission, and privacy issues). In this work, we deal with the source data-free unsupervised domain adaptation problem and propose a novel approach referred to as Virtual Domain Modeling for Domain Adaptation (VDM-DA), in which the virtual domain acts as a bridge between the source and target domains. Specifically, based on the pre-trained source model, we generate the virtual domain samples by using an approximated Gaussian Mixture Model (GMM) in the feature space, such that the virtual domain maintains a similar distribution with the source domain without access to the original source data. Moreover, we also design an effective distribution alignment method to reduce the distribution divergence between the virtual domain and the target domain by gradually improving the compactness of the target domain distribution through model learning. In this way, we successfully achieve the goal of distribution alignment between the source and target domains when training deep networks without access to the source domain data. We conduct extensive experiments on four benchmark datasets for both 2D image-based and 3D point cloud-based cross-domain object recognition tasks, where the proposed method referred to as Virtual Domain Modeling for Domain Adaptation (VDM-DA) achieves the promising performance on all datasets. Jing Zhang 0017, Wen Li 0001, Dong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Analogical Image Translation for Fog GenerationabstractImage-to-image translation is to map images from a given style to another given style. While exceptionally successful, current methods assume the availability of training images in both source and target domains, which does not always hold in practice. Inspired by humans' reasoning capability of analogy, we propose analogical image translation (AIT) that exploit the concept of gist, for the first time. Given images of two styles in the source domain: A and A', along with images B of the first style in the target domain, learn a model to translate B to B' in the target domain, such that A:A' :: B:B'. AIT is especially useful for translation scenarios in which training data of one style is hard to obtain but training data of the same two styles in another domain is available. For instance, in the case from normal conditions to extreme, rare conditions, obtaining real training images for the latter case is challenging. However, obtaining synthetic data for both cases is relatively easy. In this work, we aim at adding adverse weather effects, more specifically fog, to images taken in clear weather. To circumvent the challenge of collecting real foggy images, AIT learns the gist of translating synthetic clear-weather to foggy images, followed by adding fog effects onto real clear-weather images, without ever seeing any real foggy image. AIT achieves zero-shot image translation capability, whose effectiveness and benefit are demonstrated by the downstream task of semantic foggy scene understanding. Dengxin Dai, Wen Li 0001, Danda Pani Paudel, Luc Van Gool |
AAAI | 4 |
| 2021 | Unbiased Mean Teacher for Cross-Domain Object DetectionabstractCross-domain object detection is challenging, because object detection model is often vulnerable to data variance, especially to the considerable domain shift between two distinctive domains. In this paper, we propose a new Unbiased Mean Teacher (UMT) model for cross-domain object detection. We reveal that there often exists a considerable model bias for the simple mean teacher (MT) model in cross-domain scenarios, and eliminate the model bias with several simple yet highly effective strategies. In particular, for the teacher model, we propose a cross-domain distillation method for MT to maximally exploit the expertise of the teacher model. Moreover, for the student model, we alleviate its bias by augmenting training samples with pixel-level adaptation. Finally, for the teaching process, we employ an out-of-distribution estimation strategy to select samples that most fit the current model to further enhance the cross-domain distillation process. By tackling the model bias issue with these strategies, our UMT model achieves mAPs of 44.1%, 58.1%, 41.7%, and 43.1% on benchmark datasets Clipart1k, Watercolor2k, Foggy Cityscapes, and Cityscapes, respectively, which outperforms the existing state-of-the-art results in notable margins. Our implementation is available at https://github.com/kinredon/umt. Jinhong Deng, Wen Li 0001, Lixin Duan |
CVPR | 2 |
| 2021 | Cluster, Split, Fuse, and Update: Meta-Learning for Open Compound Domain Adaptive Semantic SegmentationabstractOpen compound domain adaptation (OCDA) is a domain adaptation setting, where target domain is modeled as a compound of multiple unknown homogeneous domains, which brings the advantage of improved generalization to unseen domains. In this work, we propose a principled meta-learning based approach to OCDA for semantic segmentation, MOCDA, by modeling the unlabeled target domain continuously. Our approach consists of four key steps. First, we cluster target domain into multiple sub-target domains by image styles, extracted in an unsupervised manner. Then, different sub-target domains are split into independent branches, for which batch normalization parameters are learnt to treat them independently. A meta-learner is thereafter deployed to learn to fuse sub-target domain-specific predictions, conditioned upon the style code. Meanwhile, we learn to online update the model by model-agnostic meta-learning (MAML) algorithm, thus to further improve generalization. We validate the benefits of our approach by extensive experiments on synthetic-to-real knowledge transfer benchmark, where we achieve the state-of-the-art performance in both compound and open domains. Danda Pani Paudel, Yawei Li 0001, Ajad Chhatkuli, Wen Li 0001, Dengxin Dai, Luc Van Gool |
CVPR | 6 |
| 2021 | The Heterogeneity Hypothesis: Finding Layer-Wise Differentiated Network ArchitecturesabstractIn this paper, we tackle the problem of convolutional neural network design. Instead of focusing on the design of the overall architecture, we investigate a design space that is usually overlooked, i.e. adjusting the channel configurations of predefined networks. We find that this adjustment can be achieved by shrinking widened baseline networks and leads to superior performance. Based on that, we articulate the "heterogeneity hypothesis": with the same training protocol, there exists a layer-wise differentiated net-work architecture (LW-DNA) that can outperform the original network with regular channel configurations but with a lower level of model complexity.The LW-DNA models are identified without extra computational cost or training time compared with the original network. This constraint leads to controlled experiments which direct the focus to the importance of layer-wise specific channel configurations. LW-DNA models come with advantages related to overfitting, i.e. the relative relationship between model complexity and dataset size. Experiments are conducted on various networks and datasets for image classification, visual tracking and image restoration. The resultant LW-DNA models consistently outperform the baseline models. Code is available at https://github.com/ofsoundof/Heterogeneity_Hypothesis.git. Yawei Li 0001, Wen Li 0001, Martin Danelljan, Kai Zhang 0008, Shuhang Gu, Luc Van Gool, Radu Timofte |
CVPR | 2 |
| 2021 | SRDAN: Scale-Aware and Range-Aware Domain Adaptation Network for Cross-Dataset 3D Object DetectionabstractGeometric characteristic plays an important role in the representation of an object in 3D point clouds. For example, large objects often contain more points, while small ones contain fewer points. The points from objects near the capture device are denser, while those from far-range objects are sparser. These issues bring new challenges to 3D object detection, especially under the domain adaptation scenarios. In this work, we propose a new cross-dataset 3D object detection method named Scale-aware and Range-aware Domain Adaptation Network (SRDAN). We take advantage of the geometric characteristics of 3D data (i.e., size and distance), and propose the scale-aware domain alignment and the range-aware domain alignment strategies to guide the distribution alignment between two domains. For scale-aware domain alignment, we design a 3D voxel-based feature pyramid network to extract multi-scale semantic voxel features, and align the features and instances with similar scales between two domains. For range-aware domain alignment, we introduce a range-guided domain alignment module to align the features of objects according to their distance to the capture device. Extensive experiments under three different scenarios demonstrate the effectiveness of our SRDAN approach, and comprehensive ablation study also validates the importance of geometric characteristics for cross-dataset 3D object detection. Wen Li 0001, Dong Xu 0001 |
CVPR | 2 |
| 2021 | mDALU: Multi-Source Domain Adaptation and Label Unification with Partial DatasetsabstractOne challenge of object recognition is to generalize to new domains, to more classes and/or to new modalities. This necessitates methods to combine and reuse existing datasets that may belong to different domains, have partial annotations, and/or have different data modalities. This paper formulates this as a multi-source domain adaptation and label unification problem, and proposes a novel method for it. Our method consists of a partially-supervised adaptation stage and a fully-supervised adaptation stage. In the former, partial knowledge is transferred from multiple source domains to the target domain and fused therein. Negative transfer between unmatching label spaces is mitigated via three new modules: domain attention, uncertainty maximization and attention-guided adversarial alignment. In the latter, knowledge is transferred in the unified label space after a label completion process with pseudolabels. Extensive experiments on three different tasks - image classification, 2D semantic image segmentation, and joint 2D-3D semantic segmentation - show that our method outperforms all competing methods significantly. Dengxin Dai, Wen Li 0001, Luc Van Gool |
ICCV | 4 |
| 2021 | BAPA-Net: Boundary Adaptation and Prototype Alignment for Cross-domain Semantic SegmentationabstractExisting cross-domain semantic segmentation methods usually focus on the overall segmentation results of whole objects but neglect the importance of object boundaries. In this work, we find that the segmentation performance can be considerably boosted if we treat object boundaries properly. For that, we propose a novel method called BAPA-Net, which is based on a convolutional neural network via Boundary Adaptation and Prototype Alignment, under the unsupervised domain adaptation setting. Specifically, we first construct additional images by pasting objects from source images to target images, and we develop a so-called boundary adaptation module to weigh each pixel based on its distance to the nearest boundary pixel of those pasted source objects. Moreover, we propose another prototype alignment module to reduce the domain mismatch by minimizing distances between the class prototypes of the source and target domains, where boundaries are removed to avoid domain confusion during prototype calculation. By integrating the boundary adaptation and prototype alignment, we are able to train a discriminative and domain-invariant model for cross-domain semantic segmentation. We conduct extensive experiments on the benchmark datasets of urban scenes (i.e., GTA5→Cityscapes and SYNTHIA→Cityscapes). And the promising results clearly show the effectiveness of our BAPA-Net method over existing state-of-the-art for cross-domain semantic segmentation. Our implementation is available at https://github.com/manmanjun/BAPA-Net. Yahao Liu, Jinhong Deng, Xinchen Gao, Wen Li 0001, Lixin Duan |
ICCV | 4 |
| 2021 | Sparse-to-dense Feature Matching: Intra and Inter domain Cross-modal Learning in Domain Adaptation for 3D Semantic SegmentationabstractDomain adaptation is critical for success when confronting with the lack of annotations in a new domain. As the huge time consumption of labeling process on 3D point cloud, domain adaptation for 3D semantic segmentation is of great expectation. With the rise of multi-modal datasets, large amount of 2D images are accessible besides 3D point clouds. In light of this, we propose to further leverage 2D data for 3D domain adaptation by intra and inter domain cross modal learning. As for intra-domain cross modal learning, most existing works sample the dense 2D pixel-wise features into the same size with sparse 3D point-wise features, resulting in the abandon of numerous useful 2D features. To address this problem, we propose Dynamic sparse-to-dense Cross Modal Learning (DsCML) to increase the sufficiency of multi-modality information interaction for domain adaptation. For inter-domain cross modal learning, we further advance Cross Modal Adversarial Learning (CMAL) on 2D and 3D data which contains different semantic content aiming to promote high-level modal complementarity. We evaluate our model under various multi-modality domain adaptation settings including day-to-night, country-to-country and dataset-to-dataset, brings large improvements over both uni-modal and multi-modal domain adaptation methods on all settings. Code is available at https://github.com/leolyj/DsCML Duo Peng, Yinjie Lei, Wen Li 0001, Yulan Guo |
ICCV | 3 |
| 2021 | Multi-Scale Enhanced Active Learning for Skeleton-Based Action RecognitionabstractSkeleton-based models have been widely used, because of their robustness to complex backgrounds and high computational efficiency. However, annotating skeleton sequences is labor-intensive. It is appealing to reduce the cost of acquiring data with accurate labels for skeleton-based models. This paper presents an active learning method for the skeleton-based action recognition model, which boosts the performance of the model with less labeled data by instructing humans to annotate the most valuable samples. The key issue in active learning is to train a model that precisely predicts the values of samples. To achieve this, we propose to enhance the ability of our model to evaluate samples by modeling actions from different granularities from multi-scale representations of skeletons. The multi-scale method is simple and easy to use, which can be treated as a plug-and-play extension to strengthen the skeleton-based models. We conduct experiments on the SHREC, NTU-60, and Kinetics-Skeleton. Extensive experimental results demonstrate the effectiveness of the proposed method. Wen Li 0001, Lixin Duan |
ICME | 3 |
| 2021 | Counterfactual Debiasing Inference for Compositional Action RecognitionabstractCompositional action recognition is a novel challenge in the computer vision community and focuses on revealing the different combinations of verbs and nouns instead of treating subject-object interactions in videos as individual instances only. Existing methods tackle this challenging task by simply ignoring appearance information or fusing object appearances with dynamic instance tracklets. However, those strategies usually do not perform well for unseen action instances. For that, in this work we propose a novel learning framework called Counterfactual Debiasing Network (CDN) to improve the model generalization ability by removing the interference introduced by visual appearances of objects/subjects. It explicitly learns the appearance information in action representations and later removes the effect of such information in a causal inference manner. Specifically, we use tracklets and video content to model the factual inference by considering both appearance information and structure information. In contrast, only video content with appearance information is leveraged in the counterfactual inference. With the two inferences, we conduct a causal graph which captures and removes the bias introduced by the appearance information by subtracting the result of the counterfactual inference from that of the factual inference. By doing that, our proposed CDN method can better recognize unseen action instances by debiasing the effect of appearances. Extensive experiments on the Something-Else dataset clearly show the effectiveness of our proposed CDN over existing state-of-the-art methods. Pengzhan Sun 0001, Bo Wu 0018, Xunsong Li, Wen Li 0001, Lixin Duan, Chuang Gan 0001 |
ACM Multimedia | 4 |
| 2021 | Move As You Like: Image Animation in E-Commerce ScenarioabstractCreative image animations are attractive in e-commerce applications, where motion transfer is one of the import ways to generate animations from static images. However, existing methods rarely transfer motion to objects other than human body or human face, and even fewer apply motion transfer in practical scenarios. In this work, we apply motion transfer on the Taobao product images in real e-commerce scenario to generate creative animations, which are more attractive than static images and they will bring more benefits. We animate the Taobao products of dolls, copper running horses and toy dinosaurs based on motion transfer method for demonstration. Borun Xu, Jiale Tao, Tiezheng Ge, Yuning Jiang 0001, Wen Li 0001, Lixin Duan |
ACM Multimedia | 6 |
| 2021 | STST: Spatial-Temporal Specialized Transformer for Skeleton-based Action RecognitionabstractSkeleton-based action recognition has been widely investigated considering their strong adaptability to dynamic circumstances and complicated backgrounds. To recognize different actions from skeleton sequences, it is essential and crucial to model the posture of the human represented by the skeleton and its changes in the temporal dimension. However, most of the existing works treat skeleton sequences in the temporal and spatial dimension in the same way, ignoring the difference between the temporal and spatial dimension in skeleton data which is not an optimal way to model skeleton sequences. The posture represented by the skeleton in each frame is proposed to be modeled individually. Meanwhile, capturing the movement of the entire skeleton in the temporal dimension is needed. So, we designed Spatial Transformer Block and Directional Temporal Transformer Block for modeling skeleton sequences in spatial and temporal dimensions respectively. Due to occlusion/sensor/raw video, etc., there are noises on both temporal and spatial dimensions in the extracted skeleton data reducing the recognition capabilities of models. To adapt to this imperfect information condition, we propose a multi-task self-supervised learning method by providing confusing samples in different situations to improve the robustness of our model. Combining the above design, we propose our Spatial-Temporal Specialized Transformer~(STST) and conduct experiments with our model on the SHREC, NTU-RGB+D, and Kinetics-Skeleton. Extensive experimental results demonstrate the improved performances and analysis of the proposed method. Bo Wu 0018, Wen Li 0001, Lixin Duan, Chuang Gan 0001 |
ACM Multimedia | 3 |
| 2021 | Scale-Aware Domain Adaptive Faster R-CNN
Wen Li 0001, Christos Sakaridis, Dengxin Dai, Luc Van Gool |
Int. J. Comput. Vis. | 3 |
| 2021 | DLOW: Domain Flow and Applications
Wen Li 0001, Dengxin Dai, Luc Van Gool |
Int. J. Comput. Vis. | 2 |
| 2021 | Self-Paced Collaborative and Adversarial Network for Unsupervised Domain AdaptationabstractThis paper proposes a new unsupervised domain adaptation approach called Collaborative and Adversarial Network (CAN), which uses the domain-collaborative and domain-adversarial learning strategies for training the neural network. The domain-collaborative learning strategy aims to learn domain specific feature representation to preserve the discriminability for the target domain, while the domain adversarial learning strategy aims to learn domain invariant feature representation to reduce the domain distribution mismatch between the source and target domains. We show that these two learning strategies can be uniformly formulated as domain classifier learning with positive or negative weights on the losses. We then design a collaborative and adversarial training scheme, which automatically learns domain specific representations from lower blocks in CNNs through collaborative learning and domain invariant representations from higher blocks through adversarial learning. Moreover, to further enhance the discriminability in the target domain, we propose Self-Paced CAN (SPCAN), which progressively selects pseudo-labeled target samples for re-training the classifiers. We employ a self-paced learning strategy such that we can select pseudo-labeled target samples in an easy-to-hard fashion. Additionally, we build upon the popular two-stream approach to extend our domain adaptation approach for more challenging video action recognition task, which additionally considers the cooperation between the RGB stream and the optical flow stream. We propose the Two-stream SPCAN (TS-SPCAN) method to select and reweight the pseudo labeled target samples of one stream (RGB/Flow) based on the information from the other stream (Flow/RGB) in a cooperative way. As a result, our TS-SPCAN model is able to exchange the information between the two streams. Comprehensive experiments on different benchmark datasets, Office-31, ImageCLEF-DA and VISDA-2017 for the object recognition task, and UCF101-10 and HMDB51-10 for the video action recognition task, show our newly proposed approaches achieve the state-of-the-art performance, which clearly demonstrates the effectiveness of our proposed approaches for unsupervised domain adaptation. Dong Xu 0001, Wanli Ouyang, Wen Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Cross-Dataset Point Cloud Recognition Using Deep-Shallow Domain Adaptation NetworkabstractIn this work, we propose a new two-view domain adaptation network named Deep-Shallow Domain Adaptation Network (DSDAN) for 3D point cloud recognition. Different from the traditional 2D image recognition task, the valuable texture information is often absent in point cloud data, making point cloud recognition a challenging task, especially in the cross-dataset scenario where the training and testing data exhibit a considerable distribution mismatch. In our DSDAN method, we tackle the challenging cross-dataset 3D point cloud recognition task from two aspects. On one hand, we propose a two-view learning framework, such that we can effectively leverage multiple feature representations to improve the recognition performance. To this end, we propose a simple and efficient Bag-of-Points feature method, as a complementary view to the deep representation. Moreover, we also propose a cross view consistency loss to boost the two-view learning framework. On the other hand, we further propose a two-level adaptation strategy to effectively address the domain distribution mismatch issue. Specifically, we apply a feature-level distribution alignment module for each view, and also propose an instance-level adaptation approach to select highly confident pseudo-labeled target samples for adapting the model to the target domain, based on which a co-training scheme is used to integrate the learning and adaptation process on the two views. Extensive experiments on the benchmark dataset show that our newly proposed DSDAN method outperforms the existing state-of-the-art methods for the cross-dataset point cloud recognition task. Wen Li 0001, Dong Xu 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Fixing Localization Errors to Improve Image Classification
Guolei Sun, Salman Khan 0001, Wen Li 0001, Hisham Cholakkal, Fahad Shahbaz Khan, Luc Van Gool |
ECCV (25) | 3 |
| 2020 | Off-Policy Reinforcement Learning for Efficient and Effective GAN Architecture Search
Yuan Tian 0014, Qin Wang 0013, Zhiwu Huang, Wen Li 0001, Dengxin Dai, Jun Wang 0012, Olga Fink |
ECCV (7) | 4 |
| 2020 | Tensorized Multi-view Subspace Representation Learning
Changqing Zhang 0002, Huazhu Fu, Jing Wang 0023, Wen Li 0001, Xiaochun Cao, Qinghua Hu |
Int. J. Comput. Vis. | 4 |
| 2020 | Discovering and incorporating latent target-domains for domain adaptation
Haoliang Li, Wen Li 0001, Shiqi Wang 0001 |
Pattern Recognit. | 2 |
| 2019 | Learning Semantic Segmentation From Synthetic Data: A Geometrically Guided Input-Output Adaptation ApproachabstractAs an alternative to manual pixel-wise annotation, synthetic data has been increasingly used for training semantic segmentation models. Such synthetic images and semantic labels can be easily generated from virtual 3D environments. In this work, we propose an approach to cross-domain semantic segmentation with the auxiliary geometric information, which can also be easily obtained from virtual environments. The geometric information is utilized on two levels for reducing domain shift: on the input level, we augment the standard image translation network with the geometric information to translate synthetic images into realistic style; on the output level, we build a task network which simultaneously performs semantic segmentation and depth estimation. Meanwhile, adversarial training is applied on the joint output space to preserve the correlation between semantics and depth. The proposed approach is validated on two pairs of synthetic to real dataset: from Virtual KITTI to KITTI, and from SYNTHIA to Cityscapes, where we achieve a clear performance gain compared to the baselines and various competing methods, demonstrating the effectiveness of the geometric information for cross-domain semantic segmentation. Wen Li 0001, Xiaoran Chen, Luc Van Gool |
CVPR | 2 |
| 2019 | DLOW: Domain Flow for Adaptation and GeneralizationabstractIn this work, we present a domain flow generation(DLOW) model to bridge two different domains by generating a continuous sequence of intermediate domains flowing from one domain to the other. The benefits of our DLOW model are two-fold. First, it is able to transfer source images into different styles in the intermediate domains. The transferred images smoothly bridge the gap between source and target domains, thus easing the domain adaptation task. Second, when multiple target domains are provided for training, our DLOW model is also able to generate new styles of images that are unseen in the training data. We implement our DLOW model based on CycleGAN. A domainness variable is introduced to guide the model to generate the desired intermediate domain images. In the inference phase, a flow of various styles of images can be obtained by varying the domainness variable. We demonstrate the effectiveness of our model for both cross-domain semantic segmentation and the style generalization tasks on benchmark datasets. Our implementation is available at https://github.com/ETHRuiGong/DLOW . Wen Li 0001, Luc Van Gool |
CVPR | 2 |
| 2019 | Sliced Wasserstein Generative ModelsabstractIn generative modeling, the Wasserstein distance (WD) has emerged as a useful metric to measure the discrepancy between generated and real data distributions. Unfortunately, it is challenging to approximate the WD of high-dimensional distributions. In contrast, the sliced Wasserstein distance (SWD) factorizes high-dimensional distributions into their multiple one-dimensional marginal distributions and is thus easier to approximate. In this paper, we introduce novel approximations of the primal and dual SWD. Instead of using a large number of random projections, as it is done by conventional SWD approximation methods, we propose to approximate SWDs with a small number of parameterized orthogonal projections in an end-to-end deep learning fashion. As concrete applications of our SWD approximations, we design two types of differentiable SWD blocks to equip modern generative frameworks---Auto-Encoders (AE) and Generative Adversarial Networks (GAN). In the experiments, we not only show the superiority of the proposed generative models on standard image synthesis benchmarks, but also demonstrate the state-of-the-art performance on challenging high resolution image and video generation in an unsupervised manner. Jiqing Wu, Zhiwu Huang, Dinesh Acharya 0001, Wen Li 0001, Janine Thoma, Danda Pani Paudel, Luc Van Gool |
CVPR | 4 |
| 2019 | Fast Image Restoration With Multi-Bin Trainable Linear UnitsabstractTremendous advances in image restoration tasks such as denoising and super-resolution have been achieved using neural networks. Such approaches generally employ very deep architectures, large number of parameters, large receptive fields and high nonlinear modeling capacity. In order to obtain efficient and fast image restoration networks one should improve upon the above mentioned requirements. In this paper we propose a novel activation function, the multi-bin trainable linear unit (MTLU), for increasing the nonlinear modeling capacity together with lighter and shallower networks. We validate the proposed fast image restoration networks for image denoising (FDnet) and super-resolution (FSRnet) on standard benchmarks. We achieve large improvements in both memory and runtime over current state-of-the-art for comparable or better PSNR accuracies. Shuhang Gu, Wen Li 0001, Luc Van Gool, Radu Timofte |
ICCV | 2 |
| 2019 | Semi-Supervised Learning by Augmented Distribution AlignmentabstractIn this work, we propose a simple yet effective semi-supervised learning approach called Augmented Distribution Alignment. We reveal that an essential sampling bias exists in semi-supervised learning due to the limited number of labeled samples, which often leads to a considerable empirical distribution mismatch between labeled data and unlabeled data. To this end, we propose to align the empirical distributions of labeled and unlabeled data to alleviate the bias. On one hand, we adopt an adversarial training strategy to minimize the distribution distance between labeled and unlabeled data as inspired by domain adaptation works. On the other hand, to deal with the small sample size issue of labeled data, we also propose a simple interpolation strategy to generate pseudo training samples. Those two strategies can be easily implemented into existing deep neural networks. We demonstrate the effectiveness of our proposed approach on the benchmark SVHN and CIFAR10 datasets. Our code is available at https://github.com/qinenergy/adanet . Qin Wang 0013, Wen Li 0001, Luc Van Gool |
ICCV | 2 |
| 2018 | ROAD: Reality Oriented Adaptation for Semantic Segmentation of Urban ScenesabstractExploiting synthetic data to learn deep models has attracted increasing attention in recent years. However, the intrinsic domain difference between synthetic and real images usually causes a significant performance drop when applying the learned model to real world scenarios. This is mainly due to two reasons: 1) the model overfits to synthetic images, making the convolutional filters incompetent to extract informative representation for real images; 2) there is a distribution difference between synthetic and real data, which is also known as the domain adaptation problem. To this end, we propose a new reality oriented adaptation approach for urban scene semantic segmentation by learning from synthetic data. First, we propose a target guided distillation approach to learn the real image style, which is achieved by training the segmentation model to imitate a pretrained real style model using real images. Second, we further take advantage of the intrinsic spatial structure presented in urban scene images, and propose a spatial-aware adaptation scheme to effectively align the distribution of two domains. These two modules can be readily integrated with existing state-of-the-art semantic segmentation networks to improve their generalizability when adapting from synthetic to real urban scenes. We evaluate the proposed method on Cityscapes dataset by adapting from GTAV and SYNTHIA datasets, where the results demonstrate the effectiveness of our method. Wen Li 0001, Luc Van Gool |
CVPR | 2 |
| 2018 | Domain Adaptive Faster R-CNN for Object Detection in the WildabstractObject detection typically assumes that training and test data are drawn from an identical distribution, which, however, does not always hold in practice. Such a distribution mismatch will lead to a significant performance drop. In this work, we aim to improve the cross-domain robustness of object detection. We tackle the domain shift on two levels: (1) the image-level shift, such as image style, illumination, etc., and (2) the instance-level shift, such as object appearance, size, etc. We build our approach based on the recent state-of-the-art Faster R-CNN model, and design two domain adaptation components, on image level and instance level, to reduce the domain discrepancy. The two domain adaptation components are based onH-divergence theory, and are implemented by learning a domain classifier in adversarial training manner. The domain classifiers on different levels are further reinforced with a consistency regularization to learn a domain-invariant region proposal network (RPN) in the Faster R-CNN model. We evaluate our newly proposed approach using multiple datasets including Cityscapes, KITTI, SIM10K, etc. The results demonstrate the effectiveness of our proposed approach for robust object detection in various domain shift scenarios. Wen Li 0001, Christos Sakaridis, Dengxin Dai, Luc Van Gool |
CVPR | 2 |
| 2018 | Appearance-and-Relation Networks for Video ClassificationabstractSpatiotemporal feature learning in videos is a fundamental problem in computer vision. This paper presents a new architecture, termed as Appearance-and-Relation Network (ARTNet), to learn video representation in an end-to-end manner. ARTNets are constructed by stacking multiple generic building blocks, called as SMART, whose goal is to simultaneously model appearance and relation from RGB input in a separate and explicit manner. Specifically, SMART blocks decouple the spatiotemporal learning module into an appearance branch for spatial modeling and a relation branch for temporal modeling. The appearance branch is implemented based on the linear combination of pixels or filter responses in each frame, while the relation branch is designed based on the multiplicative interactions between pixels or filter responses across multiple frames. We perform experiments on three action recognition benchmarks: Kinetics, UCF101, and HMDB51, demonstrating that SMART blocks obtain an evident improvement over 3D convolutions for spatiotemporal feature learning. Under the same training setting, ARTNets achieve superior performance on these three datasets to the existing state-of-the-art methods.1 Limin Wang 0002, Wei Li 0044, Wen Li 0001, Luc Van Gool |
CVPR | 3 |
| 2018 | Collaborative and Adversarial Network for Unsupervised Domain AdaptationabstractIn this paper, we propose a new unsupervised domain adaptation approach called Collaborative and Adversarial Network (CAN) through domain-collaborative and domain-adversarial training of neural networks. We add several domain classifiers on multiple CNN feature extraction blocks1, in which each domain classifier is connected to the hidden representations from one block and one loss function is defined based on the hidden presentation and the domain labels (e.g., source and target). We design a new loss function by integrating the losses from all blocks in order to learn domain informative representations from lower blocks through collaborative learning and learn domain uninformative representations from higher blocks through adversarial learning. We further extend our CAN method as Incremental CAN (iCAN), in which we iteratively select a set of pseudo-labelled target samples based on the image classifier and the last domain classifier from the previous training epoch and re-train our CAN model by using the enlarged training set. Comprehensive experiments on two benchmark datasets Office and ImageCLEF-DA clearly demonstrate the effectiveness of our newly proposed approaches CAN and iCAN for unsupervised domain adaptation. Wanli Ouyang, Wen Li 0001, Dong Xu 0001 |
CVPR | 3 |
| 2018 | Dividing and Aggregating Network for Multi-view Action Recognition
Dongang Wang, Wanli Ouyang, Wen Li 0001, Dong Xu 0001 |
ECCV (9) | 3 |
| 2018 | Semi-Supervised Optimal Transport for Heterogeneous Domain AdaptationabstractHeterogeneous domain adaptation (HDA) aims to exploit knowledge from a heterogeneous source domain to improve the learning performance in a target domain. Since the feature spaces of the source and target domains are different, the transferring of knowledge is extremely difficult. In this paper, we propose a novel semi-supervised algorithm for HDA by exploiting the theory of optimal transport (OT), a powerful tool originally designed for aligning two different distributions. To match the samples between heterogeneous domains, we propose to preserve the semantic consistency between heterogeneous domains by incorporating label information into the entropic Gromov-Wasserstein discrepancy, which is a metric in OT for different metric spaces, resulting in a new semi-supervised scheme. Via the new scheme, the target and transported source samples with the same label are enforced to follow similar distributions. Lastly, based on the Kullback-Leibler metric, we develop an efficient algorithm to optimize the resultant problem. Comprehensive experiments on both synthetic and real-world datasets demonstrate the effectiveness of our proposed method. Yuguang Yan, Wen Li 0001, Hanrui Wu, Huaqing Min, Mingkui Tan, Qingyao Wu |
IJCAI | 2 |
| 2018 | Visual Recognition in RGB Images and Videos by Learning from RGB-D DataabstractIn this work, we propose a framework for recognizing RGB images or videos by learning from RGB-D training data that contains additional depth information. We formulate this task as a new unsupervised domain adaptation (UDA) problem, in which we aim to take advantage of the additional depth features in the source domain and also cope with the data distribution mismatch between the source and target domains. To handle the domain distribution mismatch, we propose to learn an optimal projection matrix to map the samples from both domains into a common subspace such that the domain distribution mismatch can be reduced. Such projection matrix can be effectively optimized by exploiting different strategies. Moreover, we also use different ways to utilize the additional depth features. To simultaneously cope with the above two issues, we formulate a unified learning framework called domain adaptation from multi-view to single-view (DAM2S). By defining various forms of regularizers in our DAM2S framework, different strategies can be readily incorporated to learn robust SVM classifiers for classifying the target samples, and three methods are developed under our DAM2S framework. We conduct comprehensive experiments for object recognition, cross-dataset and cross-view action recognition, which demonstrate the effectiveness of our proposed methods for recognizing RGB images and videos by learning from RGB-D data. Wen Li 0001, Lin Chen 0021, Dong Xu 0001, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Domain Generalization and Adaptation Using Low Rank Exemplar SVMsabstractDomain adaptation between diverse source and target domains is challenging, especially in the real-world visual recognition tasks where the images and videos consist of significant variations in viewpoints, illuminations, qualities, etc. In this paper, we propose a new approach for domain generalization and domain adaptation based on exemplar SVMs. Specifically, we decompose the source domain into many subdomains, each of which contains only one positive training sample and all negative samples. Each subdomain is relatively less diverse, and is expected to have a simpler distribution. By training one exemplar SVM for each subdomain, we obtain a set of exemplar SVMs. To further exploit the inherent structure of source domain, we introduce a nuclear-norm based regularizer into the objective function in order to enforce the exemplar SVMs to produce a low-rank output on training samples. In the prediction process, the confident exemplar SVM classifiers are selected and reweigted according to the distribution mismatch between each subdomain and the test sample in the target domain. We formulate our approach based on the logistic regression and least square SVM algorithms, which are referred to as low rank exemplar SVMs (LRE-SVMs) and low rank exemplar least square SVMs (LRE-LSSVMs), respectively. A fast algorithm is also developed for accelerating the training of LRE-LSSVMs. We further extend Domain Adaptation Machine (DAM) to learn an optimal target classifier for domain adaptation, and show that our approach can also be applied to domain adaptation with evolving target domain, where the target data distribution is gradually changing. The comprehensive experiments for object recognition and action recognition demonstrate the effectiveness of our approach for domain generalization and domain adaptation with fixed and evolving target domains. Wen Li 0001, Zheng Xu 0002, Dong Xu 0001, Dengxin Dai, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Unsupervised Domain Adaptation for Face Anti-SpoofingabstractFace anti-spoofing (a.k.a. presentation attack detection) has recently emerged as an active topic with great significance for both academia and industry due to the rapidly increasing demand in user authentication on mobile phones, PCs, tablets, and so on. Recently, numerous face spoofing detection schemes have been proposed based on the assumption that training and testing samples are in the same domain in terms of the feature space and marginal probability distribution. However, due to unlimited variations of the dominant conditions (illumination, facial appearance, camera quality, and so on) in face acquisition, such single domain methods lack generalization capability, which further prevents them from being applied in practical applications. In light of this, we introduce an unsupervised domain adaptation face anti-spoofing scheme to address the real-world scenario that learns the classifier for the target domain based on training samples in a different source domain. In particular, an embedding function is first imposed based on source and target domain data, which maps the data to a new space where the distribution similarity can be measured. Subsequently, the Maximum Mean Discrepancy between the latent features in source and target domains is minimized such that a more generalized classifier can be learned. State-of-the-art representations including both hand-crafted and deep neural network learned features are further adopted into the framework to quest the capability of them in domain adaptation. Moreover, we introduce a new database for face spoofing detection, which contains more than 4000 face samples with a large variety of spoofing types, capture devices, illuminations, and so on. Extensive experiments on existing benchmark databases and the new database verify that the proposed approach can gain significantly better generalization capability in cross-domain scenarios by providing consistently better anti-spoofing performance. Haoliang Li, Wen Li 0001, Shiqi Wang 0001, Feiyue Huang, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | An Exemplar-Based Multi-View Domain Generalization Framework for Visual RecognitionabstractIn this paper, we propose a new exemplar-based multi-view domain generalization (EMVDG) framework for visual recognition by learning robust classifier that are able to generalize well to arbitrary target domain based on the training samples with multiple types of features (i.e., multi-view features). In this framework, we aim to address two issues simultaneously. First, the distribution of training samples (i.e., the source domain) is often considerably different from that of testing samples (i.e., the target domain), so the performance of the classifiers learnt on the source domain may drop significantly on the target domain. Moreover, the testing data are often unseen during the training procedure. Second, when the training data are associated with multi-view features, the recognition performance can be further improved by exploiting the relation among multiple types of features. To address the first issue, considering that it has been shown that fusing multiple SVM classifiers can enhance the domain generalization ability, we build our EMVDG framework upon exemplar SVMs (ESVMs), in which a set of ESVM classifiers are learnt with each one trained based on one positive training sample and all the negative training samples. When the source domain contains multiple latent domains, the learnt ESVM classifiers are expected to be grouped into multiple clusters. To address the second issue, we propose two approaches under the EMVDG framework based on the consensus principle and the complementary principle, respectively. Specifically, we propose an EMVDG_CO method by adding a co-regularizer to enforce the cluster structures of ESVM classifiers on different views to be consistent based on the consensus principle. Inspired by multiple kernel learning, we also propose another EMVDG_MK method by fusing the ESVM classifiers from different views based on the complementary principle. In addition, we further extend our EMVDG framework to exemplar-based multi-view domain adaptation (EMVDA) framework when the unlabeled target domain data are available during the training procedure. The effectiveness of our EMVDG and EMVDA frameworks for visual recognition is clearly demonstrated by comprehensive experiments on three benchmark data sets. Li Niu 0002, Wen Li 0001, Dong Xu 0001, Jianfei Cai 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Learning Discriminative Correlation Subspace for Heterogeneous Domain AdaptationabstractDomain adaptation aims to reduce the effort on collecting and annotating target data by leveraging knowledge from a different source domain. The domain adaptation problem will become extremely challenging when the feature spaces of the source and target domains are different, which is also known as the heterogeneous domain adaptation (HDA) problem. In this paper, we propose a novel HDA method to find the optimal discriminative correlation subspace for the source and target data. The discriminative correlation subspace is inherited from the canonical correlation subspace between the source and target data, and is further optimized to maximize the discriminative ability for the target domain classifier. We formulate a joint objective in order to simultaneously learn the discriminative correlation subspace and the target domain classifier. We then apply an alternating direction method of multiplier (ADMM) algorithm to address the resulting non-convex optimization problem. Comprehensive experiments on two real-world data sets demonstrate the effectiveness of the proposed method compared to the state-of-the-art methods. Yuguang Yan, Wen Li 0001, Michael Kwok-Po Ng, Mingkui Tan, Hanrui Wu, Huaqing Min, Qingyao Wu |
IJCAI | 2 |
| 2017 | Visual Recognition by Learning From Web Data via Weakly Supervised Domain GeneralizationabstractIn this paper, a weakly supervised domain generalization (WSDG) method is proposed for real-world visual recognition tasks, in which we train classifiers by using Web data (e.g., Web images and Web videos) with noisy labels. In particular, two challenging problems need to be solved when learning robust classifiers, in which the first issue is to cope with the label noise of training Web data from the source domain, while the second issue is to enhance the generalization capability of learned classifiers to an arbitrary target domain. In order to handle the first problem, the training samples within each category are partitioned into clusters, where we use one bag to denote each cluster and instances to denote the samples in each cluster. Then, we identify a proportion of good training samples in each bag and train robust classifiers by using the good training samples, which leads to a multi-instance learning (MIL) problem. In order to handle the second problem, we assume that the training samples possibly form a set of hidden domains, with each hidden domain associated with a distinctive data distribution. Then, for each category and each hidden latent domain, we propose to learn one classifier by extending our MIL formulation, which leads to our WSDG approach. In the testing stage, our approach can obtain better generalization capability by effectively integrating multiple classifiers from different latent domains in each category. Moreover, our WSDG approach is further extended to utilize additional textual descriptions associated with Web data as privileged information (PI), although testing data do not have such PI. Extensive experiments on three benchmark data sets indicate that our newly proposed methods are effective for real-world visual recognition tasks by learning from Web data. Li Niu 0002, Wen Li 0001, Dong Xu 0001, Jianfei Cai 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Fast Algorithms for Linear and Kernel SVM+abstractThe SVM+ approach has shown excellent performance in visual recognition tasks for exploiting privileged information in the training data. In this paper, we propose two efficient algorithms for solving the linear and kernel SVM+, respectively. For linear SVM+, we absorb the bias term into the weight vector, and formulate a new optimization problem with simpler constraints in the dual form. Then, we develop an efficient dual coordinate descent algorithm to solve the new optimization problem. For kernel SVM+, we further apply the l2-loss, which leads to a simpler optimization problem in the dual form with only half of dual variables when compared with the dual form of the original SVM+ method. More interestingly, we show that our new dual problem can be efficiently solved by using the SMO algorithm of the one-class SVM problem. Comprehensive experiments on three datasets clearly demonstrate that our proposed algorithms achieve significant speed-up than the state-of-the-art solvers for linear and kernel SVM+. Wen Li 0001, Dengxin Dai, Mingkui Tan, Dong Xu 0001, Luc Van Gool |
CVPR | 1 |
| 2016 | Exploiting Privileged Information from Web Data for Action and Event Recognition
Li Niu 0002, Wen Li 0001, Dong Xu 0001 |
Int. J. Comput. Vis. | 2 |
| 2016 | Co-Labeling for Multi-View Weakly Labeled LearningabstractIt is often expensive and time consuming to collect labeled training samples in many real-world applications. To reduce human effort on annotating training samples, many machine learning techniques (e.g., semi-supervised learning (SSL), multi-instance learning (MIL), etc.) have been studied to exploit weakly labeled training samples. Meanwhile, when the training data is represented with multiple types of features, many multi-view learning methods have shown that classifiers trained on different views can help each other to better utilize the unlabeled training samples for the SSL task. In this paper, we study a new learning problem called multi-view weakly labeled learning, in which we aim to develop a unified approach to learn robust classifiers by effectively utilizing different types of weakly labeled multi-view data from a broad range of tasks including SSL, MIL and relative outlier detection (ROD). We propose an effective approach called co-labeling to solve the multi-view weakly labeled learning problem. Specifically, we model the learning problem on each view as a weakly labeled learning problem, which aims to learn an optimal classifier from a set of pseudo-label vectors generated by using the classifiers trained from other views. Unlike traditional co-training approaches using a single pseudo-label vector for training each classifier, our co-labeling approach explores different strategies to utilize the predictions from different views, biases and iterations for generating the pseudo-label vectors, making our approach more robust for real-world applications. Moreover, to further improve the weakly labeled learning on each view, we also exploit the inherent group structure in the pseudo-label vectors generated from different strategies, which leads to a new multi-layer multiple kernel learning problem. Promising results for text-based image retrieval on the NUS-WIDE dataset as well as news classification and text categorization on several real-world multi-view datasets clearly demonstrate that our proposed co-labeling approach achieves state-of-the-art performance for various multi-view weakly labeled learning problems including multi-view SSL, multi-view MIL and multi-view ROD. Xinxing Xu, Wen Li 0001, Dong Xu 0001, Ivor W. Tsang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Image Classification by Cross-Media Active Learning With Privileged InformationabstractIn this paper, we propose a novel cross-media active learning algorithm to reduce the effort on labeling images for training. The Internet images are often associated with rich textual descriptions. Even though such textual information is not available in test images, it is still useful for learning robust classifiers. In light of this, we apply the recently proposed supervised learning paradigm, learning using privileged information, to the active learning task. Specifically, we train classifiers on both visual features and privileged information, and measure the uncertainty of unlabeled data by exploiting the learned classifiers and slacking function. Then, we propose to select unlabeled samples by jointly measuring the cross-media uncertainty and the visual diversity. Our method automatically learns the optimal tradeoff parameter between the two measurements, which in turn makes our algorithms particularly suitable for real-world applications. Extensive experiments demonstrate the effectiveness of our approach. Yan Yan 0006, Feiping Nie 0001, Wen Li 0001, Chenqiang Gao, Yi Yang 0001, Dong Xu 0001 |
IEEE Trans. Multim. | 3 |
| 2015 | Visual recognition by learning from web data: A weakly supervised domain generalization approachabstractIn this work, we formulate a new weakly supervised domain generalization approach for visual recognition by using loosely labeled web images/videos as training data. Specifically, we aim to address two challenging issues when learning robust classifiers: 1) coping with noise in the labels of training web images/videos in the source domain; and 2) enhancing generalization capability of learnt classifiers to any unseen target domain. To address the first issue, we partition the training samples in each class into multiple clusters. By treating each cluster as a “bag” and the samples in each cluster as “instances”, we formulate a multi-instance learning (MIL) problem by selecting a subset of training samples from each training bag and simultaneously learning the optimal classifiers based on the selected samples. To address the second issue, we assume the training web images/videos may come from multiple hidden domains with different data distributions. We then extend our MIL formulation to learn one classifier for each class and each latent domain such that multiple classifiers from each class can be effectively integrated to achieve better generalization capability. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our new approach for visual recognition by learning from web data. Li Niu 0002, Wen Li 0001, Dong Xu 0001 |
CVPR | 2 |
| 2015 | FaLRR: A fast low rank representation solverabstractLow rank representation (LRR) has shown promising performance for various computer vision applications such as face clustering. Existing algorithms for solving LRR usually depend on its two-variable formulation which contains the original data matrix. In this paper, we develop a fast LRR solver called FaLRR, by reformulating LRR as a new optimization problem with regard to factorized data (which is obtained by skinny SVD of the original data matrix). The new formulation benefits the corresponding optimization and theoretical analysis. Specifically, to solve the resultant optimization problem, we propose a new algorithm which is not only efficient but also theoretically guaranteed to obtain a globally optimal solution. Regarding the theoretical analysis, the new formulation is helpful for deriving some interesting properties of LRR. Last but not least, the proposed algorithm can be readily incorporated into an existing distributed framework of LRR for further acceleration. Extensive experiments on synthetic and real-world datasets demonstrate that our FaLRR achieves order-of-magnitude speedup over existing LRR solvers, and the efficiency can be further improved by incorporating our algorithm into the distributed framework of LRR. Shijie Xiao, Wen Li 0001, Dong Xu 0001, Dacheng Tao |
CVPR | 2 |
| 2015 | Multi-view Domain Generalization for Visual RecognitionabstractIn this paper, we propose a new multi-view domain generalization (MVDG) approach for visual recognition, in which we aim to use the source domain samples with multiple types of features (i.e., multi-view features) to learn robust classifiers that can generalize well to any unseen target domain. Considering the recent works show the domain generalization capability can be enhanced by fusing multiple SVM classifiers, we build upon exemplar SVMs to learn a set of SVM classifiers by using one positive sample and all negative samples in the source domain each time. When the source domain samples come from multiple latent domains, we expect the weight vectors of exemplar SVM classifiers can be organized into multiple hidden clusters. To exploit such cluster structure, we organize the weight vectors learnt on each view as a weight matrix and seek the low-rank representation by reconstructing this weight matrix using itself as the dictionary. To enforce the consistency of inherent cluster structures discovered from the weight matrices learnt on different views, we introduce a new regularizer to minimize the mismatch between any two representation matrices on different views. We also develop an efficient alternating optimization algorithm and further extend our MVDG approach for domain adaptation by exploiting the manifold structure of unlabeled target domain samples. Comprehensive experiments for visual recognition clearly demonstrate the effectiveness of our approaches for domain generalization and domain adaptation. Li Niu 0002, Wen Li 0001, Dong Xu 0001 |
ICCV | 2 |
| 2015 | Distance Metric Learning Using Privileged Information for Face Verification and Person Re-IdentificationabstractIn this paper, we propose a new approach to improve face verification and person re-identification in the RGB images by leveraging a set of RGB-D data, in which we have additional depth images in the training data captured using depth cameras such as Kinect. In particular, we extract visual features and depth features from the RGB images and depth images, respectively. As the depth features are available only in the training data, we treat the depth features as privileged information, and we formulate this task as a distance metric learning with privileged information problem. Unlike the traditional face verification and person re-identification tasks that only use visual features, we further employ the extra depth features in the training data to improve the learning of distance metric in the training process. Based on the information-theoretic metric learning (ITML) method, we propose a new formulation called ITML with privileged information (ITML+) for this task. We also present an efficient algorithm based on the cyclic projection method for solving the proposed ITML+ formulation. Extensive experiments on the challenging faces data sets EUROCOM and CurtinFaces for face verification as well as the BIWI RGBD-ID data set for person re-identification demonstrate the effectiveness of our proposed approach. Xinxing Xu, Wen Li 0001, Dong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | Recognizing RGB Images by Learning from RGB-D DataabstractIn this work, we propose a new framework for recognizing RGB images captured by the conventional cameras by leveraging a set of labeled RGB-D data, in which the depth features can be additionally extracted from the depth images. We formulate this task as a new unsupervised domain adaptation (UDA) problem, in which we aim to take advantage of the additional depth features in the source domain and also cope with the data distribution mismatch between the source and target domains. To effectively utilize the additional depth features, we seek two optimal projection matrices to map the samples from both domains into a common space by preserving as much as possible the correlations between the visual features and depth features. To effectively employ the training samples from the source domain for learning the target classifier, we reduce the data distribution mismatch by minimizing the Maximum Mean Discrepancy (MMD) criterion, which compares the data distributions for each type of feature in the common space. Based on the above two motivations, we propose a new SVM based objective function to simultaneously learn the two projection matrices and the optimal target classifier in order to well separate the source samples from different classes when using each type of feature in the common space. An efficient alternating optimization algorithm is developed to solve our new objective function. Comprehensive experiments for object recognition and gender recognition demonstrate the effectiveness of our proposed approach for recognizing RGB images by learning from RGB-D data. Lin Chen 0021, Wen Li 0001, Dong Xu 0001 |
CVPR | 2 |
| 2014 | Exploiting Privileged Information from Web Data for Image Categorization
Wen Li 0001, Li Niu 0002, Dong Xu 0001 |
ECCV (5) | 1 |
| 2014 | Exploiting Low-Rank Structure from Latent Domains for Domain Generalization
Zheng Xu 0002, Wen Li 0001, Li Niu 0002, Dong Xu 0001 |
ECCV (3) | 2 |
| 2014 | Incorporating Privileged Genetic Information for Fundus Image Based Glaucoma Detection
Lixin Duan, Yanwu Xu 0001, Wen Li 0001, Lin Chen 0021, Damon Wing Kee Wong, Tien Yin Wong, Jiang Liu 0001 |
MICCAI (2) | 3 |
| 2014 | Learning With Augmented Features for Supervised and Semi-Supervised Heterogeneous Domain AdaptationabstractIn this paper, we study the heterogeneous domain adaptation (HDA) problem, in which the data from the source domain and the target domain are represented by heterogeneous features with different dimensions. By introducing two different projection matrices, we first transform the data from two domains into a common subspace such that the similarity between samples across different domains can be measured. We then propose a new feature mapping function for each domain, which augments the transformed samples with their original features and zeros. Existing supervised learning methods (e.g., SVM and SVR) can be readily employed by incorporating our newly proposed augmented feature representations for supervised HDA. As a showcase, we propose a novel method called Heterogeneous Feature Augmentation (HFA) based on SVM. We show that the proposed formulation can be equivalently derived as a standard Multiple Kernel Learning (MKL) problem, which is convex and thus the global solution can be guaranteed. To additionally utilize the unlabeled data in the target domain, we further propose the semi-supervised HFA (SHFA) which can simultaneously learn the target classifier as well as infer the labels of unlabeled target samples. Comprehensive experiments on three different applications clearly demonstrate that our SHFA and HFA outperform the existing HDA methods. Wen Li 0001, Lixin Duan, Dong Xu 0001, Ivor W. Tsang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Flowing on Riemannian Manifold: Domain Adaptation by Shifting CovarianceabstractDomain adaptation has shown promising results in computer vision applications. In this paper, we propose a new unsupervised domain adaptation method called domain adaptation by shifting covariance (DASC) for object recognition without requiring any labeled samples from the target domain. By characterizing samples from each domain as one covariance matrix, the source and target domain are represented into two distinct points residing on a Riemannian manifold. Along the geodesic constructed from the two points, we then interpolate some intermediate points (i.e., covariance matrices), which are used to bridge the two domains. By utilizing the principal components of each covariance matrix, samples from each domain are further projected into intermediate feature spaces, which finally leads to domain-invariant features after the concatenation of these features from intermediate points. In the multiple source domain adaptation task, we also need to effectively integrate different types of features between each pair of source and target domains. We additionally propose an SVM based method to simultaneously learn the optimal target classifier as well as the optimal weights for different source domains. Extensive experiments demonstrate the effectiveness of our method for both single source and multiple source domain adaptation tasks. Zhen Cui 0001, Wen Li 0001, Dong Xu 0001, Shiguang Shan, Xilin Chen 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2013 | Fusing Robust Face Region Descriptors via Multiple Metric Learning for Face Recognition in the WildabstractIn many real-world face recognition scenarios, face images can hardly be aligned accurately due to complex appearance variations or low-quality images. To address this issue, we propose a new approach to extract robust face region descriptors. Specifically, we divide each image (resp. video) into several spatial blocks (resp. spatial-temporal volumes) and then represent each block (resp. volume) by sum-pooling the nonnegative sparse codes of position-free patches sampled within the block (resp. volume). Whitened Principal Component Analysis (WPCA) is further utilized to reduce the feature dimension, which leads to our Spatial Face Region Descriptor (SFRD) (resp. Spatial-Temporal Face Region Descriptor, STFRD) for images (resp. videos). Moreover, we develop a new distance metric learning method for face verification called Pairwise-constrained Multiple Metric Learning (PMML) to effectively integrate the face region descriptors of all blocks (resp. volumes) from an image (resp. a video). Our work achieves the state-of-the-art performances on two real-world datasets LFW and YouTube Faces (YTF) according to the restricted protocol. Zhen Cui 0001, Wen Li 0001, Dong Xu 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2013 | Semantically-Based Human Scanpath Estimation with HMMsabstractWe present a method for estimating human scan paths, which are sequences of gaze shifts that follow visual attention over an image. In this work, scan paths are modeled based on three principal factors that influence human attention, namely low-level feature saliency, spatial position, and semantic content. Low-level feature saliency is formulated as transition probabilities between different image regions based on feature differences. The effect of spatial position on gaze shifts is modeled as a Levy flight with the shifts following a 2D Cauchy distribution. To account for semantic content, we propose to use a Hidden Markov Model (HMM) with a Bag-of-Visual-Words descriptor of image regions. An HMM is well-suited for this purpose in that 1) the hidden states, obtained by unsupervised learning, can represent latent semantic concepts, 2) the prior distribution of the hidden states describes visual attraction to the semantic concepts, and 3) the transition probabilities represent human gaze shift patterns. The proposed method is applied to task-driven viewing processes. Experiments and analysis performed on human eye gaze data verify the effectiveness of this method. Dong Xu 0001, Qingming Huang, Wen Li 0001, Min Xu 0001, Stephen Lin 0001 |
ICCV | 4 |
| 2013 | Learning Prototype Hyperplanes for Face Verification in the WildabstractIn this paper, we propose a new scheme called Prototype Hyperplane Learning (PHL) for face verification in the wild using only weakly labeled training samples (i.e., we only know whether each pair of samples are from the same class or different classes without knowing the class label of each sample) by leveraging a large number of unlabeled samples in a generic data set. Our scheme represents each sample in the weakly labeled data set as a mid-level feature with each entry as the corresponding decision value from the classification hyperplane (referred to as the prototype hyperplane) of one Support Vector Machine (SVM) model, in which a sparse set of support vectors is selected from the unlabeled generic data set based on the learnt combination coefficients. To learn the optimal prototype hyperplanes for the extraction of mid-level features, we propose a Fisher’s Linear Discriminant-like (FLD-like) objective function by maximizing the discriminability on the weakly labeled data set with a constraint enforcing sparsity on the combination coefficients of each SVM model, which is solved by using an alternating optimization method. Then, we use the recent work called Side-Information based Linear Discriminant (SILD) analysis for dimensionality reduction and a cosine similarity measure for final face verification. Comprehensive experiments on two data sets, Labeled Faces in the Wild (LFW) and YouTube Faces, demonstrate the effectiveness of our scheme. Meina Kan, Dong Xu 0001, Shiguang Shan, Wen Li 0001, Xilin Chen 0001 |
IEEE Trans. Image Process. | 4 |
| 2012 | Batch mode Adaptive Multiple Instance Learning for computer vision tasksabstractMultiple Instance Learning (MIL) has been widely exploited in many computer vision tasks, such as image retrieval, object tracking and so on. To handle ambiguity of instance labels in positive bags, the training process of traditional MIL methods is usually computationally expensive, which limits the applications of MIL in more computer vision tasks. In this paper, we propose a novel batch mode framework, namely Batch mode Adaptive Multiple Instance Learning (BAMIL), to accelerate the instance-level MIL methods. Specifically, instead of using all training bags at once, we divide the training bags into several sets of bags (i.e., batches). At each time, we use one batch of training bags to train a new classifier which is adapted from the latest pre-learned classifier. Such batch mode framework significantly accelerates the traditional MIL methods for large scale applications and can be also used in dynamic environments such as object tracking. The experimental results show that our BAMIL is much faster than the recently developed MIL with constrained positive bags while achieves comparable performance for text-based web image retrieval. In dynamic settings, BAMIL also achieves the better overall performance for object tracking when compared with other online MIL methods. Wen Li 0001, Lixin Duan, Ivor W. Tsang, Dong Xu 0001 |
CVPR | 1 |
| 2012 | Co-labeling: A New Multi-view Learning Approach for Ambiguous ProblemsabstractWe propose a multi-view learning approach called co-labeling which is applicable for several machine learning problems where the labels of training samples are uncertain, including semi-supervised learning (SSL), multi-instance learning (MIL) and max-margin clustering (MMC). Particularly, we first unify those problems into a general ambiguous problem in which we simultaneously learn a robust classifier as well as find the optimal training labels from a finite label candidate set. To effectively utilize multiple views of data, we then develop our co-labeling approach for the general multi-view ambiguous problem. In our work, classifiers trained on different views can teach each other by iteratively passing the predictions of training samples from one classifier to the others. The predictions from one classifier are considered as label candidates for the other classifiers. To train a classifier with a label candidate set for each view, we adopt the Multiple Kernel Learning (MKL) technique by constructing the base kernel through associating the input kernel calculated from input features with one label candidate. Compared with the traditional co-training method which was specifically designed for SSL, the advantages of our co-labeling are two-fold: 1) it can be applied to other ambiguous problems such as MIL and MMC, 2) it is more robust by using the MKL method to integrate multiple labeling candidates obtained from different iterations and biases. Promising results on several real-world multi-view data sets clearly demonstrate the effectiveness of our proposed co-labeling for both MIL and SSL. Wen Li 0001, Lixin Duan, Ivor W. Tsang, Dong Xu 0001 |
ICDM | 1 |
| 2011 | Text-based image retrieval using progressive multi-instance learningabstractRelevant and irrelevant images collected from the Web (e.g., Flickr.com) have been employed as loosely labeled training data for image categorization and retrieval. In this work, we propose a new approach to learn a robust classifier for text-based image retrieval (TBIR) using relevant and irrelevant training web images, in which we explicitly handle noise in the loose labels of training images. Specifically, we first partition the relevant and irrelevant training web images into clusters. By treating each cluster as a "bag" and the images in each bag as "instances", we formulate this task as a multi-instance learning problem with constrained positive bags, in which each positive bag contains at least a portion of positive instances. We present a new algorithm called MIL-CPB to effectively exploit such constraints on positive bags and predict the labels of test instances (images). Observing that the constraints on positive bags may not always be satisfied in our application, we additionally propose a progressive scheme (referred to as Progressive MIL-CPB, or PMIL-CPB) to further improve the retrieval performance, in which we iteratively partition the top-ranked training web images from the current MIL-CPB classifier to construct more confident positive "bags "and then add these new "bags" as training data to learn the subsequent MIL-CPB classifiers. Comprehensive experiments on two challenging real-world web image data sets demonstrate the effectiveness of our approach. © 2011 IEEE. Wen Li 0001, Lixin Duan, Dong Xu 0001, Ivor W. Tsang |
ICCV | 1 |
| 2011 | Improving Web Image Search by Bag-Based RerankingabstractGiven a textual query in traditional text-based image retrieval (TBIR), relevant images are to be reranked using visual features after the initial text-based search. In this paper, we propose a new bag-based reranking framework for large-scale TBIR. Specifically, we first cluster relevant images using both textual and visual features. By treating each cluster as a "bag" and the images in the bag as "instances," we formulate this problem as a multi-instance (MI) learning problem. MI learning methods such as mi-SVM can be readily incorporated into our bag-based reranking framework. Observing that at least a certain portion of a positive bag is of positive instances while a negative bag might also contain positive instances, we further use a more suitable generalized MI (GMI) setting for this application. To address the ambiguities on the instance labels in the positive and negative bags under this GMI setting, we develop a new method referred to as GMI-SVM to enhance retrieval performance by propagating the labels from the bag level to the instance level. To acquire bag annotations for (G)MI learning, we propose a bag ranking method to rank all the bags according to the defined bag ranking score. The top ranked bags are used as pseudopositive training bags, while pseudonegative training bags can be obtained by randomly sampling a few irrelevant images that are not associated with the textual query. Comprehensive experiments on the challenging real-world data set NUS-WIDE demonstrate our framework with automatic bag annotation can achieve the best performances compared with existing image reranking methods. Our experiments also demonstrate that GMI-SVM can achieve better performances when using the manually labeled training bags obtained from relevance feedback. Lixin Duan, Wen Li 0001, Ivor W. Tsang, Dong Xu 0001 |
IEEE Trans. Image Process. | 2 |
| 2009 | Efficient Edge Matching using Improved Hierarchical Chamfer MatchingabstractMatching is a central problem in pattern recognition and computer vision, its applications includes object detection and tracking. HCMA (hierarchical chamfer matching) is a classical image matching algorithm, which utilizes the edge information to match the images robustly and the multi-resolution pyramid to accelerate the matching process. However, for images with cluttered background and high resolution, HCMA is relatively computationally expensive, which has impeded its success in practical applications, especially in real-time applications. In this paper, an improved hierarchical chamfer matching algorithm is proposed to reduce its computational cost without degrading its matching quality. According to the experimental results, the proposed improvements are able to save 75% ~ 95% of the computational time, without causing any false matching. Pengfei Xu 0005, Wen Li 0001, Zhongke Wu |
ISCAS | 3 |
| 2005 | Fast block-based image restoration employing the improved best neighborhood matching approachabstractThe best neighborhood matching (BNM) algorithm is an efficient approach for image restoration. However, its high computation overhead imposes an obstacle to its application. In this paper, a fast image restoration approach named jump and look around BNM (JLBNM) is proposed to reduce computation overhead of the BNM. The main idea of JLBNM is to employ two kinds of search mechanisms so that the whole search process can be sped up. Some optimization techniques for the restoration algorithm JLBNM are also developed, including adaptive threshold in the matching stage, the terminal threshold in the searching stage, and the application of an appropriate matching function in both the matching and recovering stages. Theoretical analysis and experiment results have shown that JLBNM not only can provide high quality for image restoration but also has low computation overhead. Wen Li 0001, David Zhang 0001, Xiangzhen Qiao |
IEEE Trans. Syst. Man Cybern. Part A | 1 |