VLDB 2026 Research / reviewers in the wild / expert
Long Lan
dblp:124/2136
· DBLP profile ↗
142ranked-venue papers
11as first author
110since 2021 · last 2026
0000-0002-4238-8985ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 85 · 5 first-author · 66 since 2021Graphics, computer vision, multimedia, augmented reality and games · 53 · 5 first-author · 40 since 2021Databases, data management, data science and information retrieval · 15 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Let Synthetic Data Shine: Domain Reassembly and Soft-Fusion for Single Domain Generalization
Hao Li 0025, Yubin Xiao, Ke Liang 0006, Mengzhu Wang, Long Lan, Kenli Li 0001, Xinwang Liu 0002 |
Int. J. Comput. Vis. | 5 |
| 2026 | FedPuzzle: Federated causal discovery from distributed heterogeneous variable sets
Yiyao Li, Yeting Guo, Ligong Cao, Haotian Wang 0001, Long Lan |
Inf. Sci. | 5 |
| 2026 | Distilling structural knowledge from CNNs to vision transformers for data-efficient visual recognition
Dingyao Chen, Xiao Teng, Xun Yang 0001, Long Lan |
Neural Networks | 5 |
| 2026 | LCA-Med: A lightweight cross-modal adaptive feature processing module for detecting imbalanced medical image distribution
Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Hudan Pan, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001 |
Neural Networks | 2 |
| 2026 | RefSAM: Efficiently adapting segmenting anything model for referring video object segmentation
Yonglin Li, Jing Zhang 0037, Xiao Teng, Xinwang Liu 0002, Long Lan |
Neural Networks | 6 |
| 2026 | On the Two Facets to Conquer Wild Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection serves as an unknown-handling mechanism for open-world classification, enabling the identification of OOD data that diverge semantically from in-distribution (ID) data. The learning strategy known as outlier exposure (OE) enhances this process by incorporating OOD data during model training, directly making models learn to discern between ID and OOD patterns. However, in practice, the collected OOD data often contain many ID semantics, of which the scenario is commonly referred to as wild OOD detection. It can markedly compromise the reliability of models in OOD detection, yet few studies have addressed this critical issue. In this paper, we theoretically analyze wild OOD detection from the instance and distribution facets, respectively, to better comprehend its challenges and accordingly introduce two general solutions. At the instance facet, ID/OOD indicators contain errors due to the wild nature, where some data are of OOD labels yet should be assigned as ID. Hence, we introduce a general framework that can dynamically estimate the true ID/OOD indicators solely based on wild OOD data, thereby mitigating their negative impacts. At the distribution facet, the wild OOD distribution is a mixture of ID and OOD distributions, where the ID sub-distribution can mislead the model. We therefore propose a resampling scheme to remove the potential ID sub-distribution, with resampling probabilities estimated from the known ID distribution, enabling OE training to better address wild OOD detection. We provide theoretical guarantees for both solutions and develop algorithms that enhance their practical efficacy, ultimately integrating them into a unified framework that leverages their complementary strengths. Ultimately, we validate our approaches through comprehensive empirical evaluations across a range of wild OOD detection scenarios, clearly demonstrating the superior performance and reliability of our methods when compared to advanced counterparts. Zhaohui Hu, Xinwang Liu 0002, Long Lan, Bo Han 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Continuous Review and Timely Correction: Enhancing the Resistance to Noisy Labels via Self-Not-True and Class-Wise DistillationabstractDeep neural networks possess remarkable learning capabilities but are vulnerable to overfitting in the presence of mislabeled data. A well-known memorization effect causes networks to first fit clean samples and later memorize noisy labels. Although early stopping can partially alleviate this issue, it cannot prevent the accumulation of incorrect knowledge or recover information lost due to mislabeled inputs. In this paper, we introduce an innovative mechanism for continuous review and timely correction of learned knowledge. Our approach allows the network to repeatedly revisit and reinforce correct information while promptly addressing any inaccuracies stemming from mislabeled data. We present a novel method called self-not-true-distillation (SNTD). This technique employs self-distillation, where the network from previous training iterations acts as a teacher, guiding the current network to review and solidify its understanding of accurate labels. Crucially, SNTD masks the true class label in the logits during this process, concentrating on the non-true classes to correct any erroneous knowledge that may have been acquired. We also recognize that different data classes follow distinct learning trajectories. A single teacher network might struggle to effectively guide the learning of all classes at once, which necessitates selecting different teacher networks for each specific class. Additionally, the influence of the teacher network's guidance varies throughout the training process. To address these challenges, we propose SNTD+, which integrates a class-wise distillation strategy along with a dynamic weight adjustment mechanism. Together, these enhancements significantly bolster SNTD's robustness in tackling complex scenarios characterized by label noise. Long Lan, Xinghao Wu, Bo Han 0003, Xinwang Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Object style diffusion for generalized object detection in urban scene
Hao Li 0025, Xiangyuan Yang, Mengzhu Wang, Long Lan, Ke Liang 0006, Xinwang Liu 0002, Kenli Li 0001 |
Pattern Recognit. | 4 |
| 2026 | Towards to real world vehicle privacy protection: A new dataset and benchmark
Jiayi Lin 0010, Chengming Zou, Long Lan, Yong Luo 0002, Yue Yu 0001, Yaowei Wang 0001, Wei Zeng 0006, Yonghong Tian 0001 |
Pattern Recognit. | 3 |
| 2026 | Reliable Exploration Strategy for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification addresses the task of associating individuals across disjoint camera views in the absence of annotated training data. While current methods often focus on hard samples to learn discriminative features, these hard samples are more prone to label noise, which can negatively impact model performance. To address this issue, we propose an information-complementary hybrid contrastive learning framework that consists of two branches for extracting richer semantic information. Specifically, the first branch performs cluster-level contrastive learning to capture general semantic patterns, while the second branch employs relation-guided instance-level contrastive learning, which leverages local structures to mine hard samples for more discriminative representations. To mitigate label noise during hard sample mining, we introduce a temporal-guided dynamic weighting module that uses temporal clustering results to evaluate the reliability of the mined hard samples. This module generates instance-level weights to reduce the impact of unreliable samples, ensuring more stable training. Extensive experiments on four popular benchmarks demonstrate the superiority of our method compared to state-of-the-art approaches. Xiao Teng, Long Lan |
IEEE Signal Process. Lett. | 4 |
| 2026 | Probability-Guided Contrastive Learning for Long-Tailed Domain GeneralizationabstractAfter training on a specific source domain, models can leverage domain generalization (DG) techniques to achieve superior and broader performance on new, unseen target domains. Existing DG often utilizes contrastive learning to learn domain-invariant features. The goal of contrastive learning is to learn effective representations of data, causing samples from the same category to cluster together in feature space, while samples from different categories are dispersed. Traditional contrastive learning is limited to a finite set of contrastive pairs for DG. To handle this problem, we consider sampling from an infinite number of contrastive pairs using a mixture of von Mises-Fisher (vMF) distributions on the unit hypersphere. We propose a novel method called Probability-guided Contrastive Learning (PgCL), which selects contrastive pairs based on estimated data distributions of samples from each category in feature space. Additionally, we derive the exact analytical formula for the expected contrastive loss. We conduct an empirical investigation of the error bounds of PgCL and demonstrate its performance by comparing it with several leading methods across a range of DG datasets. Mengzhu Wang, Houcheng Su, Shanshan Wang 0008, Long Lan, Liang Yang 0002, Li Shen 0008 |
IEEE Trans. Big Data | 5 |
| 2026 | Phrase Grounding-Based Style Transfer for Single-Domain Generalized Object DetectionabstractSingle-domain generalized object detection aims to enhance a model’s generalization to multiple unseen target domains using only data from a single source domain during training. This is a practical yet challenging scenario, as it requires the model to address domain shift without incorporating target domain data into the training process. In this paper, we propose a novel phrase-grounding-based style transfer (PGST) approach for the task. Specifically, we first define textual prompts to describe objects for potential unseen target domains. Then, we leverage the grounded language-image pre-training (GLIP) model to capture the styles of these target domains and perform style transfer from the source to the target domains. The style-transferred visual features from the source domain are semantically rich and closely approximate those of their hypothetical counterparts in the target domain. Finally, we employ these style-transferred visual features to fine-tune GLIP. By introducing these imaginary counterparts, the detector can be effectively generalized to unseen target domains using only a single source domain during training. Our method significantly improves mean average precision (mAP), with an average increase of 8.8% across five diverse weather-driving benchmarks. Notably, our approach outperforms or matches the performance of domain-adaptive object detection methods, which require target domain data for training, in several challenging scenarios. Wei Wang 0335, Cong Wang 0018, Mengzhu Wang, Xiang Zhang 0008, Long Lan, Xinwang Liu 0002, Kenli Li 0001, Xiaochun Cao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and SegmentationabstractReferring multi-object tracking (RMOT) is an emerging cross-modal task that aims to locate an arbitrary number of target objects and maintain their identities referred by a language expression in a video. This intricate task involves the reasoning of linguistic and visual modalities, along with the temporal association of target objects. However, the seminal work relies on loose feature fusion and neglects long-term information. In this study, we introduce a compact Transformer-based method, termed TenRMOT. We conduct feature fusion at both encoding and decoding stages to fully exploit the advantages of Transformer architecture. Specifically, we incrementally perform cross-modal fusion layer-by-layer during the encoding phase. In the decoding phase, we utilize language-guided queries to probe memory features for accurate prediction of the desired objects. Moreover, we introduce a query update module that explicitly leverages temporal prior information of the tracked objects to enhance the consistency of their trajectories. In addition, we introduce a novel task called Referring Multi-Object Tracking and Segmentation (RMOTS) and construct a new dataset named Ref-KITTI Segmentation. Our dataset consists of 18 videos with 818 expressions, and each expression averages 10.7 masks, which poses a greater challenge compared to the typical single mask in most existing referring video segmentation datasets. TenRMOT demonstrates superior performance on both the referring multi-object tracking and the segmentation tasks. Changcheng Xiao, Qiong Cao, Xiang Zhang 0008, Tao Wang 0006, Canqun Yang, Long Lan |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | C-WOE: Clustering for Out-of-Distribution Detection Learning With Wild Outlier ExposureabstractOut-of-distribution (OOD) detection plays a crucial role as a mechanism for handling anomalies in computer vision systems. Among existing approaches, outlier exposure (OE), which trains the model with an additional auxiliary OOD dataset, has demonstrated strong effectiveness. However, acquiring clean and well-curated auxiliary OOD data is often infeasible, particularly within large and complex systems. Alternatively, wild outliers, i.e., unlabeled samples collected directly in deployment environments, are abundant and easy to obtain, and recent studies have shown that they can substantially benefit OOD detection learning. Nevertheless, wild outliers typically contain a mixture of in-distribution (ID) and OOD samples. Directly using them as auxiliary OOD data unavoidably exposes the model to adverse supervision signals arising from the contained ID samples. Yet existing methods still lack an effective strategy that can fully leverage wild outliers while suppressing the negative influence introduced by their ID subset. To this end, we propose a simple yet effective method named Clustering for Wild Outlier Exposure (C-WOE), which alleviates the adverse effect of the ID samples contained within wild outliers by reweighting them. Specifically, C-WOE assigns higher weights to real OOD samples and lower weights to ID samples and dynamically updates these weights during training. Theoretically, we establish solid guarantees for the proposed method. Empirically, extensive experiments conducted on various real-world benchmarks and simulated datasets demonstrate that C-WOE notably achieves superior performance compared with state-of-the-art methods, validating its reliability in image processing applications. Long Lan, Zhaohui Hu, Tongliang Liu, Xinwang Liu 0002 |
IEEE Trans. Image Process. | 1 |
| 2026 | Exploring Direction Alignment and Discrepancy Standardization for Knowledge DistillationabstractKnowledge Distillation (KD) is a widely popular model compression technique that can effectively transfer knowledge from a pre-trained, large-scale teacher model to a more compact and lightweight student model. Traditional KD methods aim to improve the student’s representation capability by mimicking the teacher’s features, e.g., minimizing the \(\mathcal{L}_{2}\) distance between their intermediate features. However, due to the capacity gap between the student and the teacher, student often struggles to precisely mimic the features of the teacher. To address this challenge, we propose to boost the knowledge distillation for the visual recognition tasks via Direction Alignment and Discrepancy Standardization ( DADS) , which exploits the feature scaling technique to distill from both the feature direction and feature discrepancy. To this end, we devise an efficient feature alignment module to align the dimensions of teacher and student features. Moreover, we align the direction of student features and teacher features, which are pre-processed by normalization. Furthermore, we leverage the Kullback–Leibler (KL) divergence to refine the features alignment, minimizing discrepancy in the distribution of features across samples, which is pre-processed by \(\mathcal{Z}\) -score standardization. In this way, our proposed approach can effectively transfer the knowledge from the teacher to the student, facilitating the downstream visual recognition applications, such as image classification and semantic segmentation. Extensive experimental analyses clearly validate the effectiveness of DADS . Compared with previous KD methods, our approach sets a new benchmark, achieving state-of-the-art results on visual recognition tasks. Dingyao Chen, Xiao Teng, Xiang Zhang 0008, Xun Yang 0001, Long Lan |
ACM Trans. Knowl. Discov. Data | 5 |
| 2025 | Relieving Universal Label Noise for Unsupervised Visible-Infrared Person Re-Identification by Inferring from NeighborsabstractUnsupervised visible-infrared person re-identification (USL-VI-ReID) is of great research and practical significance yet remains challenging due to the absence of annotations. Existing approaches aim to learn modality-invariant representations in an unsupervised setting. However, these methods often encounter label noise within and across modalities due to suboptimal clustering results and considerable modality discrepancies, which impedes effective training. To address these challenges, we propose a straightforward yet effective solution for USL-VI-ReID by mitigating universal label noise using neighbor information. Specifically, we introduce the Neighbor-guided Universal Label Calibration (N-ULC) module, which replaces explicit hard pseudo labels in both homogeneous and heterogeneous spaces with soft labels derived from neighboring samples to reduce label noise. Additionally, we present the Neighbor-guided Dynamic Weighting (N-DW) module to enhance training stability by minimizing the influence of unreliable samples. Extensive experiments on the RegDB and SYSU-MM01 datasets demonstrate that our method outperforms existing USL-VI-ReID approaches, despite its simplicity. Xiao Teng, Long Lan, Dingyao Chen, Kele Xu |
AAAI | 2 |
| 2025 | MagicNaming: Consistent Identity Generation by Finding a "Name Space" in T2I Diffusion ModelsabstractLarge-scale text-to-image diffusion models, (e.g., DALL-E, SDXL) are capable of generating famous persons by simply referring to their names. Is it possible to make such models generate generic identities as simple as the famous ones, e.g., just use a name? In this paper, we explore the existence of a ``Name Space'', where any point in the space corresponds to a specific identity. Fortunately, we find some clues in the feature space spanned by text embedding of celebrities' names. Specifically, we first extract the embeddings of celebrities' names in the Laion5B dataset with the text encoder of diffusion models. Such embeddings are used as supervision to learn an encoder that can predict the name (actually an embedding) of a given face image. We experimentally find that such name embeddings work well in promising the generated image with good identity consistency. Note that like the names of celebrities, our predicted name embeddings are disentangled from the semantics of text inputs, making the original generation capability of text-to-image models well-preserved. Moreover, by simply plugging such name embeddings, all variants (e.g., from Civitai) derived from the same base model (i.e., SDXL) readily become identity-aware text-to-image models. Heliang Zheng, Long Lan, Wanrong Huang, Yuhua Tang |
AAAI | 4 |
| 2025 | XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?abstractThe astonishing breakthrough of multimodal large language models (MLLMs) has necessitated new benchmarks to quantitatively assess their capabilities, reveal their limitations, and indicate future research directions. However, this is challenging in the context of remote sensing (RS), since the imagery features ultra-high resolution that incorporates extremely complex semantic relationships. Existing benchmarks usually adopt notably smaller image sizes than real-world RS scenarios, suffer from limited annotation quality, and consider insufficient dimensions of evaluation. To address these issues, we present XLRS-Bench: a comprehensive benchmark for evaluating the perception and reasoning capabilities of MLLMs in ultra-high-resolution RS scenarios. XLRS-Bench boasts the largest average image size (8500×8500) observed thus far, with all evaluation samples meticulously annotated manually, assisted by a novel semi-automatic captioner on ultra-high-resolution RS images. On top of the XLRS-Bench, 16 sub-tasks are defined to evaluate MLLMs’ 10 kinds of perceptual capabilities and 6 kinds of reasoning capabilities, with a primary emphasis on advanced cognitive processes that facilitate real-world decision-making and the capture of spatiotemporal changes. The results of both general and RS-focused MLLMs on XLRS-Bench indicate that further efforts are needed for real-world RS applications. We have open-sourced XLRS-Bench to support further research in developing more powerful MLLMs for remote sensing. Fengxiang Wang 0004, Hongzhen Wang, Zonghao Guo, Di Wang 0023, Yulin Wang 0002, Mingshuo Chen, Long Lan, Wenjing Yang 0002, Jing Zhang 0037, Zhiyuan Liu 0001, Maosong Sun 0001 |
CVPR | 8 |
| 2025 | Text-guided Multimodal Fusion for the Multimodal Emotion and Intent Joint UnderstandingabstractEmotion and Intent Joint Understanding in Multi-modal Conversation is a challenging task in the field of affective computing, aiming to decode the semantic information manifested in the multimodal conversational while simultaneously inferring the emotions and intents of the utterance. To address this challenge, we propose the Text-guided Multimodal Emotion-Intent Joint Recognition method. By leveraging the text modality to guide the fusion process, it effectively reduces the noise introduced by other modalities. To strengthen the text modality’s guiding role, we use large language models (LLMs) for multi-turn targeted data augmentation and oversampling strategies to address data imbalance. Our approach achieved first place in Track 1 (English) of the ICASSP 2025 MEIJU Challenge, demonstrating its effectiveness in practical applications. Yu Zhang 0133, Bin Chen 0006, Hongfei Ye, Zijian Gao, Tianjiao Wan, Long Lan, Kele Xu |
ICASSP | 6 |
| 2025 | Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling
Fengxiang Wang 0004, Hongzhen Wang, Di Wang 0023, Zonghao Guo, Zhenyu Zhong, Long Lan, Wenjing Yang 0002, Jing Zhang 0037 |
ICCV | 6 |
| 2025 | Effective and Efficient Time-Varying Counterfactual Prediction with State-Space ModelsabstractTime-varying counterfactual prediction (TCP) from observational data supports the answer of when and how to assign multiple sequential treatments, yielding importance in various applications. Despite the progress achieved by recent advances, e.g., LSTM or Transformer based causal approaches, their capability of capturing interactions in long sequences remains to be improved in both prediction performance and running efficiency. In parallel with the development of TCP, the success of the state-space models (SSMs) has achieved remarkable progress toward long-sequence modeling with saved running time. Consequently, studying how Mamba simultaneously benefits the effectiveness and efficiency of TCP becomes a compelling research direction. In this paper, we propose to exploit advantages of the SSMs to tackle the TCP task, by introducing a counterfactual Mamba model with Covariate-based Decorrelation towards Selective Parameters (Mamba-CDSP). Motivated by the over-balancing problem in TCP of the direct covariate balancing methods, we propose to de-correlate between the current treatment and the representation of historical covariates, treatments, and outcomes, which can mitigate the confounding bias while preserve more covariate information. In addition, we show that the overall de-correlation in TCP is equivalent to regularizing the selective parameters of Mamba over each time step, which leads our approach to be effective and lightweight. We conducted extensive experiments on both synthetic and real-world datasets, demonstrating that Mamba-CDSP not only outperforms baselines by a large margin, but also exhibits prominent running efficiency. Haotian Wang 0001, Haoxuan Li 0001, Hao Zou 0001, Haoang Chi, Long Lan, Wanrong Huang, Wenjing Yang 0002 |
ICLR | 5 |
| 2025 | Wave-wise Discriminative Tracking by Phase-Amplitude Separation, Augmentation and MixtureabstractDistinguishing key features in complex visual tasks is challenging. A novel approach treats image patches (tokens) as waves. By using both phase and amplitude, it captures richer semantics and specific invariances compared to pixel-based methods, and allows for feature fusion across regions for a holistic image representation. Based on this, we propose the Wave-wise Discriminative Transformer Tracker (WDT). During tracking, WDT represents features via phase-amplitude separation, enhancement, and mixture. First, we designed a Mutual Exclusive Phase-Amplitude Extractor (MEPAE) to separate phase and amplitude features with distinct semantics, representing spatial target info and background brightness respectively. Then, Wave-wise Feature Augmentation is carried out with two submodules: Phase-Amplitude Feature Augmentation and Mixture. The augmentation module disrupts the separated features in the same batch, and the mixture module recombines them to generate positive and negative waves. The original features are aggregated into the original wave. Positive waves have the same phase but different amplitudes, and negative waves have different phase components. Finally, self-supervised and tracking-supervised losses guide the global and local representation learning for original, positive, and negative waves, enhancing wave-level discrimination. Experiments on five benchmarks prove the effectiveness of our method. Huibin Tan, Mingyu Cao, Xihuai He, Hao Li 0025, Long Lan, Mengzhu Wang |
IJCAI | 7 |
| 2025 | Breaking the Gradient Barrier: Unveiling Large Language Models for Strategic ClassificationabstractStrategic classification (SC) explores how individuals or entities modify their features strategically to achieve favorable classification outcomes. However, existing SC methods, which are largely based on linear models or shallow neural networks, face significant limitations in terms of scalability and capacity when applied to real-world datasets with significantly increasing scale, especially in financial services and the internet sector.
In this paper, we investigate how to leverage large language models to design a more scalable and efficient SC framework, especially in the case of growing individuals engaged with decision-making processes. Specifically, we introduce GLIM, a gradient-free SC method grounded in in-context learning.
During the feed-forward process of self-attention, GLIM implicitly simulates the typical bi-level optimization process of SC, including both the feature manipulation and decision rule optimization.
Without fine-tuning the LLMs, our proposed GLIM enjoys the advantage of cost-effective adaptation in dynamic strategic environments. Theoretically, we prove GLIM can support pre-trained LLMs to adapt to a broad range of strategic manipulations. We validate our approach through experiments with a collection of pre-trained LLMs on real-world and synthetic datasets in financial and internet domains, demonstrating that our GLIM exhibits both robustness and efficiency, and offering an effective solution for large-scale SC tasks. Xinpeng Lv, Yunxin Mao, Haoxuan Li 0001, Ke Liang 0006, Jinxuan Yang, Wanrong Huang, Haoang Chi, Long Lan, Yuanlong Chen, Wenjing Yang 0002, Haotian Wang 0001 |
NeurIPS | 9 |
| 2025 | GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K ResolutionabstractUltra-high-resolution (UHR) remote sensing (RS) imagery offers valuable data for Earth observation but pose challenges for existing multimodal foundation models due to two key bottlenecks: (1) limited availability of UHR training data, and (2) token explosion caused by the large image size. To address data scarcity, we introduce **SuperRS-VQA** (avg. 8,376$\times$8,376) and **HighRS-VQA** (avg. 2,000$\times$1,912), the highest-resolution vision-language datasets in RS to date, covering 22 real-world dialogue tasks. To mitigate token explosion, our pilot studies reveal significant redundancy in RS images: crucial information is concentrated in a small subset of object-centric tokens, while pruning background tokens (e.g., ocean or forest) can even improve performance. Motivated by these findings, we propose two strategies: *Background Token Pruning* and *Anchored Token Selection*, to reduce the memory footprint while preserving key semantics. Integrating these techniques, we introduce **GeoLLaVA-8K**, the first RS-focused multimodal large language model capable of handling inputs up to 8K$\times$8K resolution, built on the LLaVA framework. Trained on SuperRS-VQA and HighRS-VQA, GeoLLaVA-8K sets a new state-of-the-art on the XLRS-Bench. Datasets and code were released at https://github.com/MiliLab/GeoLLaVA-8K. Fengxiang Wang 0004, Mingshuo Chen, Di Wang 0023, Haotian Wang 0001, Zonghao Guo, Zefan Wang, Boqi Shan, Long Lan, Yulin Wang 0002, Hongzhen Wang, Wenjing Yang 0002, Bo Du 0001, Jing Zhang 0037 |
NeurIPS | 9 |
| 2025 | RoMA: Scaling up Mamba-based Foundation Models for Remote SensingabstractRecent advances in self-supervised learning for Vision Transformers (ViTs) have fueled breakthroughs in remote sensing (RS) foundation models. However, the quadratic complexity of self-attention poses a significant barrier to scalability, particularly for large models and high-resolution images. While the linear-complexity Mamba architecture offers a promising alternative, existing RS applications of Mamba remain limited to supervised tasks on small, domain-specific datasets. To address these challenges, we propose RoMA, a framework that enables scalable self-supervised pretraining of Mamba-based RS foundation models using large-scale, diverse, unlabeled data. RoMA enhances scalability for high-resolution images through a tailored auto-regressive learning strategy, incorporating two key innovations: 1) a rotation-aware pretraining mechanism combining adaptive cropping with angular embeddings to handle sparsely distributed objects with arbitrary orientations, and 2) multi-scale token prediction objectives that address the extreme variations in object scales inherent to RS imagery. Systematic empirical studies validate that Mamba adheres to RS data and parameter scaling laws, with performance scaling reliably as model and data size increase. Furthermore, experiments across scene classification, object detection, and semantic segmentation tasks demonstrate that RoMA-pretrained Mamba models consistently outperform ViT-based counterparts in both accuracy and computational efficiency. The source code and pretrained models have be released at https://github.com/MiliLab/RoMA. Fengxiang Wang 0004, Yulin Wang 0002, Mingshuo Chen, Haotian Wang 0001, Hongzhen Wang, Haiyan Zhao 0001, Yangang Sun, Di Wang 0023, Long Lan, Wenjing Yang 0002, Jing Zhang 0037 |
NeurIPS | 10 |
| 2025 | Uncertainty Quantification for Black-Box LLMs via Star Graphs Connectivity: Exploring Alternatives for Semantic Density
Zhaoye Li, Huibin Tan, Long Lan, Yize Sui |
ECML/PKDD (4) | 4 |
| 2025 | NT-FAN: A simple yet effective noise-tolerant few-shot adaptation network
Wenjing Yang 0002, Haoang Chi, Yibing Zhan, Xiaoguang Ren, Dapeng Tao, Long Lan |
Artif. Intell. | 7 |
| 2025 | A visual state space Model-Based Cross-Domain adaptive detection method for imbalanced medical image distribution
Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Hudan Pan, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001 |
Appl. Intell. | 2 |
| 2025 | Dragon Boat Optimization: A Meta-Heuristic for Intelligent SystemsabstractABSTRACT Dragon boat racing, a popular aquatic folklore team sport, is traditionally held during the Dragon Boat Festival. Inspired by this event, we propose a novel human‐based meta‐heuristic algorithm called dragon boat optimization (DBO) in this paper. It models the unique behaviours of each crew member on the dragon boat during the race by introducing social psychology mechanisms (social loafing, social incentive). Throughout this process, the focus is on the interaction and collaboration among the crew members, as well as their decision‐making in various situations. During each iteration, DBO implements different state updating strategies. By accurately modelling the crew's behaviour and employing adaptive state update strategies, DBO consistently achieves high optimization performance, as validated by comprehensive testing on 29 benchmark functions and 2 structural design problems. Experimental results indicate that DBO outperforms 7 and 16 state‐of‐the‐art meta‐heuristic algorithms across these test functions and problems, respectively. Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001 |
Expert Syst. J. Knowl. Eng. | 2 |
| 2025 | Self-supervised re-identification for online joint multi-object trackingabstractRecently, the bottleneck of multi-object tracking is shifting from detection performance to association performance. However, research on association algorithms requires a large number of identity labels, which are more expensive than detection labels. To circumvent the need for identity labels, we propose a Self-supervised Re-identification module for online joint Multi-Object Tracking (SR-MOT). Specifically, we design an appearance discriminator to judge identities based solely on detection hypotheses and then associate the same identity with the final trajectory. To train the discriminator without using identity labels, we construct negative pairs by the detections that appear in the same video frame, as they definitely belong to different identities. Positive pairs are naturally constructed through several useful data augmentation strategies at the box level. In addition, our proposed method balances conflicting detection and re-ID tasks by using different output features and dynamically adjusts detection and re-ID loss weights based on the information content of the loss distribution to promote balance between the two tasks from the feature level and optimization methods. In our evaluation on the MOT Challenge benchmark, we show that our SR-MOT performs comparably to supervised methods and is significantly superior to other unsupervised methods. Our proposed method provides a practical solution for multi-object tracking without the need for identity labels, making it more accessible for real-world applications. Shuman Li, Longqi Yang 0002, Huibin Tan, Binglin Wang, Wanrong Huang, Hengzhu Liu, Wenjing Yang 0002, Long Lan |
Knowl. Inf. Syst. | 8 |
| 2025 | From Concrete to Abstract: Multi-View Clustering on Relational KnowledgeabstractMulti-view clustering (MVC) is a fast-growing research direction. However, most existing MVC works focus on concrete objects (e.g., cats, desks) but ignore abstract objects (e.g., knowledge, thoughts), which are also important parts of our daily lives and more correlated to cognition. Relational knowledge, as a typical abstract concept, describes the relationship between entities. For example, "Cats like eating fishes," as relational knowledge, reveals the relationship "eating" between "cats" and "fishes." To fill this gap, we first point out that MVC on relational knowledge is considered an important scenario. Then, we construct 8 new datasets to lay research grounds for them. Moreover, a simple yet effective relational knowledge MVC paradigm (RK-MVC) is proposed by compensating the omitted sample-global correlations from the structural knowledge information. Concretely, the basic consensus features are first learned via adopted MVC backbones, and sample-global correlations are generated in both coarse-grained and fine-grained manners. In particular, the sample-global correlation learning module can be easily extended to various MVC backbones. Finally, both basic consensus features and sample-global correlation features are weighted fused as the target consensus feature. We adopt 9 typical MVC backbones in this paper for comparison from 7 aspects, demonstrating the promising capacity of our RK-MVC. Ke Liang 0006, Lingyuan Meng, Hao Li 0025, Jun Wang 0118, Long Lan, Miaomiao Li 0001, Xinwang Liu 0002, Huaimin Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Sample Adaptive Localized Simple Multiple Kernel K-Means and its Application in Parcellation of Human Cerebral CortexabstractSimple multiple kernel k-means (SMKKM) introduces a new minimization-maximization learning paradigm for multi-view clustering and makes remarkable achievements in some applications. As one of its variants, localized SMKKM (LSMKKM) is recently proposed to capture the variation among samples, focusing on reliable pairwise samples, which should keep together and cut off unreliable, farther pairwise ones. Though demonstrating effectiveness, we observe that LSMKKM indiscriminately utilizes the variation of each sample, resulting in unsatisfying clustering performance. To overcome this limitation, we propose a sample adaptive localized SMKKM (SAL-SMKKM) algorithm where the weight of the local alignment for each sample can be adaptively adjusted, resulting in a more challenging tri-level minimization-minimization-maximization. To deal with it, we reformulate it into a minimization problem of an optimal function characterized by minimization-maximization dynamics, prove its differentiability, and develop a reduced gradient descent method to optimize it. We then theoretically analyze the clustering performance of the proposed SAL-SMKKM by deriving its generalization error bound. In addition, we empirically evaluate the clustering performance of the proposed SAL-SMKKM on several benchmark datasets. Experiment results clearly indicate that proposed algorithms consistently outperform state-of-the-art ones. Finally, we apply the proposed SAL-SMKKM to the multi-modal parcellation of the human cerebral cortex, which is essential and helpful to understanding brain organization and function. As seen, SAL-SMKKM achieves accurate parcellation in an automatic and objective manner without any manual intervention, which once again demonstrates its validity and effectiveness in practical applications. Xinwang Liu 0002, Yi Zhang 0104, Li Liu 0002, Chang Tang, Long Lan, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Unknown-Aware Bilateral Dependency Optimization for Defending Against Model Inversion AttacksabstractBy abusing access to a well-trained classifier, model inversion (MI) attacks pose a significant threat as they can recover the original training data, leading to privacy leakage. Previous studies mitigated MI attacks by imposing regularization to reduce the dependency between input features and outputs during classifier training, a strategy known as unilateral dependency optimization. However, this strategy contradicts the objective of minimizing the supervised classification loss, which inherently seeks to maximize the dependency between input features and outputs. Consequently, there is a trade-off between improving the model's robustness against MI attacks and maintaining its classification performance. To address this issue, we propose the bilateral dependency optimization strategy (BiDO), a dual-objective approach that minimizes the dependency between input features and latent representations, while simultaneously maximizing the dependency between latent representations and labels. BiDO is remarkable for its privacy-preserving capabilities. However, models trained with BiDO exhibit diminished capabilities in out-of-distribution (OOD) detection compared to models trained with standard classification supervision. Given the open-world nature of deep learning systems, this limitation could lead to significant security risks, as encountering OOD inputs-whose label spaces do not overlap with the in-distribution (ID) data used during training-is inevitable. To address this, we leverage readily available auxiliary OOD data to enhance the OOD detection performance of models trained with BiDO. This leads to the introduction of an upgraded framework, unknown-aware BiDO (BiDO+), which mitigates both privacy and security concerns. As a highlight, with comparable model utility, BiDO-HSIC+ reduces the FPR95 by 55.02% and enhances the AUCROC by 9.52% compared to BiDO-HSIC, while also providing superior MI robustness. Xiong Peng, Feng Liu 0003, Nannan Wang 0001, Long Lan, Tongliang Liu, Yiu-Ming Cheung, Bo Han 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | WildVideo: Benchmarking LMMs for Understanding Video-Language InteractionabstractWe introduce WildVideo, an open-world benchmark dataset designed to address how to assess hallucination of Large Multi-modal Models (LMMs) for understanding video-language interaction in the wild. Our WildVideo comprehensively tests the perceptual, cognitive, and contextual comprehension hallucination of LMMs through both single-turn and multi-turn open-ended question-answering (QA) tasks on videos captured from two human perspectives (i.e. first-person view and third-person view). We define 9 distinct tasks that challenge LMMs across multi-level perceptual tasks (e.g., static and dynamic perception), multi-aspect cognitive tasks (e.g., commonsense, world knowledge), and multi-faceted contextual comprehension tasks (e.g., contextual ellipsis, cross-turn retrieval). The benchmark consists of 1,318 meticulously curated videos, supplemented with 13,704 single-turn QA pairs and 1,585 multi-turn dialogues (up to 5 turns). We evaluated 14 commonly-used LMMs on WildVideo, revealing significant hallucination issues of current LMMs, highlighting substantial gaps in their current capabilities. Songyuan Yang, Weijiang Yu, Wenjing Yang 0002, Xinwang Liu 0002, Huibin Tan, Long Lan, Nong Xiao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Graph Convolutional Mixture-of-Experts Learner Network for Long-Tailed Domain GeneralizationabstractThe goal of single domain generalization is to use data from a single domain (source domain) to train a model, which is then deployed over several unknown domains for testing (target domains). This study introduces a practical approach diverging from traditional DG, which typically relies on multiple source domains. We focus on Single Long-Tailed Domain Generalization, which refers to a scenario in the context of long-tail distribution, where although minority classes may have fewer samples in a single domain, these minority classes could become more prevalent and dominant in other domains. We introduce the Graph Convolutional Mixture-of-Experts Learners Network for Long-Tailed Domain Generalization (GCML) as a solution to this problem. Our approach presents two novel tactics. Initially, we utilize an expert learning technique that is skill-diverse. In order to properly manage the unknown target domain, this entails training multiple specialists inside a single long-tailed source domain and combining their knowledge. Then, we use a graph convolutional network to facilitate domain generalization, leveraging joint data structure modeling to learn more domain-invariant feature. Experiments conducted on four established benchmarks reveal that our GCML algorithm outperforms contemporary domain generalization techniques, demonstrating its efficacy in this complex task. Mengzhu Wang, Houcheng Su, Shanshan Wang 0008, Li Shen 0008, Long Lan, Liang Yang 0002, Xiaochun Cao |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Efficient Prompt Tuning of Large Vision-Language Model for Fine-Grained Ship ClassificationabstractRemote-sensing fine-grained ship classification (RS-FGSC) poses a significant challenge due to the high similarity between classes and the limited availability of labeled data, limiting the effectiveness of traditional supervised classification methods. Recent advancements in large pretrained vision-language models (VLMs) have demonstrated impressive capabilities in few-shot or zero-shot learning, particularly in understanding image content. This study delves into harnessing the potential of VLMs to enhance classification accuracy for unseen ship categories, which holds considerable significance in scenarios with restricted data due to cost or privacy constraints. Directly fine-tuning VLMs for RS-FGSC often encounters the challenge of overfitting the seen classes, resulting in suboptimal generalization to unseen classes, which highlights the difficulty in differentiating complex backgrounds and capturing distinct ship features. To address these issues, we introduce a novel prompt tuning technique that employs a hierarchical, multigranularity prompt design. Our approach integrates remote sensing ship priors through bias terms, learned from a small trainable network. This strategy enhances the model’s generalization capabilities while improving its ability to discern intricate backgrounds and learn discriminative ship features. Furthermore, we contribute to the field by introducing a comprehensive dataset, FGSCM-52, significantly expanding existing datasets with more extensive data and detailed annotations for less common ship classes. Extensive experimental evaluations demonstrate the superiority of our proposed method over current state-of-the-art techniques. The source code will be made publicly available. Long Lan, Fengxiang Wang 0004, Xiangtao Zheng, Zengmao Wang, Xinwang Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Towards Efficient Partially Relevant Video Retrieval With Active Moment DiscoveringabstractPartially relevant video retrieval (PRVR) is a practical yet challenging task in text-to-video retrieval, where videos are untrimmed and contain much background content. The pursuit here is of both effective and efficient solutions to capture the partial correspondence between text queries and untrimmed videos. Existing PRVR methods, which typically focus on modeling multi-scale clip representations, however, suffer from content independence and information redundancy, impairing retrieval performance. To overcome these limitations, we propose a simple yet effective approach with active moment discovering (AMDNet). We are committed to discovering video moments that are semantically consistent with their queries. By using learnable span anchors to capture distinct moments and applying masked multi-moment attention to emphasize salient moments while suppressing redundant backgrounds, we achieve more compact and informative video representations. To further enhance moment modeling, we introduce a moment diversity loss to encourage different moments of distinct regions and a moment relevance loss to promote semantically query-relevant moments, which cooperate with a partially relevant retrieval loss for end-to-end optimization. Extensive experiments on two large-scale video datasets (i.e., TVR and ActivityNet Captions) demonstrate the superiority and efficiency of our AMDNet. In particular, AMDNet is about 15.5 times smaller (#parameters) while 6.0 points higher (SumR) than the up-to-date method GMMFormer on TVR. Peipei Song, Long Lan, Weidong Chen 0013, Dan Guo 0001, Xun Yang 0001, Meng Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | STFormer: Spatial-Temporal-Aware Transformer for Video Instance SegmentationabstractVideo instance segmentation (VIS) is a challenging task, requiring handling object classification, segmentation, and tracking in videos. Existing Transformer-based VIS approaches have shown remarkable success, combining encoded features and instance queries as decoder inputs. However, their decoder inputs are low-resolution due to computational cost, resulting in a loss of fine-grained information, sensitivity to background interference, and poor handling of small objects. Moreover, the queries are randomly initialized without location information, hindering convergence efficiency and accurate object instance localization. To address these issues, we propose a novel VIS approach, STFormer, with a spatial-temporal feature aggregation (STFA) module and spatial-temporal-aware Transformer (STT). Specifically, STFA obtains robust high-resolution masked features efficiently for the decoder, while STT's location-guided instance query (LGIQ) improves initial instance queries. STFormer preserves more fine-grained information, improves convergence efficiency, and localizes object instance features accurately. Extensive experiments on YouTube-VIS 2019, YouTube-VIS 2021, and OVIS datasets show that STFormer outperforms mainstream VIS methods. Wei Wang 0335, Mengzhu Wang, Huibin Tan, Long Lan, Zhigang Luo, Xinwang Liu 0002, Kenli Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Smooth-Guided Implicit Data Augmentation for Domain GeneralizationabstractThe training process of a domain generalization (DG) model involves utilizing one or more interrelated source domains to attain optimal performance on an unseen target domain. Existing DG methods often use auxiliary networks or require high computational costs to improve the model's generalization ability by incorporating a diverse set of source domains. In contrast, this work proposes a method called Smooth-Guided Implicit Data Augmentation (SGIDA) that operates in the feature space to capture the diversity of source domains. To amplify the model's generalization capacity, a distance metric learning (DML) loss function is incorporated. Additionally, rather than depending on deep features, the suggested approach employs logits produced from cross entropy (CE) losses with infinite augmentations. A theoretical analysis shows that logits are effective in estimating distances defined on original features, and the proposed approach is thoroughly analyzed to provide a better understanding of why logits are beneficial for DG. Moreover, to increase the diversity of the source domain, a sampling-based method called smooth is introduced to obtain semantic directions from interclass relations. The effectiveness of the proposed approach is demonstrated through extensive experiments on widely used DG, object detection, and remote sensing datasets, where it achieves significant improvements over existing state-of-the-art methods across various backbone networks. Mengzhu Wang, Junze Liu, Ge Luo 0003, Shanshan Wang 0008, Wei Wang 0335, Long Lan, Ye Wang 0023, Feiping Nie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Scaling Few-Shot Learning for the Open WorldabstractFew-shot learning (FSL) aims to enable learning models with the ability to automatically adapt to novel (unseen) domains in open-world scenarios. Nonetheless, there exists a significant disparity between the vast number of new concepts encountered in the open world and the restricted available scale of existing FSL works, which primarily focus on a limited number of novel classes. Such a gap hinders the practical applicability of FSL in realistic scenarios. To bridge this gap, we propose a new problem named Few-Shot Learning with Many Novel Classes (FSL-MNC) by substantially enlarging the number of novel classes, exceeding the count in the traditional FSL setup by over 500-fold. This new problem exhibits two major challenges, including the increased computation overhead during meta-training and the degraded classification performance by the large number of classes during meta-testing. To overcome these challenges, we propose a Simple Hierarchy Pipeline (SHA-Pipeline). Due to the inefficiency of traditional protocols of EML, we re-design a lightweight training strategy to reduce the overhead brought by much more novel classes. To capture discriminative semantics across numerous novel classes, we effectively reconstruct and leverage the class hierarchy information during meta-testing. Experiments show that the proposed SHA-Pipeline significantly outperforms not only the ProtoNet baseline but also the state-of-the-art alternatives across different numbers of novel classes. Wenjing Yang 0002, Haotian Wang 0001, Haoang Chi, Long Lan, Ji Wang 0001 |
AAAI | 5 |
| 2024 | Learning to Learn Better Visual PromptsabstractPrompt tuning provides a low-cost way of adapting vision-language models (VLMs) for various downstream vision tasks without requiring updating the huge pre-trained parameters. Dispensing with the conventional manual crafting of prompts, the recent prompt tuning method of Context Optimization (CoOp) introduces adaptable vectors as text prompts. Nevertheless, several previous works point out that the CoOp-based approaches are easy to overfit to the base classes and hard to generalize to novel classes. In this paper, we reckon that the prompt tuning works well only in the base classes because of the limited capacity of the adaptable vectors. The scale of the pre-trained model is hundreds times the scale of the adaptable vector, thus the learned vector has a very limited ability to absorb the knowledge of novel classes. To minimize this excessive overfitting of textual knowledge on the base class, we view prompt tuning as learning to learn (LoL) and learn the prompt in the way of meta-learning, the training manner of dividing the base classes into many different subclasses could fully exert the limited capacity of prompt tuning and thus transfer it power to recognize the novel classes. To be specific, we initially perform fine-tuning on the base class based on the CoOp method for pre-trained CLIP. Subsequently, predicated on the fine-tuned CLIP model, we carry out further fine-tuning in an N-way K-shot manner from the perspective of meta-learning on the base classes. We finally apply the learned textual vector and VLM for unseen classes.Extensive experiments on benchmark datasets validate the efficacy of our meta-learning-informed prompt tuning, affirming its role as a robust optimization strategy for VLMs. Fengxiang Wang 0004, Wanrong Huang, Shaowu Yang, Long Lan |
AAAI | 5 |
| 2024 | Diversifying Cross-Domain Few-Shot Learning via Multimodal Image EditingabstractStanding out as one of the most widely used tools in Cross-Domain Few-Shot Learning (CDFSL), data augmentation forms the bedrock of numerous recent advancements. However, the current augmentations in CDFSL are limited in their ability to modify high-level semantic attributes, resulting in a lack of diversity along key semantic dimensions. One of the most promising tools to edit images with key semantic attributes, e.g. backgrounds, is image-to-image generation via large multimodal models (LMMs). Given the promising image editing results of recent LMMs, we delve into leveraging LMMs to augment data diversity for CDFSL. We propose a novel method named, Multimodal Few-shot Image Editing (MFIE), which uses LMMs to automatically translate class-specific images into class-agnostic natural language descriptions for various key semantic attributes in target domains and editing origin images based on class-agnostic natural language descriptions. To filter out corrupted data that disturbs the class-specific information, we apply semantic filtering using image-language similarity. Experiments on Meta-Datset show that MFIE surpasses SOTA CDFSL algorithms. Wenjing Yang 0002, Long Lan, Mingyang Geng, Haotian Wang 0001, Haoang Chi, Xueqiong Li, Ji Wang 0001 |
ICASSP | 3 |
| 2024 | Continuous Review and Timely Correction: Enhancing the Resistance to Noisy Labels via Self-Not-True DistillationabstractDeep neural networks possess substantial learning capacities and robust expressive power, making them prone to overfitting mislabeled data. Fortunately, the memorization effect shows that the networks tend to memorize the clean data first, and then gradually memorize the mislabeled data. Correspondingly, early stopping is proposed and has proven to be effective in mitigating overfitting. However, the networks can still overfit some mislabeled data in the early training stage, resulting in forgotten knowledge of clean data. In addition, early stopping lacks correction of errors caused by mislabeled data. In this paper, we propose that the network should continuously review the knowledge it learned earlier to enhance clean data memorization while timely correcting the incorrect knowledge learned from the mislabeled data. To implement these two ideas, we first introduce self-distillation into training, which employs a teacher network from the previous stage to guide the current network, enhancing clean data memorization. Based on this, we further propose the not-true distillation. Before distilling knowledge from the teacher network, we mask the true class (i.e. label class) in the logits, focusing only on not-true classes to correct the accumulated incorrect knowledge. Extensive experiments on simulated and realworld benchmarks adequately validate the superior performance of our method. Xinghao Wu, Yuhua Tang, Long Lan |
ICASSP | 5 |
| 2024 | Contrastive Transformer Cross-Modal Hashing for Video-Text Retrieval
Xiaobo Shen 0001, Qianxin Huang, Long Lan, Yuhui Zheng |
IJCAI | 3 |
| 2024 | Enhancing Unsupervised Visible-Infrared Person Re-Identification with Bidirectional-Consistency Gradual MatchingabstractUnsupervised visible-infrared person re-identification (USL-VI-ReID) is of great research and practical significance yet remains challenging due to significant modality discrepancy and lack of annotations. Many existing approaches utilize variants of bipartite graph global matching algorithms to address this issue, aiming to establish cross-modality correspondences. However, these methods may encounter mismatches due to significant modality gaps and limited model representation. To mitigate this, we propose a simple yet effective framework for USL-VI-ReID, which gradually establishes associations between different modalities. To measure the confidence whether samples from different modalities belong to the same identity, we introduce a bidirectional-consistency criterion, which not only considers direct relationships between samples from different modalities but also incorporates potential hard negative samples from the same modality. Additionally, we propose a cross-modality correlation preserving module to further enhance the semantic representation of the model by maintaining consistency in correlations across modalities. Extensive experiments conducted on the public SYSU-MM01 and RegDB datasets demonstrate the superiority of our method over existing USL-VI-ReID approaches across various settings, despite the simplicity of our method. Xiao Teng, Kele Xu, Long Lan |
ACM Multimedia | 4 |
| 2024 | Tracing Training Progress: Dynamic Influence Based Selection for Active LearningabstractActive learning (AL) aims to select highly informative data points from an unlabeled dataset for annotation, mitigating the need for extensive human labeling effort. However, classical AL methods heavily rely on human expertise to design the sampling strategy, inducing limited scalability and generalizability. Many efforts have sought to address this limitation by directly connecting sample selection with model performance improvement, typically through influence function. Nevertheless, these approaches often ignore the dynamic nature of model behavior during training optimization, despite empirical evidence highlights the importance of dynamic influence to track the sample contribution. This oversight can lead to suboptimal selection, hindering the generalizability of model. In this study, we explore the dynamic influence based data selection strategy by tracing the impact of unlabeled instances on model performance throughout the training process. Our theoretical analyses suggest that selecting samples with higher projected gradients along the accumulated optimization direction at each checkpoint leads to improved performance. Furthermore, to capture a wider range of training dynamics without incurring excessive computational or memory costs, we introduce an additional dynamic loss term designed to encapsulate more generalized training progress information. These insights are integrated into a universal and task-agnostic AL framework termed Dynamic Influence Scoring for Active Learning (DISAL). Comprehensive experiments across various tasks have demonstrated that DISAL significantly surpasses existing state-of-the-art AL methods, demonstrating its ability to facilitate more efficient and effective learning in different domains. Tianjiao Wan, Kele Xu, Long Lan, Zijian Gao, Bo Ding 0001, Huaimin Wang 0001 |
ACM Multimedia | 3 |
| 2024 | MambaTrack: A Simple Baseline for Multiple Object Tracking with State Space ModelabstractTracking by detection has been the prevailing paradigm in the field of Multi-object Tracking (MOT). These methods typically rely on the Kalman Filter to estimate the future locations of objects, assuming linear object motion. However, they fall short when tracking objects exhibiting nonlinear and diverse motion in scenarios like dancing and sports. In addition, there has been limited focus on utilizing learning-based motion predictors in MOT. To address these challenges, we resort to exploring data-driven motion prediction methods. Inspired by the great expectation of state space models (SSMs), such as Mamba, in long-term sequence modeling with near-linear complexity, we introduce a Mamba-based motion model named Mamba moTion Predictor (MTP). MTP is designed to model the complex motion patterns of objects like dancers and athletes. Specifically, MTP takes the spatial-temporal location dynamics of objects as input, captures the motion pattern using a bi-Mamba encoding layer, and predicts the next motion. In real-world scenarios, objects may be missed due to occlusion or motion blur, leading to premature termination of their trajectories. To tackle this challenge, we further expand the application of MTP. We employ it in an autoregressive way to compensate for missing observations by utilizing its own predictions as inputs, thereby contributing to more consistent trajectories. Our proposed tracker, MambaTrack, demonstrates advanced performance on benchmarks such as Dancetrack and SportsMOT, which are characterized by complex motion and severe occlusion. Changcheng Xiao, Qiong Cao, Zhigang Luo, Long Lan |
ACM Multimedia | 4 |
| 2024 | Self-distillation Enhanced Vertical Wavelet Spatial Attention for Person Re-identification
Huibin Tan, Long Lan, Xiao Teng |
MMM (2) | 3 |
| 2024 | Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?abstractCausal reasoning capability is critical in advancing large language models (LLMs) towards artificial general intelligence (AGI). While versatile LLMs appear to have demonstrated capabilities in understanding contextual causality and providing responses that obey the laws of causality, it remains unclear whether they perform genuine causal reasoning akin to humans. However, current evidence indicates the contrary. Specifically, LLMs are only capable of performing shallow (level-1) causal reasoning, primarily attributed to the causal knowledge embedded in their parameters, but they lack the capacity for genuine human-like (level-2) causal reasoning. To support this hypothesis, methodologically, we delve into the autoregression mechanism of transformer-based LLMs, revealing that it is not inherently causal. Empirically, we introduce a new causal Q&A benchmark named CausalProbe 2024, whose corpus is fresh and nearly unseen for the studied LLMs. Empirical results show a significant performance drop on CausalProbe 2024 compared to earlier benchmarks, indicating that LLMs primarily engage in level-1 causal reasoning.To bridge the gap towards level-2 causal reasoning, we draw inspiration from the fact that human reasoning is usually facilitated by general knowledge and intended goals. Inspired by this, we propose G$^2$-Reasoner, a LLM causal reasoning method that incorporates general knowledge and goal-oriented prompts into LLMs' causal reasoning processes. Experiments demonstrate that G$^2$-Reasoner significantly enhances LLMs' causal reasoning capability, particularly in fresh and fictitious contexts. This work sheds light on a new path for LLMs to advance towards genuine causal reasoning, going beyond level-1 and making strides towards level-2. Haoang Chi, Wenjing Yang 0002, Feng Liu 0003, Long Lan, Xiaoguang Ren, Tongliang Liu, Bo Han 0003 |
NeurIPS | 5 |
| 2024 | Instance-Level Scaling and Dynamic Margin-Alignment Knowledge Distillation
Xiao Teng, Zheng Qin 0002, Long Lan, Jing Zhang 0037 |
PRCV (11) | 6 |
| 2024 | DBTN: An adaptive neural network for multiple-disease detection via imbalanced medical images distribution
Xiang Li 0089, Long Lan, Chang-Yong Sun, Shaowu Yang, Shuihua Wang, Wenjing Yang 0002, Heng Liu 0001, Yudong Zhang 0001 |
Appl. Intell. | 2 |
| 2024 | Discriminative object tracking by domain contrast
Huayue Cai, Xiang Zhang 0008, Long Lan, Changcheng Xiao, Chuanfu Xu, Jie Liu 0002, Zhigang Luo |
Comput. Vis. Image Underst. | 3 |
| 2024 | EAFP-Med: An efficient adaptive feature processing module based on prompts for medical image detectionabstractThe rapid proliferation of medical imaging technologies presents a significant challenge for cross-domain adaptive image detection, as lesion representations can vary dramatically across technologies. To address this issue, we draw inspiration from large language models to propose EAFP-Med, an efficient adaptive feature processing module based on prompts for medical image detection. EAFP-Med incorporates a prompt-driven dynamic parameter update mechanism, empowering it to extract cross-domain multi-scale lesion features from medical images of diverse modalities adaptively. This exceptional flexibility liberates it from the constraints of any particular imaging technique, fostering great adaptability. Furthermore, EAFP-Med can also serve as a feature preprocessing module connected to any model front-end to enhance the lesion features in input images. Moreover, we propose a novel adaptive disease detection model named EAFP-Med ST, which utilizes the Swin Transformer V2 – Tiny (SwinV2-T) as its backbone and connects it to EAFP-Med. We have compared our method to nine state-of-the-art methods. Experimental results show that the overall accuracy of EAFP Med ST on chest X-ray, brain magnetic resonance imaging, and skin image datasets is 98.47%, 97.60%, and 99.06%, respectively, superior to all the compared state-of-the-art methods. Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Does Confusion Really Hurt Novel Class Discovery?
Haoang Chi, Wenjing Yang 0002, Feng Liu 0003, Long Lan, Bo Han 0003 |
Int. J. Comput. Vis. | 4 |
| 2024 | TIG-CL: Teacher-Guided Individual- and Group-Aware Contrastive Learning for Unsupervised Person Reidentification in Internet of ThingsabstractUnsupervised person reidentification (Re-ID) has attracted widespread due to its potential in Internet of Things applications, such as intelligent visual surveillance, it refers to retrieving the same individual across different camera views without using labeled data. To tackle the problem, a prevalent technique adopted by existing methods involves generating pseudo labels through clustering algorithms. However, this approach can result in merging individuals with different identities into the same group (i.e., cluster) during the training process. As a result, the resulting group centers may obscure the inherent characteristics of individual identities, thereby hindering the model from learning discriminative representations. To address the issue, we present a teacher-guided individual- and group-aware contrastive learning framework. Specifically, we propose a departure from the traditional approach of relying solely on contrastive learning between individual features and their corresponding group centers. Instead, we also exploit the relationship among individuals to construct contrast pairs and facilitate the learning of more discriminative features. This strategy enables the model to learn more about the individual characteristics that distinguish different persons, thus enhancing its ability to reidentify individuals accurately. Moreover, our method introduces a novel hybrid distillation module that enables simultaneous probability distillation at the group level and relationship distillation at the individual level. Guided by the teacher model, this module leads to improved feature representations of the student model. Extensive experimental results verify the effectiveness of our approach on four popular Re-ID data sets. The code will be made publicly available. Xiao Teng, Xueqiong Li, Xinwang Liu 0002, Long Lan |
IEEE Internet Things J. | 5 |
| 2024 | IoUformer: Pseudo-IoU prediction with transformer for visual tracking
Huayue Cai, Long Lan, Jing Zhang 0037, Xiang Zhang 0008, Yibing Zhan, Zhigang Luo |
Neural Networks | 2 |
| 2024 | MotionTrack: Learning motion predictor for multiple object tracking
Changcheng Xiao, Qiong Cao, Long Lan, Xiang Zhang 0008, Zhigang Luo, Dacheng Tao |
Neural Networks | 4 |
| 2024 | Analysis of Video Quality Datasets via Design of Minimalistic Video Quality ModelsabstractBlind video quality assessment (BVQA) plays an indispensable role in monitoring and improving the end-users' viewing experience in various real-world video-enabled media applications. As an experimental field, the improvements of BVQA models have been measured primarily on a few human-rated VQA datasets. Thus, it is crucial to gain a better understanding of existing VQA datasets in order to properly evaluate the current progress in BVQA. Towards this goal, we conduct a first-of-its-kind computational analysis of VQA datasets via designing minimalistic BVQA models. By minimalistic, we restrict our family of BVQA models to build only upon basic blocks: a video preprocessor (for aggressive spatiotemporal downsampling), a spatial quality analyzer, an optional temporal quality analyzer, and a quality regressor, all with the simplest possible instantiations. By comparing the quality prediction performance of different model variants on eight VQA datasets with realistic distortions, we find that nearly all datasets suffer from the easy dataset problem of varying severity, some of which even admit blind image quality assessment (BIQA) solutions. We additionally justify our claims by comparing our model generalization capabilities on these VQA datasets, and by ablating a dizzying set of BVQA design choices related to the basic building blocks. Our results cast doubt on the current progress in BVQA, and meanwhile shed light on good practices of constructing next-generation VQA datasets and models. Wei Sun 0029, Wen Wen 0007, Xiongkuo Min, Long Lan, Guangtao Zhai, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Tackling Noisy Labels With Network Parameter Additive DecompositionabstractGiven data with noisy labels, over-parameterized deep networks suffer overfitting mislabeled data, resulting in poor generalization. The memorization effect of deep networks shows that although the networks have the ability to memorize all noisy data, they would first memorize clean training data, and then gradually memorize mislabeled training data. A simple and effective method that exploits the memorization effect to combat noisy labels is early stopping. However, early stopping cannot distinguish the memorization of clean data and mislabeled data, resulting in the network still inevitably overfitting mislabeled data in the early training stage. In this paper, to decouple the memorization of clean data and mislabeled data, and further reduce the side effect of mislabeled data, we perform additive decomposition on network parameters. Namely, all parameters are additively decomposed into two groups, i.e., parameters w are decomposed as w=σ+γ. Afterward, the parameters σ are considered to memorize clean data, while the parameters γ are considered to memorize mislabeled data. Benefiting from the memorization effect, the updates of the parameters σ are encouraged to fully memorize clean data in early training, and then discouraged with the increase of training epochs to reduce interference of mislabeled data. The updates of the parameters γ are the opposite. In testing, only the parameters σ are employed to enhance generalization. Extensive experiments on both simulated and real-world benchmarks confirm the superior performance of our method. Xiaobo Xia, Long Lan, Xinghao Wu, Jun Yu 0001, Wenjing Yang 0002, Bo Han 0003, Tongliang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | SiamATTRPN: Enhance Visual Tracking With Channel and Spatial AttentionabstractVisual tracking is an important research topic in the field of computer vision. The current Siamese tracker based on the region proposal network (SiamRPN) has achieved promising tracking results in terms of efficiency and performance. However, through our empirical study, we have observed that deep features learned by SiamRPN are of substandard quality, as the salient regions within the deep features fail to correspond accurately with meaningful objects. To address this limitation, we propose an approach to enhance the quality of the learned deep features through the incorporation of an attention mechanism. Attention mechanisms have been shown to be effective in distinguishing similar objects, as they suppress background objects while highlighting target information that is most relevant. As a result, a new tracking method with channel and spatial attention termed SiamATTRPN is explored. To verify the effectiveness of SiamATTRPN, experiments on benchmark datasets demonstrate that our proposed tracker outperforms the baseline tracker significantly. Huayue Cai, Xiang Zhang 0008, Long Lan, Wenxin Shen, Junyang Chen 0001, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | DeIoU: Toward Distinguishable Box Prediction in Densely Packed Object DetectionabstractThe Intersection over Union (IoU) has been widely employed in various stages of object detection owing to its ability to quantify the similarity between boxes objectively. However, in densely packed scenes full of crowded and small-sized objects, adjacent positive boxes often exhibit high levels of overlap. This overlap interference compromises the consistency between quality evaluation and confidence, leading to ambiguous box prediction within the previous IoU-based models. To address this issue, we design a novel learning paradigm tailored for Dense scenes based on IoU, called DeIoU. This approach effectively suppresses unnecessary overlap between predicted boxes and thereby enhances representation learning for non-salient objects. Specifically, it consists of a dense box regression loss${\mathcal {L}}_{DeIoU}$and a one-to-many (O2M) label matching strategy guided by DeIoU. These components focus on calibrating the position and shape prediction quality during the model training, learning distinguishable object features by penalizing overlap interference between neighboring boxes. Extensive experiments on four object detection datasets including SKU-110K, CrowdHuman, MS COCO 2017, and DIOR, demonstrate that our DeIoU-based learning strategy outperforms other state-of-the-art methods. Notably, the proposed method delivers a substantial improvement (average$1.3~{AP}$and$1.8~MR^{-2}$) across popular detectors on SKU-110K and CrowdHuman while exhibiting distinct competitiveness on small objects within natural scenes. Linfei Wang, Yibing Zhan, Long Lan, Dapeng Tao, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Perceptual Quality Assessment of Virtual Reality Videos in the WildabstractInvestigating how people perceive virtual reality (VR) videos in the wild (i.e., those captured by everyday users) is a crucial and challenging task in VR-related applications due to complexauthenticdistortionslocalizedinspaceandtime.Existingpanoramic video databases only consider synthetic distortions, assume fixed viewing conditions, and are limited in size. To overcome these shortcomings, we construct the VR Video Quality in the Wild (VRVQW) database, containing 502 user-generated videos with diverse content and distortion characteristics. Based on VRVQW, we conduct a formal psychophysical experiment to record the scanpaths and perceived quality scores from 139 participants under two different viewing conditions. We provide a thorough statistical analysis of the recordeddata, observing significantimpact of viewing conditions on both human scanpaths and perceived quality. Moreover, we develop an objective quality assessment model for VR videos based on pseudocylindrical representation and convolution. Results on the proposed VRVQW show that our method is superior to existing video quality assessment models.We have made the database and code available at https://github.com/ limuhit/VR-Video-Quality-in-the-Wild. Wen Wen 0007, Mu Li 0005, Yiru Yao, Xiangjie Sui, Yabin Zhang 0002, Long Lan, Yuming Fang 0001, Kede Ma |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Joint Spatial-Spectral Optimization for the High-Magnification Fusion of Hyperspectral and Multispectral ImagesabstractThe fusion of hyperspectral and multispectral images is an important strategy for enhancing the spatial resolution of hyperspectral images. With the rapid advancement of multispectral imaging technology, the disparity in spatial resolution between multispectral and hyperspectral images is increasing. In certain scenarios, termed high-magnification, this difference can exceed$32\times $. Previous methods do not perform well under high-magnification fusion, and naturally, a challenge arises in achieving effective high-magnification super-resolution fusion. In light of the above analysis, this article introduces a novel algorithm for high-magnification super-resolution fusion of hyperspectral and multispectral images based on the joint optimization of spatial and spectral information. Specifically, our algorithm consists of three stages: 1) a fast preliminary fusion stage based on the Moore-Penrose inverse and singular value correlation priors for the rapid acquisition of preliminary solutions; 2) a joint spatial-spectral optimization stage where a coupled optimization framework is constructed to achieve integrated optimization of spatial and spectral information; and 3) an error backpropagation optimization stage where an effective error optimization term is introduced to further refine the fusion performance. We conducted extensive experiments on widely employed publicly available simulated datasets and real datasets. The experimental results unequivocally indicate that our proposed methodology consistently exhibits superior fusion performance compared with state-of-the-art methods, even under the condition of${\geq }60\times $super-resolution. Yibing Zhan, Zhengbin Pang, Tong Zhou 0008, Xueqiong Li, Long Lan, Yuanxi Peng |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Out-of-Distribution Generalization With Causal Feature SeparationabstractDriven by empirical risk minimization, machine learning algorithm tends to exploit subtle statistical correlations existing in the training environment for prediction, while the spurious correlations are unstable across environments, leading to poor generalization performance. Accordingly, the problem of the Out-of-distribution (OOD) generalization aims to exploit an invariant/stable relationship between features and outcomes that generalizes well on all possible environments. To address the spurious correlation induced by the selection bias, in this article, we propose a novel Clique-based Causal Feature Separation (CCFS) algorithm by explicitly incorporating the causal structure to identify causal features of outcome for OOD generalization. Specifically, the proposed CCFS algorithm identifies the largest clique in the learned causal skeleton. Theoretically, we guarantee that either the largest clique or the rest of the causal skeleton is exactly the set of all causal features of the outcome. Finally, we separate the causal features from the non-causal ones with a sample-reweighting decorrelator for OOD prediction. Extensive experiments validate the effectiveness of the proposed CCFS method on both causal feature identification and OOD generalization tasks. Haotian Wang 0001, Kun Kuang 0001, Long Lan, Zige Wang, Wanrong Huang, Fei Wu 0001, Wenjing Yang 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Highly Efficient Active Learning With Tracklet-Aware Co-Cooperative Annotators for Person Re-IdentificationabstractSupervised person re-identification (ReID) has attracted widespread attentions in the computer vision community due to its great potential in real-world applications. However, the demand of human annotation heavily limits the application as it is costly to annotate identical pedestrians appearing from different cameras. Thus, how to reduce the annotation cost while preserving the performance remains challenging and has been studied extensively. In this article, we propose a tracklet-aware co-cooperative annotators' framework to reduce the demand of human annotation. Specifically, we partition the training samples into different clusters and associate adjacent images in each cluster to produce the robust tracklet which decreases the annotation requirements significantly. Besides, to further reduce the cost, we introduce a powerful teacher model in our framework to implement the active learning strategy and select the most informative tracklets for human annotator, the teacher model itself, in our setting, also acts as an annotator to label the relatively certain tracklets. Thus, our final model could be well-trained with both confident pseudo-labels and human-given annotations. Extensive experiments on three popular person ReID datasets demonstrate that our approach could achieve competitive performance compared with state-of-the-art methods in both active learning and unsupervised learning (USL) settings. Xiao Teng, Long Lan, Xueqiong Li, Yuhua Tang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Enhanced Dcf Tracker Regularized by Reliable Sample ConstructionabstractDiscriminative correlation filter (DCF) is a highly efficient tracking technique using the circulant shifted samples of search images to update the template, so the reliability of input samples determines template quality. In this paper, we rethink the reliability problem of input samples in advance during template updating and propose an enhanced DCF tracking method regularized by a novel sparse representation based reliable sample construction term, called enhanced sparse correlation filter (ESCF). Specifically, the reconstructed reliable samples are the sparse representation of circulant shifted samples of unfiltered input samples, in which the target will approach the center to preserve target visual cues into the template when using the cosine window. Besides, we jointly perform template learning and reliable sample construction into a unified learning paradigm to benefit from each other, which further can be carried out in the frequency domain without incurring excessive time cost by skillful decomposition. Experiments on several popular visual tracking datasets verify the efficacy of ESCF and show that ESCF performs favorably against several well-established representative counterparts. Mingyu Cao, Mengzhu Wang, Long Lan, Wenjing Yang 0002, Huibin Tan |
ICASSP | 4 |
| 2023 | Domain Specified Optimization for Deployment AuthorizationabstractThis paper explores Deployment Authorization (DPA) as a means of restricting the generalization capabilities of vision models on certain domains to protect intellectual property. Nevertheless, the current advancements in DPA are predominantly confined to fully supervised settings. Such settings require the accessibility of annotated images from any unauthorized domain, rendering the DPA approaches impractical for real-world applications due to its exorbitant costs.To address this issue, we propose Source-Only Deployment Authorization (SDPA), which assumes that only authorized domains are accessible during training phases, and the model’s performance on unauthorized domains must be suppressed in inference stages. Drawing inspiration from distributional robust statistics, we present a lightweight method called Domain-Specified Optimization (DSO) for SDPA that degrades the model’s generalization over a divergence ball. DSO comes with theoretical guarantees on the convergence property and its authorization performance. As a complementary of SDPA, we also propose Target-Combined Deployment Authorization (TPDA), where unauthorized domains are partially accessible, and simplify the DSO method to a perturbation operation on the pseudo predictions, referred to as Target-Dependent Domain-Specified Optimization (TDSO). We demonstrate the effectiveness of our proposed DSO and TDSO methods through extensive experiments on six image benchmarks, achieving dominant performance on both SDPA and TDPA settings. Haotian Wang 0001, Haoang Chi, Wenjing Yang 0002, Mingyang Geng, Long Lan, Jing Zhang 0037, Dacheng Tao |
ICCV | 6 |
| 2023 | MagicFusion: Boosting Text-to-Image Generation Performance by Fusing Diffusion ModelsabstractThe advent of open-source AI communities has produced a cornucopia of powerful text-guided diffusion models that are trained on various datasets. While few explorations have been conducted on ensembling such models to combine their strengths. In this work, we propose a simple yet effective method called Saliency-aware Noise Blending (SNB) that can empower the fused text-guided diffusion models to achieve more controllable generation. Specifically, we experimentally find that the responses of classifier-free guidance are highly related to the saliency of generated images. Thus we propose to trust different models in their areas of expertise by blending the predicted noises of two diffusion models in a saliency-aware manner. SNB is training-free and can be completed within a DDIM sampling process. Additionally, it can automatically align the semantics of two noise spaces without requiring additional annotations such as masks. Extensive experiments show the impressive effectiveness of SNB in various applications. The project page is available at https://magicfusion.github.io/. Heliang Zheng, Long Lan, Wenjing Yang 0002 |
ICCV | 4 |
| 2023 | CoCo: A Coupled Contrastive Framework for Unsupervised Domain Adaptive Graph ClassificationabstractAlthough graph neural networks (GNNs) have achieved impressive achievements in graph classification, they often need abundant task-specific labels, which could be extensively costly to acquire. A credible solution is to explore additional labeled graphs to enhance unsupervised learning on the target domain. However, how to apply GNNs to domain adaptation remains unsolved owing to the insufficient exploration of graph topology and the significant domain discrepancy. In this paper, we propose Coupled Contrastive Graph Representation Learning (CoCo), which extracts the topological information from coupled learning branches and reduces the domain discrepancy with coupled contrastive learning. CoCo contains a graph convolutional network branch and a hierarchical graph kernel network branch, which explore graph topology in implicit and explicit manners. Besides, we incorporate coupled branches into a holistic multi-view contrastive learning framework, which not only incorporates graph representations learned from complementary views for enhanced understanding, but also encourages the similarity between cross-domain example pairs with the same semantics for domain alignment. Extensive experiments on popular datasets show that our CoCo outperforms these competing baselines in different settings generally. Li Shen 0008, Mengzhu Wang, Long Lan, Zeyu Ma 0001, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
ICML | 4 |
| 2023 | Semantics-Enriched Cross-Modal Alignment for Complex-Query Video Moment RetrievalabstractVideo moment retrieval (VMR) aims to search for a video segment that matches the search intent in a query sentence, which has received increasing attention in recent years, due to its practical values in various fields. Existing efforts devoted to this interesting yet challenging task typically encode the query sentence and video segments into unstructured global representations for cross-modal interaction and fusion, which may fail to accurately capture the search intent in complex queries with multi-granularity semantics. Xiang Zhang 0008, Xun Yang 0001, Yibing Zhan, Long Lan, Jianfeng Dong, Hongzhou Wu |
ACM Multimedia | 5 |
| 2023 | Null-text Guidance in Diffusion Models is Secretly a Cartoon-style CreatorabstractClassifier-free guidance is an effective sampling technique in diffusion models that has been widely adopted. The main idea is to extrapolate the model in the direction of text guidance and away from null-text guidance. In this paper, we demonstrate that null-text guidance in diffusion models is secretly a cartoon-style creator, i.e., the generated images can be efficiently transformed into cartoons by simply perturbing the null-text guidance. Specifically, we proposed two disturbance methods, i.e., Rollback disturbance (Back-D) and Image disturbance (Image-D), to construct misalignment between the noisy images used for predicting null-text guidance and text guidance (subsequently referred to as null-text noisy image and text noisy imageb respectively) in the sampling process. Back-D achieves cartoonization by altering the noisb level of the null-text noisy image via replacing xt with xl + Δ t. Image-D, alternatively, produces high-fidelity, diverse cartoons by defining xt as a clean input image, which further improves the incorporation of finer image details. Through comprehensive experiments, we delved into the principle of noise disturbing for null-text and uncovered that the efficacy of disturbance depends on the correlation between the null-text noisy image and the source image. Moreover, the proposed methods, which can generate cartoon images and cartoonize specific ones, are training-free and easily integrated as a plug-and-play component in any classifier-free guided diffusion model. The project page is available at https://nulltextforcartoon.github.io/. Heliang Zheng, Long Lan, Wanrong Huang, Wenjing Yang 0002 |
ACM Multimedia | 4 |
| 2023 | SODA: Robust Training of Test-Time Data AdaptorsabstractAdapting models deployed to test distributions can mitigate the performance degradation caused by distribution shifts. However, privacy concerns may render model parameters inaccessible. One promising approach involves utilizing zeroth-order optimization (ZOO) to train a data adaptor to adapt the test data to fit the deployed models. Nevertheless, the data adaptor trained with ZOO typically brings restricted improvements due to the potential corruption of data features caused by the data adaptor. To address this issue, we revisit ZOO in the context of test-time data adaptation. We find that the issue directly stems from the unreliable estimation of the gradients used to optimize the data adaptor, which is inherently due to the unreliable nature of the pseudo-labels assigned to the test data. Based on this observation, we propose pseudo-label-robust data adaptation (SODA) to improve the performance of data adaptation. Specifically, SODA leverages high-confidence predicted labels as reliable labels to optimize the data adaptor with ZOO for label prediction. For data with low-confidence predictions, SODA encourages the adaptor to preserve data information to mitigate data corruption. Empirical results indicate that SODA can significantly enhance the performance of deployed models in the presence of distribution shifts without requiring access to model parameters. Zige Wang, Yonggang Zhang 0003, Zhen Fang 0001, Long Lan, Wenjing Yang 0002, Bo Han 0003 |
NeurIPS | 4 |
| 2023 | Self-aware circular response-guided attention for robust siamese tracking
Huibin Tan, Mengzhu Wang, Tianyi Liang 0001, Yuhua Tang, Long Lan, Wenjing Yang 0002 |
Appl. Intell. | 6 |
| 2023 | Domain-specific feature recalibration and alignment for multi-source unsupervised domain adaptationabstractAbstract Traditional unsupervised domain adaptation (UDA) usually assumes that the source domain has labels and the target domain has no labels. In a real environment, labelled source domain data usually comes from multiple different distributions. To handle this problem, multi‐source unsupervised domain adaptation (MUDA) is proposed. Multi‐source unsupervised domain adaptation aims to adapt the model trained on multi‐labelled source domains to the unlabelled target domain. In this paper, a novel MUDA method by domain‐specific feature recalibration and alignment (FRA) is proposed. Specifically, to achieve feature recalibration, the authors leverage channel attention to pick out significant channels and spatial attention to focus on important features in different channels. Such integration of channel and spatial attention can lead to effective domain‐specific feature recalibration that may be of great importance to MUDA. In addition, to achieve better MUDA, the authors propose domain‐specific feature alignment which consists of Maximum Mean Discrepancy and JS‐divergence loss. Maximum Mean Discrepancy can reduce the difference between the source domain and target domain. Meanwhile, JS‐divergence loss may ensure the prediction consistency of different classifiers in the source domains. Four experiments have proved that FRA can achieve significantly better results in popular benchmarks for MUDA. Mengzhu Wang, Dingyao Chen, Fangzhou Tan, Tianyi Liang 0001, Long Lan, Xiang Zhang 0008, Zhigang Luo |
IET Comput. Vis. | 5 |
| 2023 | Online intervention siamese tracking
Huayue Cai, Long Lan, Jing Zhang 0037, Xiang Zhang 0008, Changcheng Xiao, Zhigang Luo |
Inf. Sci. | 2 |
| 2023 | SiamDF: Tracking training data-free siamese tracker
Huayue Cai, Long Lan, Jing Zhang 0037, Xiang Zhang 0008, Zhigang Luo |
Neural Networks | 2 |
| 2023 | Meta attention for Off-Policy Actor-Critic
Jiateng Huang, Wanrong Huang, Long Lan |
Neural Networks | 3 |
| 2023 | Class-specific and self-learning local manifold structure for domain adaptation
Wei Wang 0335, Mengzhu Wang, Long Lan, Quannan Zu, Xiang Zhang 0008, Cong Wang 0018 |
Pattern Recognit. | 4 |
| 2023 | Reducing bi-level feature redundancy for unsupervised domain adaptation
Mengzhu Wang, Shanshan Wang 0008, Wei Wang 0335, Li Shen 0008, Xiang Zhang 0008, Long Lan, Zhigang Luo |
Pattern Recognit. | 6 |
| 2023 | Discriminative Geometric-Structure-Based Deep Hashing for Large-Scale Image RetrievalabstractDeep hashing reaps the benefits of deep learning and hashing technology, and has become the mainstream of large-scale image retrieval. It generally encodes image into hash code with feature similarity preserving, that is, geometric-structure preservation, and achieves promising retrieval results. In this article, we find that existing geometric-structure preservation manner inadequately ensures feature discrimination, while improving feature discrimination of hash code essentially determines hash learning retrieval performance. This fact principally spurs us to propose a discriminative geometric-structure-based deep hashing method (DGDH), which investigates three novel loss terms based on class centers to induce the so-called discriminative geometrical structure. In detail, the margin-aware center loss assembles samples in the same class to the corresponding class centers for intraclass compactness, then a linear classifier based on class center serves to boost interclass separability, and the radius loss further puts different class centers on a hypersphere to tentatively reduce quantization errors. An efficient alternate optimization algorithm with guaranteed desirable convergence is proposed to optimize DGDH. We theoretically analyze the robustness and generalization of the proposed method. The experiments on five popular benchmark datasets demonstrate superior image retrieval performance of the proposed DGDH over several state of the arts. Guohua Dong, Xiang Zhang 0008, Xiaobo Shen 0001, Long Lan, Zhigang Luo, Xiaomin Ying |
IEEE Trans. Cybern. | 4 |
| 2023 | Learning to Purification for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification is a challenging and promising task in computer vision. Nowadays unsupervised person re-identification methods have achieved great progress by training with pseudo labels. However, how to purify feature and label noise is less explicitly studied in the unsupervised manner. To purify the feature, we take into account two types of additional features from different local views to enrich the feature representation. The proposed multi-view features are carefully integrated into our cluster contrast learning to leverage more discriminative cues that the global feature easily ignored and biased. To purify the label noise, we propose to take advantage of the knowledge of teacher model in an offline scheme. Specifically, we first train a teacher model from noisy pseudo labels, and then use the teacher model to guide the learning of our student model. In our setting, the student model could converge fast with the supervision of the teacher model thus reduce the interference of noisy labels as the teacher model greatly suffered. After carefully handling the noise and bias in the feature learning, our purification modules are proven to be very effective for unsupervised person re-identification. Extensive experiments on two popular person re-identification datasets demonstrate the superiority of our method. Especially, our approach achieves a state-of-the-art accuracy 85.8% @mAP and 94.5% @Rank-1 on the challenging Market-1501 benchmark with ResNet-50 under the fully unsupervised setting. Code has been available at: https://github.com/tengxiao14/Purification_ReID. Long Lan, Xiao Teng, Jing Zhang 0037, Xiang Zhang 0008, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2023 | Contrastive Transformer Hashing for Compact Video RepresentationabstractVideo hashing learns compact representation by mapping video into low-dimensional Hamming space and has achieved promising performance in large-scale video retrieval. It is challenging to effectively exploit temporal and spatial structure in an unsupervised setting. To fulfill this gap, this paper proposes Contrastive Transformer Hashing (CTH) for effective video retrieval. Specifically, CTH develops a bidirectional transformer autoencoder, based on which visual reconstruction loss is proposed. CTH is more powerful to capture bidirectional correlations among frames than conventional unidirectional models. In addition, CTH devises multi-modality contrastive loss to reveal intrinsic structure among videos. CTH constructs inter-modality and intra-modality triplet sets and proposes multi-modality contrastive loss to exploit inter-modality and intra-modality similarities simultaneously. We perform video retrieval tasks on four benchmark datasets, i.e., UCF101, HMDB51, SVW30, FCVID using the learned compact hash representation, and extensive empirical results demonstrate the proposed CTH outperforms several state-of-the-art video hashing methods. Xiaobo Shen 0001, Yun-Hao Yuan 0001, Xichen Yang, Long Lan, Yuhui Zheng |
IEEE Trans. Image Process. | 5 |
| 2023 | Local-to-Global Deep Clustering on Approximate Uniform ManifoldabstractDeep clustering usually treats the clustering assignments as supervisory signals to learn a more compact representation with deep neural networks, under the guidance of clustering-oriented losses. Nevertheless, we observe that, without reliable supervision, such losses for global clustering would destroy the locally geometric structure underlying data. In this paper, we propose a local-to-global deep clustering method based on approximate uniform manifold (LGC-AUM) to address this issue in a two-stage fashion. In the local stage, an intra-manifold preservation loss is proposed to preserve intra-manifold structures locally on basis of approximate uniform manifold, and an inter-manifold discrimination loss is for global inter-manifold structure. Thus, this stage serves to learn more discriminative structure-preserving features by reducing the correlations between different manifolds, which paves the way for the final clustering. Build off the learned features, the second stage explores a clustering loss based on approximate uniform manifold to establish stable network training for effective clustering with two auxiliary distributions. Experiments on five benchmark datasets verify the efficacy of our LGC-AUM as compared to several well-behaved clustering counterparts. Xiang Zhang 0008, Long Lan, Zhigang Luo |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Privacy-Preserving Action RecognitionabstractAs the amount of data shared on the network increases, these data pose a threat to our privacy. This paper focuses on the privacy-preserving issues of action recognition for humans. Generally, the face is considered the most identifiable visual cue for a human. However, removing face information is not enough for many privacy-preserving scenes. Thus, we replace the human body with his poses and explore the pose presentation in the action recognition task. In privacy scenes, many human actions could not access in advance. To recognize these unseen actions, we study the zero-shot action recognition in the strict condition of privacy preservation. Specifically, we propose to use unified actor score (UAS) to enhance the action recognition accuracy. The experimental results show that UAS outperforms most of the state-of-the-art methods in standard datasets without sacrificing privacy. Chengming Zou, Ducheng Yuan, Long Lan, Haoang Chi |
ICASSP | 3 |
| 2022 | Meta Discovery: Learning to Discover Novel Classes given Very Limited Data
Haoang Chi, Feng Liu 0003, Wenjing Yang 0002, Long Lan, Tongliang Liu, Bo Han 0003, Gang Niu 0001, Mingyuan Zhou, Masashi Sugiyama |
ICLR | 4 |
| 2022 | Toward to Real Low-Resolution Person Re-identification: A New Dataset and BaselineabstractPerson re-identification(re-id) aims at querying and identifying the same target pedestrian in multiple non-overlapping cameras. However, in real scenarios, many person images captured by surveillance cameras tend to have low resolution due to camera hardware, shooting distance, viewing angle, etc. Many existing re-id methods mainly address the crossresolution person re-id problem, i.e., the high-low resolution mismatch problem. In contrast, the problem of matching low-resolution (LR) gallery images and query images with each other has been less studied. In this paper, we address this problem by developing a new gun-ball camera-based person re-id dataset and designing a LR re-id baseline model for this dataset to tackle the LR person matching problem. Extensive experiments validate the effectiveness of the proposed LR baseline model. Dongting Sun, Long Lan, Zhigang Luo |
ICME | 3 |
| 2022 | Counterfactual Causal Adversarial Networks for Domain Adaptation
Yan Jia 0001, Xiang Zhang 0008, Long Lan, Zhigang Luo |
ICONIP (6) | 3 |
| 2022 | Logit Distillation via Student Diversity
Dingyao Chen, Long Lan, Mengzhu Wang, Xiang Zhang 0008, Tianyi Liang 0001, Zhigang Luo |
ICONIP (5) | 2 |
| 2022 | Self-Reinforcing Feedback Domain Adaptation Channel
Yan Jia 0001, Xiang Zhang 0008, Long Lan, Zhigang Luo |
ICONIP (1) | 3 |
| 2022 | Frustratingly Easy Knowledge Distillation via Attentive Similarity MatchingabstractKnowledge distillation is an effective approach to transferring knowledge from the large teacher network to its small proxy student one, thereby letting the proxy student work on those resource-limited mobile devices. Most previous arts manually select the paired intermediate layers of teacher and student networks to align their pertinent features by dimension reduction. This sort of approach may confront information loss and insufficient layer-wise alignment that limit knowledge transferability. In this paper, we propose a simple and effective knowledge distillation method named attentive similarity matching (ASM). ASM at first concatenates the teacher’s intermediate features and the student’s ones together to enhance similarity representation of all the student’s layers, without involving dimension reduction, then align all cross-layer advanced similarities in an attentively weighted manner for semantic calibration. Experiments of image classification on three popular datasets show the effectiveness of the proposed method as compared to its previous cousins. Dingyao Chen, Huibin Tan, Long Lan, Xiang Zhang 0008, Tianyi Liang 0001, Zhigang Luo |
ICPR | 3 |
| 2022 | Bilateral Dependency Optimization: Defending Against Model-inversion AttacksabstractThrough using only a well-trained classifier, model-inversion (MI) attacks can recover the data used for training the classifier, leading to the privacy leakage of the training data. To defend against MI attacks, previous work utilizes a unilateral dependency optimization strategy, i.e., minimizing the dependency between inputs (i.e., features) and outputs (i.e., labels) during training the classifier. However, such a minimization process conflicts with minimizing the supervised loss that aims to maximize the dependency between inputs and outputs, causing an explicit trade-off between model robustness against MI attacks and model utility on classification tasks. In this paper, we aim to minimize the dependency between the latent representations and the inputs while maximizing the dependency between latent representations and the outputs, named a bilateral dependency optimization (BiDO) strategy. In particular, we use the dependency constraints as a universally applicable regularizer in addition to commonly used losses for deep neural networks (e.g., cross-entropy), which can be instantiated with appropriate dependency criteria according to different tasks. To verify the efficacy of our strategy, we propose two implementations of BiDO, by using two different dependency measures: BiDO with constrained covariance (BiDO-COCO) and BiDO with Hilbert-Schmidt Independence Criterion (BiDO-HSIC). Experiments show that BiDO achieves the state-of-the-art defense performance for a variety of datasets, classifiers, and MI attacks while suffering a minor classification-accuracy drop compared to the well-trained classifier with no defense, which lights up a novel road to defend against MI attacks. Xiong Peng, Feng Liu 0003, Jingfeng Zhang, Long Lan, Junjie Ye 0002, Tongliang Liu, Bo Han 0003 |
KDD | 4 |
| 2022 | Joint Modality Synergy and Spatio-temporal Cue Purification for Moment LocalizationabstractCurrently, many approaches to the sentence query based moment location (SQML) task emphasize (inter-)modality interaction between video and language query via transformer-based cross-attention or contrastive learning. However, they could still face two issues: 1) modality interaction could be unexpectedly friendly to modality specific learning that merely learns modality specific patterns, and 2) modality interaction easily confuses spatio-temporal cues and ultimately makes time cues in the original video ambiguous. In this paper, we propose a modality synergy with spatio-temporal cue purification method (MS2P) for SQML to address the above two issues. Particularly, a conceptually simple modality synergy strategy is explored to keep features modality specific while absorbing the other modality complementary information with both carefully designed cross-attention unit and non-contrastive learning. As a result, modality specific semantics can be calibrated progressively in a safer way. To preserve time cues in original video, we further purify video representation into spatial and temporal parts to enhance localization resolution by the proposed two light-weight sentence-aware filtering operations. Experiments on Charades-STA, TACoS, and ActivityNet Caption datasets show our model outperforms the state-of-the-art approaches by a large margin. Long Lan, Huibin Tan, Xiang Zhang 0008, Xurui Ma, Zhigang Luo |
ICMR | 2 |
| 2022 | APT-36K: A Large-scale Benchmark for Animal Pose Estimation and TrackingabstractAnimal pose estimation and tracking (APT) is a fundamental task for detecting and tracking animal keypoints from a sequence of video frames. Previous animal-related datasets focus either on animal tracking or single-frame animal pose estimation, and never on both aspects. The lack of APT datasets hinders the development and evaluation of video-based animal pose estimation and tracking methods, limiting the applications in real world, e.g., understanding animal behavior in wildlife conservation. To fill this gap, we make the first step and propose APT-36K, i.e., the first large-scale benchmark for animal pose estimation and tracking. Specifically, APT-36K consists of 2,400 video clips collected and filtered from 30 animal species with 15 frames for each video, resulting in 36,000 frames in total. After manual annotation and careful double-check, high-quality keypoint and tracking annotations are provided for all the animal instances. Based on APT-36K, we benchmark several representative models on the following three tracks: (1) supervised animal pose estimation on a single frame under intra- and inter-domain transfer learning settings, (2) inter-species domain generalization test for unseen animals, and (3) animal pose estimation with animal tracking. Based on the experimental results, we gain some empirical insights and show that APT-36K provides a useful animal pose estimation and tracking benchmark, offering new challenges and opportunities for future research. The code and dataset will be made publicly available at https://github.com/pandorgan/APT-36K. Yuxiang Yang 0001, Yufei Xu, Jing Zhang 0037, Long Lan, Dacheng Tao |
NeurIPS | 5 |
| 2022 | Online Multiple-Pedestrian Tracking With Detection-Pair-Based Graph Convolutional NetworksabstractThe typical Internet of Things application, unattended driving systems, will need the ability to recognize relevant traffic participants and detect dangerous situations ahead of time. An important component of these systems is one that is able to distinguish pedestrians and track their motion to make intelligent driving decisions. This article develops a high-accuracy multiple pedestrian tracking algorithm which is vital for intelligent transportation. Here, we use the off-the-shelf detectors and explore the benefits of modeling pedestrian interactions, such as the interaction of two pedestrians simultaneously matched to two pedestrians in another frame, for robust detection association. Explicitly studying interactions is nontrivial. Previous works often manually selected interacting detections (or “tracklets”) to simplify the association process. In this article, we propose a novel association method based on deep graph convolutional affinity networks (DGCANs) and extend detection-level interactions to the association-level, which treats a potential association of a detection pair as a node in the graph, and explicitly modeling the interactions among potential associations. Specifically, with the novel node, two corresponding edges are readily designed to model the compatible and colliding interactions between related associations. Our proposed method, by redefining nodes and edges, enables us to blend sufficient interaction cues from appearance and motion and learns a robust affinity measure in an end-to-end fashion. Using the Hungarian algorithm as an online tracker, our method archives state-of-the-art performance on benchmark data sets 2-D MOT15, MOT16, and MOT17. Weijiang Feng, Long Lan, Michael Buro, Zhigang Luo |
IEEE Internet Things J. | 2 |
| 2022 | Heterogeneous Pseudo-Supervised Learning for Few-shot Person Re-Identification
Long Lan, Wenjing Yang 0002 |
Neural Networks | 2 |
| 2022 | Label Propagated Nonnegative Matrix Factorization for ClusteringabstractSemi-supervised learning (SSL) that utilizes plenty of unlabeled examples to boost the performance of learning from limited labeled examples is a powerful learning paradigm with widely real-world applications such as information retrieval and document clustering. Label propagation (LP) is a popular SSL method which propagates labels through the dataset along high density areas defined by unlabeled examples, but it is fragile to bridge examples. Semi-supervised K-Means uses labeled examples to initialize clustering centers to separate different examples, however, semi-supervised K-Means fails in the situation of imbalanced issues, that is, the example size of each class varies significantly. This paper proposes a novel label propagated nonnegative matrix factorization method (LPNMF) to handle clean labeled but biased data and its extension LPNMF-E to handle noisy labeled data based on the framework of NMF. LPNMF decomposes the whole dataset into the product of a basis matrix and a coefficient matrix. To propagate labels to unlabeled examples, LPNMF regards the class indicators of labeled examples as their coefficients and iteratively updates both basis matrix and coefficients of unlabeled examples. LPNMF absorbs the merits from both semi-supervised K-Means and label propagation to handle their respective shortages. Specifically, on the one hand, LPNMF learns representative clustering centers based on the distribution of the dataset, similar to semi-supervised K-means, and thus is robust to the bridge examples. On the other hand, LPNMF pushes labels according to the affinity between examples, similar to label propagation, and thus relieves the biased problem. Moreover, we introduce a LPNMF extension to handle the noisy label case. LPNMF-E relaxes the constraint of labeled examples. Since the label of each labeled example also obtains label information from the global distribution of the whole dataset and local manifold of its neighbors, LPNMF-E outputs reliable class indicators even if a portion of examples are incorrectly labeled. Theoretical analyses for the generalization ability of our proposed models are also provided. Experimental results on both clean and noisy labeled datasets confirm the effectiveness of LPNMF and LPNMF-E compared with both LP and the representative semi-supervised K-Means algorithms. Long Lan, Tongliang Liu, Xiang Zhang 0008, Chuanfu Xu, Zhigang Luo |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Deep Co-Image-Label Hashing for Multi-Label Image RetrievalabstractDeep supervised hashing has greatly improved retrieval performance with the powerful learning capability of deep neural network. In multi-label image retrieval, existing deep hashing simply indicates whether two images are similar by constructing a similarity matrix. However, it ignores the dependency among multiple labels that has been shown important in multi-label application. To fulfill this gap, this paper proposes Deep Co-Image-Label Hashing (DCILH) to discover label dependency. Specifically, DCILH regards image and label as two views, and maps the two views into a common deep Hamming space. DCILH proposes to learn prototype for each label, and preserve similarity among images, labels, and prototypes. To exploit label dependency, DCILH further employs the label-correlation aware loss on the predicted labels, such that predicted output on positive label is enforced to be larger than that on negative label. Extensive experiments on several multi-label benchmarks demonstrate the proposed DCILH outperforms state-of-the-art deep supervised hashing on large-scale multi-label image retrieval. Xiaobo Shen 0001, Guohua Dong, Yuhui Zheng, Long Lan, Ivor W. Tsang, Quan-Sen Sun |
IEEE Trans. Multim. | 4 |
| 2022 | Redundancy, Context, and Preference: An Empirical Study of Duplicate Pull Requests in OSS ProjectsabstractOSS projects are being developed by globally distributed contributors, who often collaborate through the pull-based model today. While this model lowers the barrier to entry for OSS developers by synthesizing, automating and optimizing the contribution process, coordination among an increasing number of contributors remains as a challenge due to the asynchronous and self-organized nature of distributed development. In particular, duplicate contributions, where multiple different contributors unintentionally submit duplicate pull requests to achieve the same goal, are an elusive problem that may waste effort in automated testing, code review and software maintenance. While the issue of duplicate pull requests has been highlighted, to what extent duplicate pull requests affect the development in OSS communities has not been well investigated. In this paper, we conduct a mixed-approach study to bridge this gap. Based on a comprehensive dataset constructed from 26 popular GitHub projects, we obtain the following findings: (a) Duplicate pull requests result in redundant human and computing resources, exerting a significant impact on the contribution and evaluation process. (b) Contributors’ inappropriate working patterns and the drawbacks of their collaborating environment might result in duplicate pull requests. (c) Compared to non-duplicate pull requests, duplicate pull requests have significantly different features, e.g., being submitted by inexperienced contributors, being fixing bugs, touching cold files, and solving tracked issues. (d) Integrators choosing between duplicate pull requests prefer to accept those with early submission time, accurate and high-quality implementation, broad coverage, test code, high maturity, deep discussion, and active response. Finally, actionable suggestions and implications are proposed for OSS practitioners. Yue Yu 0001, Minghui Zhou 0001, Tao Wang 0006, Gang Yin, Long Lan, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 6 |
| 2021 | Model Compression for a Plasticity Neural Network in a Maze Exploration Scenario
Baolun Yu, Wanrong Huang, Long Lan, Yuhua Tang |
ICONIP (5) | 3 |
| 2021 | InterBN: Channel Fusion for Adversarial Unsupervised Domain AdaptationabstractA classifier trained on one dataset rarely works on other datasets obtained under different conditions because of domain shifting. Such a problem is usually solved by domain adaptation methods. In this paper, we propose a novel unsupervised domain adaptation (UDA) method based on Interchangeable Batch Normalization (InterBN) to fuse different channels in deep neural networks for adversarial domain adaptation.Specifically, we first observe that the channels with small batch normalization scaling factor have less influence on the whole domain adaption, followed by a theoretical proof that the scaling factors for some channels will definitely come close to zero when imposing a sparsity regularization. Then, we replace the channels that have smaller scaling factors in the source domain with the mean of the channels which have larger scaling factors in the target domain or vice versa. Such a simple but effective channel fusion scheme can drastically increase the domain adaption ability.Extensive experimental results show that our InterBN significantly outperforms the current adversarial domain adaptation methods by a large margin on four visual benchmarks. In particular, InterBN achieves a remarkable improvement of 7.7% over the conditional adversarial adaptation networks (CDAN) on VisDA-2017 benchmark. Mengzhu Wang, Wei Wang 0335, Baopu Li, Xiang Zhang 0008, Long Lan, Huibin Tan, Tianyi Liang 0001, Wei Yu 0029, Zhigang Luo |
ACM Multimedia | 5 |
| 2021 | TOHAN: A One-step Approach towards Few-shot Hypothesis AdaptationabstractIn few-shot domain adaptation (FDA), classifiers for the target domain are trained with \emph{accessible} labeled data in the source domain (SD) and few labeled data in the target domain (TD). However, data usually contain private information in the current era, e.g., data distributed on personal phones. Thus, the private data will be leaked if we directly access data in SD to train a target-domain classifier (required by FDA methods). In this paper, to prevent privacy leakage in SD, we consider a very challenging problem setting, where the classifier for the TD has to be trained using few labeled target data and a well-trained SD classifier, named few-shot hypothesis adaptation (FHA). In FHA, we cannot access data in SD, as a result, the private information in SD will be protected well. To this end, we propose a target-oriented hypothesis adaptation network (TOHAN) to solve the FHA problem, where we generate highly-compatible unlabeled data (i.e., an intermediate domain) to help train a target-domain classifier. TOHAN maintains two deep networks simultaneously, in which one focuses on learning an intermediate domain and the other takes care of the intermediate-to-target distributional adaptation and the target-risk minimization. Experimental results show that TOHAN outperforms competitive baselines significantly. Haoang Chi, Feng Liu 0003, Wenjing Yang 0002, Long Lan, Tongliang Liu, Bo Han 0003, William Kwok-Wai Cheung, James T. Kwok |
NeurIPS | 4 |
| 2021 | ANF: Attention-Based Noise Filtering Strategy for Unsupervised Few-Shot Classification
Guangsen Ni, Wenjing Yang 0002, Long Lan |
PRICAI (3) | 6 |
| 2021 | Enhancing the association in multi-object tracking via neighbor graphabstractMost modern multi-object tracking (MOT) systems for videos follow the tracking-by-detection paradigm, where objects of interest are first located in each frame then associated correspondingly to form their intact trajectories. In this setting, the appearance features of objects usually provide the most important cues for data association, but it is very susceptible to occlusions, illumination variations, and inaccurate detections, thus easily resulting in incorrect trajectories. To address this issue, in this study we propose to make full use of the neighboring information. Our motivations derive from the observations that people tend to move in a group. As such, when an individual target's appearance is remarkably changed, the observer can still identify it with its neighbor context. To model the contextual information from neighbors, we first utilize the spatiotemporal relations among trajectories to efficiently select suitable neighbors for targets. Subsequently, we construct neighbor graph for each target and corresponding neighbors then employ the graph convolutional networks (GCNs) to model their relations and learn the graph features. To the best of our knowledge, it is the first time to explicitly leverage neighbor cues via GCN in MOT. Finally, standardized evaluations on the MOT16 and MOT17 data sets demonstrate that our approach can remarkably reduce the identity switches whilst achieve state-of-the-art overall performance. Tianyi Liang 0001, Long Lan, Xiang Zhang 0008, Xindong Peng, Zhigang Luo |
Int. J. Intell. Syst. | 2 |
| 2021 | Semantic-consistent cross-modal hashing for large-scale image retrieval
Xuesong Gu, Guohua Dong, Xiang Zhang 0008, Long Lan, Zhigang Luo |
Neurocomputing | 4 |
| 2021 | A generic MOT boosting framework by combining cues from SOT, tracklet and re-identification
Tianyi Liang 0001, Long Lan, Xiang Zhang 0008, Zhigang Luo |
Knowl. Inf. Syst. | 2 |
| 2021 | A robust quadruple adaptation network in few-shot scenarios
Haoang Chi, Shengang Li, Wenjing Yang 0002, Long Lan |
Knowl. Based Syst. | 4 |
| 2021 | Learning deep discriminative embeddings via joint rescaled features and log-probability centers
Huayue Cai, Xiang Zhang 0008, Long Lan, Guohua Dong, Chuanfu Xu, Xinwang Liu 0002, Zhigang Luo |
Pattern Recognit. | 3 |
| 2021 | Unsupervised Discriminative Deep Hashing With Locality and Globality PreservationabstractDeep hashing has greatly improved retrieval performance with the powerful learning capability of deep neural network. However, deep unsupervised hashing can hardly achieve impressive performance due to the lack of the semantic supervision. This letter proposes Unsupervised Discriminative Deep Hashing (UD2H) to fulfill this gap. UD2H is formulated to jointly perform hash code learning and clustering, and trained in an asymmetric manner to improve the efficiency. The cluster labels supervise the training of deep model to enable hash code discriminative. Based on the outputs of the deep model, UD2H adaptively constructs a similarity graph that considers the local and global structures. Experiments on three benchmark datasets show that the proposed UD$^2$H outperforms the state-of-the-art unsupervised deep hashing methods. Zhuyi Ni, Zexuan Ji, Long Lan, Yun-Hao Yuan 0001, Xiaobo Shen 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Near-Online Multi-Pedestrian Tracking via Combining Multiple Consistent Appearance CuesabstractAn important cue for multi-pedestrian tracking in video is the consistent appearance of an individual for quite a while. In this paper, we address multi-pedestrian tracking by learning a robust appearance model from the paradigm of tracking by detection. To separate detections of different pedestrians while assembling detections of the same pedestrian, we take advantage of the cue of consistent appearance and exploit three types of evidence from the recent, past and near-future. Existing online approaches only exploit the detection-to-detection and sequence-to-detection metrics, which focus on the recent and past appearance patterns respectively, while the future pedestrian appearance is simply ignored. This drawback is remedied in this paper by further considering the sequence-to-sequence metric, which resorts to near-future appearance presentation. Adaptive combination weights are learned to fuse these three different metrics. Moreover, we propose a novel Focal Triplet Loss to make the model focus more on hard examples than the easy ones. We demonstrate that this can significantly enhance the discriminating power of the model compared with treating every sample equally. Effectiveness and efficiency of the proposed method is verified by conducting comprehensive ablation studies and comparing with many competitive (offline/online/near-online) counterparts on the MOT16 and MOT17 Challenges. Weijiang Feng, Long Lan, Yong Luo 0002, Yue Yu 0001, Xiang Zhang 0008, Zhigang Luo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Nocal-Siam: Refining Visual Features and Response With Advanced Non-Local Blocks for Real-Time Siamese TrackingabstractSiamese trackers contain two core stages, i.e., learning the features of both target and search inputs at first and then calculating response maps via the cross-correlation operation, which can also be used for regression and classification to construct typical one-shot detection tracking framework. Although they have drawn continuous interest from the visual tracking community due to the proper trade-off between accuracy and speed, both stages are easily sensitive to the distracters in search branch, thereby inducing unreliable response positions. To fill this gap, we advance Siamese trackers with two novel non-local blocks named Nocal-Siam, which leverages the long-range dependency property of the non-local attention in a supervised fashion from two aspects. First, a target-aware non-local block (T-Nocal) is proposed for learning the target-guided feature weights, which serve to refine visual features of both target and search branches, and thus effectively suppress noisy distracters. This block reinforces the interplay between both target and search branches in the first stage. Second, we further develop a location-aware non-local block (L-Nocal) to associate multiple response maps, which prevents them inducing diverse candidate target positions in the future coming frame. Experiments on five popular benchmarks show that Nocal-Siam performs favorably against well-behaved counterparts both in quantity and quality. Huibin Tan, Xiang Zhang 0008, Long Lan, Wenju Zhang, Zhigang Luo |
IEEE Trans. Image Process. | 4 |
| 2020 | Robust Normalized Squares Maximization for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) attempts to transfer specific knowledge from one domain with labeled data to another domain without labels. Recently, maximum squares loss has been proposed to tackle UDA problem but it does not consider the prediction diversity which has proven beneficial to UDA. In this paper, we propose a novel normalized squares maximization (NSM) loss in which the maximum squares is normalized by the sum of squares of class sizes. The normalization term enforces the class sizes of predictions to be balanced to explicitly increase the diversity. Theoretical analysis shows that the optimal solution to NSM is one-hot vectors with balanced class sizes, i.e., NSM encourages both discriminate and diverse predictions. We further propose a robust variant of NSM, RNSM, by replacing the square loss with L2,1-norm to reduce the influence of outliers and noises. Experiments of cross-domain image classification on two benchmark datasets illustrate the effectiveness of both NSM and RNSM. RNSM achieves promising performance compared to state-of-the-art methods. The code is available at https://github.com/wj-zhang/NSM. Wenju Zhang, Xiang Zhang 0008, Qing Liao 0001, Wenjing Yang 0002, Long Lan, Zhigang Luo |
CIKM | 5 |
| 2020 | Towards Making Unsupervised Graph Hashing RobustabstractUnsupervised hashing without supervision easily deteriorates in the case of grossly corrupted data. Motivated by robust optimization, this paper proposes a dual-graph regularized robust hashing (DGRH) based on both manifold smoothness and robust estimators in a more intuitive manner. Orthogonal to existing robust hashing methods, DGRH directly removes the outliers of datasets with M-estimator to exert robustness. In specific, it intends to recover low-rank representation from corrupted data via l1loss while preserving neighborhood relationships among samples with dual-graph regularization. Although DGRH seems a simple extension of robust PCA on graphs with hashing trick, it is easy to implement yet effective. Theory analysis is provided to support our claim. Experiments of image retrieval on three popular benchmark datasets show the efficacy of DGRH as compared to several well-behaved representative counterparts. Xuesong Gu, Guohua Dong, Xiang Zhang 0008, Long Lan, Zhigang Luo |
ICME | 4 |
| 2020 | Pairwise Similarity Regularization for Adversarial Domain AdaptationabstractDomain adaptation aims at learning a predictive model that can generalize to a new target domain different from the source (training) domain. To mitigate the domain gap, adversarial training has been developed to learn domain invariant representations. State-of-the-art methods further make use of pseudo labels generated by the source domain classifier to match conditional feature distributions between the source and target domains. However, if the target domain is more complex than the source domain, the pseudo labels are unreliable to characterize the class-conditional structure of the target domain data, undermining prediction performance. To resolve this issue, we propose a Pairwise Similarity Regularization (PSR) approach that exploits cluster structures of the target domain data and minimizes the divergence between the pairwise similarity of clustering partition and that of pseudo predictions. Therefore, PSR guarantees that two target instances in the same cluster have the same class prediction and thus eliminate the negative effect of unreliable pseudo labels. Extensive experimental results show that our PSR method significantly boosts the current adversarial domain adaptation methods by a large margin on four visual benchmarks. In particular, PSR achieves a remarkable improvement of more than 5% over the state-of-the-art on several hard-to-transfer tasks. Haotian Wang 0001, Wenjing Yang 0002, Ji Wang 0001, Ruxin Wang 0002, Long Lan, Mingyang Geng |
ACM Multimedia | 5 |
| 2020 | Semi-online Multi-people Tracking by Re-identification
Long Lan, Xinchao Wang, Gang Hua 0001, Thomas S. Huang, Dacheng Tao |
Int. J. Comput. Vis. | 1 |
| 2020 | Learning sequence-to-sequence affinity metric for near-online multi-object tracking
Weijiang Feng, Long Lan, Xiang Zhang 0008, Zhigang Luo |
Knowl. Inf. Syst. | 2 |
| 2020 | Enhancing unsupervised domain adaptation by discriminative relevance regularization
Wenju Zhang, Xiang Zhang 0008, Long Lan, Zhigang Luo |
Knowl. Inf. Syst. | 3 |
| 2020 | Object-aware semantics of attention for image captioning
Long Lan, Xiang Zhang 0008, Guohua Dong, Zhigang Luo |
Multim. Tools Appl. | 2 |
| 2020 | GateCap: Gated spatial and semantic attention model for image captioning
Long Lan, Xiang Zhang 0008, Zhigang Luo |
Multim. Tools Appl. | 2 |
| 2020 | Maximum Mean and Covariance Discrepancy for Unsupervised Domain Adaptation
Wenju Zhang, Xiang Zhang 0008, Long Lan, Zhigang Luo |
Neural Process. Lett. | 3 |
| 2019 | Attentional Residual Dense Factorized Network for Real-Time Semantic Segmentation
Long Lan, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo |
ICANN (3) | 2 |
| 2019 | Person re-identification via adaptive verification loss
Hui Tian 0005, Xiang Zhang 0008, Long Lan, Zhigang Luo |
Neurocomputing | 3 |
| 2019 | Label guided correlation hashing for large-scale cross-modal retrieval
Guohua Dong, Xiang Zhang 0008, Long Lan, Zhigang Luo |
Multim. Tools Appl. | 3 |
| 2019 | Stacked Marginal Time Warping for Temporal Alignment
Xiang Zhang 0008, Liquan Nie, Long Lan, Xuhui Huang, Zhigang Luo |
Neural Process. Lett. | 3 |
| 2019 | Nonnegative Constrained Graph Based Canonical Correlation Analysis for Multi-view Feature Learning
Huibin Tan, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo |
Neural Process. Lett. | 3 |
| 2018 | Margin-Embedding Canonical Correlation Analysis with Feature Selection for Person Re-IdentificationabstractCanonical correlation analysis (CCA) is a classical subspace learning method of capturing the common semantic information underlying multi-view data. It has been used in person re-identification (re-ID) task by treating the task of matching identical individuals across non-overlapping multi-cameras as a multi-view learning problem. However, CCA-based reID methods still achieve unsatisfactory results because few jointly consider discriminative margin information and selecting importantly relevant features. To address this issue, we propose a novel l2,1-norm regularized margin-embedding CCA ( l2,1-MCCA), which learns a generalized discriminative subspace by employing more discriminative margin information. Moreover, the new method enforces the l2,1-norm regularization term over the learned subspace to identify the relevant features. Both lightweight and effective schemes can benefit from each other and endeavor to enlarge the interclass variations whilst reducing the intra-class variations. Experiments on three popular datasets show the efficacy of l2,1-MCCA as compared with recently representative re-ID methods. Linfei Ma, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo |
ICIP | 3 |
| 2018 | Graph-Laplacian Correlated Low-Rank Representation for Subspace ClusteringabstractSubspace clustering seeks to segment a given unlabeled data into clusters with the hope of each cluster corresponding to a union of low-dimensional subspaces. Among them, low-rank representation (LRR) is a promising potential method which intends to build a good affinity matrix by using the self-expression of inputs. However, it completely ignores the important data locality. Although several works in this regard have considered the local geometric structure through the Laplacian regularizer, they also neglect the correlation of the data. In this paper, we propose a graph-Laplacian correlated low-rank representation model (GCLRR) to address such an issue. Particularly, GCLRR factorizes the self-expression as the product of two low-dimensional matrices, of which one is the latent representation of the self-expression. On the basis of the latent representation, the Laplacian regularizer is integrated with the orthogonal constraint together and behaves like the spectral clustering. Moreover, we devise a Frobenius norm based trace loss and use it to constrain both the latent representation and the self-expression to capture the correlation of the data. Our improved trace loss is more efficient than the original one. More importantly, GCLRR provides an effective unified framework to seamlessly integrate both aspects above. Then, we optimize GCLRR in the frame of alternating direction method (ADM) and fortunately derive the analytical solution to each subproblem. Experiments of motion segmentation and image clustering confirm the efficacy of the proposed GCLRR. Huayue Cai, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo |
ICIP | 4 |
| 2018 | Discrete Graph Hashing via Affine TransformationabstractIn unsupervised graph-based hashing for large-scale image retrieval, many efforts have been made to bridge the gap between the learned graph embedding and the corresponding binary codes. Relatively, few studies focus on the issue of the discrimination of graph embedding. In this paper, we firstly devise a discrete graph hashing model (DGH) that smooths graph embedding and simultaneously solving binary codes under the balanced discrete constraint, which equals a novel method of jointly learning graph embedding and spectral rotation, theoretically. To further induce discriminant graph embedding, we substitute affine transformation for spectral rotation in our DGH (abbreviated as ADGH). This is because affine transformation can accommodate both rotational angle and distance of graph embedding, while respecting the neighborhood structure among most samples. Besides, each subproblem of ADGH can yield the closed-form solution. Experiments of image retrieval on three benchmark datasets show that ADGH outperforms the representative hashing methods in quantity. Guohua Dong, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo |
ICME | 3 |
| 2018 | Cross-Layer Convolutional Siamese Network for Visual Tracking
Yanyin Chen, Huibin Tan, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo |
ICONIP (2) | 5 |
| 2018 | Background Subtraction via 3D Convolutional Neural NetworksabstractBackground subtraction can be treated as the binary classification problem of highlighting the foreground region in a video whilst masking the background region, and has been broadly applied in various vision tasks such as video surveillance and traffic monitoring. However, it still remains a challenging task due to complex scenes and for lack of the prior knowledge about the temporal information. In this paper, we propose a novel background subtraction model based on 3D convolutional neural networks (3D CNNs) which combines temporal and spatial information to effectively separate the foreground from all the sequences in an end-to-end manner. Different from conventional models, we view background subtraction as three-class classification problem, i.e., the foreground, the background and the boundary. This design can obtain more reasonable results than existing baseline models. Experiments on the Change Detection 2012 dataset verify the potential of our model in both quantity and quality. Yongqiang Gao, Huayue Cai, Xiang Zhang 0008, Long Lan, Zhigang Luo |
ICPR | 4 |
| 2018 | Flexible ranking extreme learning machine based on matrix-centering transformationabstractExisting ranking ELM algorithms bias to imbalanced queries since they equally treat each pairwise error. In this study we propose a flexible ranking ELM method based on matrix-centering transformation to replace the traditional graph Laplacian matrix based methods. Specifically, we introduce a useful query-level normalized loss function and enforce the matrix-centering transformation to it to avoid training a bias model. Fortunately, by this setting, we can also greatly simplify the learning process of ELM because of the symmetry and idempotence of the centering matrix. Based on the proposed framework, three different ranking ELM variants are implemented: (a) a regularized ranking ELM model; (b) an enhanced incremental ranking ELM model; and (c) an online sequential ranking ELM model. Experimental results demonstrate that our proposed ranking ELM algorithms can obtain comparable or better performances than the state-of-the-art ranking algorithms. Shizhao Chen, Kai Chen 0020, Chuanfu Xu, Long Lan |
IJCNN | 4 |
| 2018 | Multi-granularity Hierarchical Attention Siamese Network for Visual TrackingabstractSpeed and accuracy are the two most important focuses for many visual tracking methods. Recently, siamese networks based trackers have shown very promising potentials in both aspects, which develop a twin network to measure the responses between target and hypotheses with a fully convolutional operation. However, the learned response maps are vulnerable to background clutters and scale changes as they ignore priori knowledge such as the object salience and multi-granularity cues. To explore the benefits of priori, this paper devises a multi-granularity hierarchical attention siamese network tracker (MHA-Siam) to further enhance the tracking stability without sacrificing real-time speed. Particularly, the channel-wise attention mechanism is exploited here to filter out the background while remain the salient object region; then, the response maps of the coarse-to-finer multi-layer features are fused to capture multi-granularity location information helpful for improvement in tracking stability. To make full use of them, MHA-Siam imposes the element-wise max-and-sum operation on them to induce a reliable response map for accurate location. Experiments of visual tracking on OTB benchmark shows the superiority of MHA-Siam with the competitive efficiency to its counterpart trackers. Xiang Zhang 0008, Huibin Tan, Long Lan, Zhigang Luo, Xuhui Huang |
IJCNN | 4 |
| 2018 | Ranking-Embedded Transfer Canonical Correlation Analysis for Person Re-IdentificationabstractPerson re-identification (re-ID) seeks to match the identical individuals across different cameras and is still a challenging visual task due to substantial variances of person appearance in complex scenarios. Different from most of conventional person re-ID methods, which generally reduce person re-ID task to either a multi-view learning problem or a multi- domain learning problem alone, this paper treats such a task as a multi-view multi-domain (MVMD) learning problem to exploit the both benefits by refreshing canonical correlation analysis (CCA) with two improvements, termed as ranking-embedded transfer CCA (RTCCA). Specifically, to bridge the semantic gap between different views, we first embed a ranking weight matrix into CCA to strength the correlations among the multi-view images of the same identity and simultaneously to weaken that of different identities. Furthermore, we utilize the well-known distribution metric maximum mean discrepancy (MMD) as a regularization term to reduce the domain shift between training set and testing set. More importantly, the two improvements benefit from each other and the joint merit can further boost the re-ID performance. Experiments on three benchmarks verify the efficacy of the proposed RTCCA when compared with the recently representative baseline person re-ID methods. Linfei Ma, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo |
IJCNN | 3 |
| 2018 | Low-Rank Matrix Recovery via Continuation-Based Approximate Low-Rank Minimization
Xiang Zhang 0008, Yongqiang Gao, Long Lan, Xuhui Huang, Zhigang Luo |
PRICAI (1) | 3 |
| 2018 | Interacting Tracklets for Multi-Object TrackingabstractIn this paper, we propose to exploit the interactions between non-associable tracklets to facilitate multi-object tracking. We introduce two types of tracklet interactions, close interaction and distant interaction. The close interaction imposes physical constraints between two temporally overlapping tracklets and more importantly, allows us to learn local classifiers to distinguish targets that are close to each other in the spatiotemporal domain. The distant interaction, on the other hand, accounts for the higher-order motion and appearance consistency between two temporally isolated tracklets. Our approach is modeled as a binary labeling problem and solved using the efficient Quadratic Pseudo-Boolean Optimization (QPBO). It yields promising tracking performance on the challenging PETS09 and MOT16 dataset. Our code will be made publicly available upon the acceptance of the manuscript. Long Lan, Xinchao Wang, Shiliang Zhang, Dacheng Tao, Wen Gao 0001, Thomas S. Huang |
IEEE Trans. Image Process. | 1 |
| 2016 | Online Multi-Object Tracking by Quadratic Pseudo-Boolean Optimization
Long Lan, Dacheng Tao, Chen Gong 0002, Naiyang Guan, Zhigang Luo |
IJCAI | 1 |
| 2015 | Labelwalking nonnegative matrix factorizationabstractSemi-supervised learning (SSL) utilizes plenty of unlabeled examples to boost the performance of learning from limited labeled examples. Due to its great discriminant power, SSL has been widely applied to various real-world tasks such as information retrieval, pattern recognition, and speech separa- tion. Label propagation (LP) is a popular SSL method which propagates labels through the dataset along high density areas defined by unlabeled examples, LP assumes nearby examples should share the same label, thus, it unavoidably pushes the labels to the wrong examples, especially when different la- beled examples are not strictly separated. Seed K-means uses labeled examples to initialize class centers, and avoid getting stuck in poor local optima comparing to traditional K-means, however the hard constraint of each example's membership makes Seed K-means failed in many real world applications. This paper proposes a novel label walking nonnegative matrix factorization method (LWNMF) to handle labeled examples in SSL based on the framework of NMF. LWNMF decomposes the whole dataset into the product of a basis matrix and a coefficient matrix, and to travel labels to unlabeled examples, LWNMF regards the class indicators of labeled examples as their coefficients and iteratively updates both basis matrix and coefficients of unlabeled examples. Since LWNMF learns comprehensive class centroids, labels iteratively walk to unlabeled examples through these significant centroids. Long Lan, Naiyang Guan, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo |
ICASSP | 1 |
| 2014 | Transductive nonnegative matrix factorization for semi-supervised high-performance speech separationabstractRegarding the non-negativity property of the magnitude spectrogram of speech signals, nonnegative matrix factorization (NMF) has obtained promising performance for speech separation by independently learning a dictionary on the speech signals of each known speaker. However, traditional NM-F fails to represent the mixture signals accurately because the dictionaries for speakers are learned in the absence of mixture signals. In this paper, we propose a new transductive NMF algorithm (TNMF) to jointly learn a dictionary on both speech signals of each speaker and the mixture signals to be separated. Since TNMF learns a more descriptive dictionary by encoding the mixture signals than that learned by NMF, it significantly boosts the separation performance. Experiments results on a popular TIMIT dataset show that the proposed TNMF-based methods outperform traditional NMF-based methods for separating the monophonic mixtures of speech signals of known speakers. Naiyang Guan, Long Lan, Dacheng Tao, Zhigang Luo, Xuejun Yang |
ICASSP | 2 |
| 2014 | Soft-constrained nonnegative matrix factorization via normalizationabstractSemi-supervised clustering aims at boosting the clustering performance on unlabeled samples by using labels from a few labeled samples. Constrained NMF (CNMF) is one of the most significant semi-supervised clustering methods, and it factorizes the whole dataset by NMF and constrains those labeled samples from the same class to have identical encodings. In this paper, we propose a novel soft-constrained NMF (SCNMF) method by softening the hard constraint in CNMF. Particularly, SCNMF factorizes the whole dataset into two lower-dimensional factor matrices by using multiplicative update rule (MUR). To utilize the labels of labeled samples, SCNMF iteratively normalizes both factor matrices after updating them with MURs to make encodings of labeled samples close to their label vectors. It is therefore reasonable to believe that encodings of unlabeled samples are also close to their corresponding label vectors. Such strategy significantly boosts the clustering performance even when the labeled samples are rather limited, e.g., each class owns only a single labeled sample. Since the normalization procedure never increases the computational complexity of MUR, SCNMF is quite efficient and effective in practices. Experimental results on face image datasets illustrate both efficiency and effectiveness of SCNMF compared with both NMF and CNMF. Long Lan, Naiyang Guan, Xiang Zhang 0008, Dacheng Tao, Zhigang Luo |
IJCNN | 1 |
| 2014 | Box-constrained projective nonnegative matrix factorization via augmented Lagrangian methodabstractProjective non-negative matrix factorization (P-NMF) projects a set of examples onto a subspace spanned by a non-negative basis whose transpose is regarded as the projection matrix. Since PNMF learns a natural parts-based representation, it has been successfully used in text mining and pattern recognition. However, it is non-trivial to analyze the convergence of the optimization algorithms for PNMF because its objective function is non-convex. In this paper, we propose a Box-constrained PNMF (BPNMF) method to overcome this deficiency of PNMF. In particular, BPNMF introduces an auxiliary variable, i.e., the coefficients of examples, and incorporates the following two types of constraints: 1) each entry of the basis is non-negative and upper-bounded, i.e., box-constrained, and 2) the coefficients equal to the projected points of the examples. The first box constraint makes the basis to be bound and the second equality constraint keeps its equivalence to PNMF. Similar to PNMF, BPNMF is difficult because the objective function is non-convex. To solve BPNMF, we developed an efficient algorithm in the frame of augmented Lagrangian multiplier (ALM) method and proved that the ALM-based algorithm converges to local minima. Experimental results on two face image datasets demonstrate the effectiveness of BPNMF compared with the representative methods. Xiang Zhang 0008, Naiyang Guan, Long Lan, Dacheng Tao, Zhigang Luo |
IJCNN | 3 |
| 2012 | Graph Based Semi-supervised Non-negative Matrix Factorization for Document ClusteringabstractNon-negative matrix factorization (NMF) approximates a non-negative matrix by the product of two low-rank matrices and achieves good performance in clustering. Recently, semi-supervised NMF (SS-NMF) further improves the performance by incorporating part of the labels of few samples into NMF. In this paper, we proposed a novel graph based SS-NMF (GSS-NMF). For each sample, GSS-NMF minimizes its distances to the same labeled samples and maximizes the distances against different labeled samples to incorporate the discriminative information. Since both labeled and unlabeled samples are embedded in the same reduced dimensional space, the discriminative information from the labeled samples is successfully transferred to the unlabeled samples, and thus it greatly improves the clustering performance. Since the traditional multiplicative update rule converges slowly, we applied the well-known projected gradient method to optimizing GSS-NMF and the proposed algorithm can be applied to optimizing other manifold regularized NMF efficiently. Experimental results on two popular document datasets, i.e., Reuters21578 and TDT-2, show that GSS-NMF outperforms the representative SS-NMF algorithms. Naiyang Guan, Xuhui Huang, Long Lan, Zhigang Luo, Xiang Zhang 0008 |
ICMLA (1) | 3 |
| 2012 | Sparse Representation Based Discriminative Canonical Correlation Analysis for Face RecognitionabstractCanonical correlation analysis (CCA) has been widely used in pattern recognition and machine learning. However, both CCA and its extensions sometimes cannot give satisfactory results. In this paper, we propose a new CCA-type method termed sparse representation based discriminative CCA (SPDCCA) by incorporating sparse representation and discriminative information simultaneously into traditional CCA. In particular, SPDCCA not only preserves the sparse reconstruction relationship within data based on sparse representation, but also preserves the maximum-margin based discriminative information, and thus it further enhances the classification performance. Experimental results on Yale, Extended Yale B, and ORL datasets show that SPDCCA outperforms both CCA and its extensions including KCCA, LPCCA and LDCCA in face recognition. Naiyang Guan, Xiang Zhang 0008, Zhigang Luo, Long Lan |
ICMLA (1) | 4 |
| 2012 | Semi-supervised Non-negative Patch Alignment FrameworkabstractNon-negative matrix factorization (NMF) learns the latent semantic space more direct and reliable than the latent semantic indexing (LSI) and the spectral clustering methods, thus performs well in document clustering. Recently, semi-supervised NMF such as N2S2L, CNMF and unsupervised method such as GNMF significantly improve the face recognition performance, but they are designed for classification. In this paper, we combine both geometric structure and label information with NMF under the non-negative patch alignment framework (NPAF) to form SS-NPAF. Due to this combination, it greatly improves the clustering performance. To optimize SS-NPAF, we apply the well-known projected gradient method to overcome the slow convergence problem of the mostly used multiplicative update rule. Experimental results on two popular document datasets, i.e., Reuters21578 and TDT-2, show that SS-NPAF outperforms the representative SS-NMF algorithms. Long Lan, Xuhui Huang, Naiyang Guan, Zhigang Luo, Xiang Zhang 0008 |
ICMLA (1) | 1 |