EDBT 2026 Demo / reviewers in the wild / expert
Hao Luo 0004
dblp:14/3727-4
· DBLP profile ↗
40ranked-venue papers
4as first author
34since 2021 · last 2026
0000-0002-6405-4011ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 1 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 20 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DyDiT++: Diffusion Transformers With Timestep and Spatial Dynamics for Efficient Visual GenerationabstractDiffusion Transformer (DiT), an emerging diffusion model for visual generation, has demonstrated superior performance but suffers from substantial computational costs. Our investigations reveal that these costs primarily stem from the static inference paradigm, which inevitably introduces redundant computation in certain diffusion timesteps and spatial regions. To overcome this inefficiency, we propose Dynamic Diffusion Transformer (DyDiT), an architecture that dynamically adjusts its computation along both timestep and spatial dimensions. Specifically, we introduce a Timestep-wise Dynamic Width (TDW) approach that adapts model width conditioned on the generation timesteps. In addition, we design a Spatial-wise Dynamic Token (SDT) strategy to avoid redundant computation at unnecessary spatial locations. TDW and SDT can be seamlessly integrated into DiT and significantly accelerate the generation process. Building on these designs, we present an extended version, DyDiT++, with improvements in three key aspects. First, it extends the generation mechanism of DyDiT beyond diffusion to flow matching, demonstrating that our method can also accelerate flow-matching-based generation, enhancing its versatility. Furthermore, we enhance DyDiT to tackle more complex visual generation tasks, including video generation and text-to-image generation, thereby broadening its real-world applications. Finally, to address the high cost of full fine-tuning and democratize technology access, we investigate the feasibility of training DyDiT in a parameter-efficient manner and introduce timestep-based dynamic LoRA (TD-LoRA). Extensive experiments on diverse visual generation models, including DiT, SiT, Latte, and FLUX, demonstrate the effectiveness of DyDiT++. Remarkably, with $< $<3% additional fine-tuning iterations, our approach reduces the FLOPs of DiT-XL by 51%, yielding 1.73× realistic speedup on hardware, and achieves a competitive FID score of 2.07 on ImageNet. Wangbo Zhao, Yizeng Han, Jiasheng Tang, Kai Wang 0036, Hao Luo 0004, Yibing Song, Gao Huang 0001, Fan Wang 0019, Yang You 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | AnyI2V: Animating Any Conditional Image with Motion Control
Ziye Li, Hao Luo 0004, Xincheng Shuai, Henghui Ding |
ICCV | 2 |
| 2025 | Preacher: Paper-to-Video Agentic System
Ling Yang 0006, Hao Luo 0004, Fan Wang 0019, Hongyan Li 0002, Mengdi Wang 0001 |
ICCV | 3 |
| 2025 | Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation
Xincheng Shuai, Henghui Ding, Zhenyuan Qin, Hao Luo 0004, Xingjun Ma, Dacheng Tao |
ICCV | 4 |
| 2025 | PlayerOne: Egocentric World SimulatorabstractWe introduce PlayerOne, the first egocentric realistic world simulator, facilitating immersive and unrestricted exploration within vividly dynamic environments. Given an egocentric scene image from the user, PlayerOne can accurately construct the corresponding world and generate egocentric videos that are strictly aligned with the real-scene human motion of the user captured by an exocentric camera. PlayerOne is trained in a coarse-to-fine pipeline that first performs pretraining on large-scale egocentric text-video pairs for coarse-level egocentric understanding, followed by finetuning on synchronous motion-video data extracted from egocentric-exocentric video datasets with our automatic construction pipeline. Besides, considering the varying importance of different components, we design a part-disentangled motion injection scheme, enabling precise control of part-level movements. In addition, we devise a joint reconstruction framework that progressively models both the 4D scene and video frames, ensuring scene consistency in the long-form video generation. Experimental results demonstrate its great generalization ability in precise control of varying human movements and world-consistent modeling of diverse scenarios. It marks the first endeavor into egocentric real-world simulation and can pave the way for the community to delve into fresh frontiers of world modeling and its diverse applications. Yuanpeng Tu, Hao Luo 0004, Xi Chen 0119, Xiang Bai, Fan Wang 0019, Hengshuang Zhao |
NeurIPS | 2 |
| 2025 | Compressing Vision Transformer from the View of Model Property in Frequency Domain
Zhenyu Wang 0008, Xuemei Xie, Hao Luo 0004, Weisheng Dong, Yongxu Liu 0001, Fan Wang 0019, Guangming Shi |
Int. J. Comput. Vis. | 3 |
| 2025 | Exploring the Coordination of Frequency and Attention in Masked Image ModelingabstractRecently, masked image modeling (MIM), which learns visual representations by reconstructing the masked patches of an image, has become a popular self-supervised paradigm. However, the pre-training of MIM always takes massive time due to the large-scale data and large-size backbones. We mainly attribute it to the random patch masking in previous MIM works, which fails to leverage the crucial semantic information for effective visual representation learning. To tackle this issue, we propose the Frequency & Attention-driven Masking and Throwing Strategy (FAMT), which can detect semantic patches and reduce the number of training patches to boost model performance and training efficiency simultaneously. Specifically, FAMT utilizes the self-attention mechanism to extract semantic information from the image for masking during training in an unsupervised manner. However, attention alone could sometimes focus on inappropriate areas regarding the semantic information. Thus, we are motivated to incorporate the information from the frequency domain into the self-attention mechanism to derive the sampling weights for masking, which captures semantic patches for visual representation learning. Furthermore, we introduce a patch throwing strategy based on the derived sampling weights to reduce the training cost. FAMT can be seamlessly integrated as a plug-and-play module and surpasses previous works, e.g. reducing the training phase time by nearly 50% and improving the linear probing accuracy of MAE by $1.8$ % ~ $ 6.3$ % across various datasets, including CIFAR-10/100, Tiny ImageNet, and ImageNet-1K. FAMT also demonstrates superior performance in downstream detection and segmentation tasks. Jie Gui, Tuo Chen, Minjing Dong, Zhengqi Liu, Hao Luo 0004, James T. Kwok, Yuan Yan Tang |
IEEE Trans. Image Process. | 5 |
| 2024 | BVT-IMA: Binary Vision Transformer with Information-Modified AttentionabstractAs a compression method that can significantly reduce the cost of calculations and memories, model binarization has been extensively studied in convolutional neural networks. However, the recently popular vision transformer models pose new challenges to such a technique, in which the binarized models suffer from serious performance drops. In this paper, an attention shifting is observed in the binary multi-head self-attention module, which can influence the information fusion between tokens and thus hurts the model performance. From the perspective of information theory, we find a correlation between attention scores and the information quantity, further indicating that a reason for such a phenomenon may be the loss of the information quantity induced by constant moduli of binarized tokens. Finally, we reveal the information quantity hidden in the attention maps of binary vision transformers and propose a simple approach to modify the attention values with look-up information tables so that improve the model performance. Extensive experiments on CIFAR-100/TinyImageNet/ImageNet-1k demonstrate the effectiveness of the proposed information-modified attention on binary vision transformers. Zhenyu Wang 0008, Hao Luo 0004, Xuemei Xie, Fan Wang 0019, Guangming Shi |
AAAI | 2 |
| 2024 | Enhancing Hyperspectral Images via Diffusion Model and Group-Autoencoder Super-resolution NetworkabstractExisting hyperspectral image (HSI) super-resolution (SR) methods struggle to effectively capture the complex spectral-spatial relationships and low-level details, while diffusion models represent a promising generative model known for their exceptional performance in modeling complex relations and learning high and low-level visual features. The direct application of diffusion models to HSI SR is hampered by challenges such as difficulties in model convergence and protracted inference time. In this work, we introduce a novel Group-Autoencoder (GAE) framework that synergistically combines with the diffusion model to construct a highly effective HSI SR model (DMGASR). Our proposed GAE framework encodes high-dimensional HSI data into low-dimensional latent space where the diffusion model works, thereby alleviating the difficulty of training the diffusion model while maintaining band correlation and considerably reducing inference time. Experimental results on both natural and remote sensing hyperspectral datasets demonstrate that the proposed method is superior to other state-of-the-art methods both visually and metrically. Zhaoyang Wang 0003, Mingyang Zhang 0002, Hao Luo 0004, Maoguo Gong |
AAAI | 4 |
| 2024 | CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge TransferabstractCross-lingual cross-modal retrieval has garnered increasing attention recently, which aims to achieve the alignment between vision and target language (V-T) without using any annotated V-T data pairs. Current methods employ machine translation (MT) to construct pseudo-parallel data pairs, which are then used to learn a multi-lingual and multi-modal embedding space that aligns visual and target-language representations. However, the large heterogeneous gap between vision and text, along with the noise present in target language translations, poses significant challenges in effectively aligning their representations. To address these challenges, we propose a general framework, Cross-Lingual to Cross-Modal (CL2CM), which improves the alignment between vision and target language using cross-lingual transfer. This approach allows us to fully leverage the merits of multi-lingual pre-trained models (e.g., mBERT) and the benefits of the same modality structure, i.e., smaller gap, to provide reliable and comprehensive semantic correspondence (knowledge) for the cross-modal network. We evaluate our proposed approach on two multilingual image-text datasets, Multi30K and MSCOCO, and one video-text dataset, VATEX. The results clearly demonstrate the effectiveness of our proposed method and its high potential for large-scale retrieval. Fan Wang 0019, Jianfeng Dong, Hao Luo 0004 |
AAAI | 4 |
| 2024 | Accelerating Parallel Sampling of Diffusion ModelsabstractDiffusion models have emerged as state-of-the-art generative models for image generation. However, sampling from diffusion models is usually time-consuming due to the inherent autoregressive nature of their sampling process. In this work, we propose a novel approach that accelerates the sampling of diffusion models by parallelizing the autoregressive process. Specifically, we reformulate the sampling process as solving a system of triangular nonlinear equations through fixed-point iteration. With this innovative formulation, we explore several systematic techniques to further reduce the iteration steps required by the solving process. Applying these techniques, we introduce ParaTAA, a universal and training-free parallel sampling algorithm that can leverage extra computational and memory resources to increase the sampling speed. Our experiments demonstrate that ParaTAA can decrease the inference steps required by common sequential sampling algorithms such as DDIM and DDPM by a factor of 4$\sim$14 times. Notably, when applying ParaTAA with 100 steps DDIM for Stable Diffusion, a widely-used text-to-image diffusion model, it can produce the same images as the sequential sampling in only 7 inference steps. The code is available at https://github.com/TZW1998/ParaTAA-Diffusion. Zhiwei Tang, Jiasheng Tang, Hao Luo 0004, Fan Wang 0019, Tsung-Hui Chang |
ICML | 3 |
| 2024 | DiffAug: Enhance Unsupervised Contrastive Learning with Domain-Knowledge-Free Diffusion-based Data AugmentationabstractUnsupervised Contrastive learning has gained prominence in fields such as vision, and biology, leveraging predefined positive/negative samples for representation learning. Data augmentation, categorized into hand-designed and model-based methods, has been identified as a crucial component for enhancing contrastive learning. However, hand-designed methods require human expertise in domain-specific data while sometimes distorting the meaning of the data. In contrast, generative model-based approaches usually require supervised or large-scale external data, which has become a bottleneck constraining model training in many domains. To address the problems presented above, this paper proposes DiffAug, a novel unsupervised contrastive learning technique with diffusion mode-based positive data generation. DiffAug consists of a semantic encoder and a conditional diffusion model; the conditional diffusion model generates new positive samples conditioned on the semantic encoding to serve the training of unsupervised contrast learning. With the help of iterative training of the semantic encoder and diffusion model, DiffAug improves the representation ability in an uninterrupted and unsupervised manner. Experimental evaluations show that DiffAug outperforms hand-designed and SOTA model-based augmentation methods on DNA sequence, visual, and bio-feature datasets. The code for review is released at DiffAug CODE. Zelin Zang, Hao Luo 0004, Kai Wang 0036, Fan Wang 0019, Stan Z. Li, Yang You 0001 |
ICML | 2 |
| 2024 | Adaptive Query Selection for Camouflaged Instance SegmentationabstractCamouflaged instance segmentation is a challenging task due to the various aspects such as color, structure, lighting, etc., of object instances embedded in complex backgrounds. Although the current DETR-based scheme simplifies the pipeline, it suffers from a large number of object queries, leading to many false positive instances. To address this issue, we propose an adaptive query selection mechanism. Our research reveals that a large number of redundant queries scatter the extracted features of the camouflaged instances. To remove these redundant queries with weak correlation, we evaluate the importance of the object query from the perspectives of information entropy and volatility. Moreover, we observed that occlusion and overlapping instances significantly impact the accuracy of the selection mechanism. Therefore, we design a boundary location embedding mechanism that incorporates fake instance boundaries to obtain better location information for more accurate query instance matching. We conducted extensive experiments on two challenging camouflaged instance segmentation datasets, namely COD10K and NC4K, and demonstrated the effectiveness of our proposed model. Compared with the OSFormer, our model significantly improves the performance by 3.8% AP and 5.6% AP with less computational cost, achieving the state-of-the-art of 44.8 AP and 48.1 AP with ResNet-50 on the COD10K and NC4K test-dev sets, respectively. Pichao Wang, Hao Luo 0004, Fan Wang 0019 |
ACM Multimedia | 3 |
| 2024 | SCT: A Simple Baseline for Parameter-Efficient Fine-Tuning via Salient Channels
Hengyuan Zhao, Pichao Wang, Hao Luo 0004, Fan Wang 0019, Zheng Shou 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | A Survey on Self-Supervised Learning: Algorithms, Applications, and Future TrendsabstractDeep supervised learning algorithms typically require a large volume of labeled data to achieve satisfactory performance. However, the process of collecting and labeling such data can be expensive and time-consuming. Self-supervised learning (SSL), a subset of unsupervised learning, aims to learn discriminative features from unlabeled data without relying on human-annotated labels. SSL has garnered significant attention recently, leading to the development of numerous related algorithms. However, there is a dearth of comprehensive studies that elucidate the connections and evolution of different SSL variants. This paper presents a review of diverse SSL methods, encompassing algorithmic aspects, application domains, three key trends, and open research questions. First, we provide a detailed introduction to the motivations behind most SSL algorithms and compare their commonalities and differences. Second, we explore representative applications of SSL in domains such as image processing, computer vision, and natural language processing. Lastly, we discuss the three primary trends observed in SSL research and highlight the open questions that remain. Jie Gui, Tuo Chen, Jing Zhang 0037, Qiong Cao, Zhenan Sun, Hao Luo 0004, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Dynamic gradient reactivation for backward compatible person re-identification
Xiao Pan 0001, Hao Luo 0004, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009, Jianming Zhang 0005, Jianyang Gu, Peike Li |
Pattern Recognit. | 2 |
| 2024 | Modality Bias Calibration Network via Information Disentanglement for Visible-Infrared Person ReidentificationabstractVisible–infrared person reidentification (VI-ReID) in social surveillance systems involves analyzing social behavior using nonoverlapping cross-modality camera sets. It often has poor retrieval performance under modality gap. One way to alleviate such the modality discrepancy is to learn shared person features that are generalizable across different modalities. However, because of significant differences in color between the visible and infrared images, the learned share features are always inclined to specific information of corresponding modality. To this end, we propose a modality bias calibration network (MBCNet) that filters out identity-irrelevant interference and recalibrates the learned modality-shared features. Specifically, to emphasize the modality-shared cues, we employ a feature decomposition module in the feature-level to filter out style variations and extract identity-relevant discriminative cues from the residual feature. In order to achieve a better disentanglement, a dual ranking entropy constraint is further proposed to ensure that the learned features contain only identity-relevant information and discard style-relevant information. Simultaneously, we design a decorrelated orthogonality Loss to ensure the disentangled features are not correlated with each other. Through comprehensive experiments, we demonstrate that MBCNet significantly improves the cross-modality retrieval performance in social surveillance systems and effectively addresses the modality bias training issue. Hao Luo 0004, Xiantao Peng, Wei Jiang 0009 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Region Generation and Assessment Network for Occluded Person Re-IdentificationabstractPerson Re-identification (ReID) plays a more and more crucial role in recent years with a wide range of applications. Existing ReID methods are suffering from the challenges of misalignment and occlusions, which degrade the performance dramatically. Most methods tackle such challenges by utilizing external tools to locate body parts or exploiting matching strategies. Nevertheless, the inevitable domain gap between the datasets utilized for external tools and the ReID datasets and the complicated matching process make these methods unreliable and sensitive to noises. In this paper, we propose a Region Generation and Assessment Network (RGANet) to effectively and efficiently detect the human body regions and highlight the important regions. In the proposed RGANet, we first devise a Region Generation Module (RGM) which utilizes the pre-trained CLIP to locate the human body regions using semantic prototypes extracted from text descriptions. Learnable prompt is designed to eliminate domain gap between CLIP datasets and ReID datasets. Then, to measure the importance of each generated region, we introduce a Region Assessment Module (RAM) that assigns confidence scores to different regions and reduces the negative impact of the occlusion regions by lower scores. The RAM consists of a discrimination-aware indicator and an invariance-aware indicator, where the former indicates the capability to distinguish from different identities and the latter represents consistency among the images of the same class of human body regions. Extensive experimental results for six widely-used benchmarks including three tasks (occluded, partial, and holistic) demonstrate the superiority of RGANet against state-of-the-art methods. Shuting He, Kai Wang 0036, Hao Luo 0004, Fan Wang 0019, Wei Jiang 0009, Henghui Ding |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | VGSG: Vision-Guided Semantic-Group Network for Text-Based Person SearchabstractText-based Person Search (TBPS) aims to retrieve images of target pedestrian indicated by textual descriptions. It is essential for TBPS to extract fine-grained local features and align them crossing modality. Existing methods utilize external tools or heavy cross-modal interaction to achieve explicit alignment of cross-modal fine-grained features, which is inefficient and time-consuming. In this work, we propose a Vision-Guided Semantic-Group Network (VGSG) for text-based person search to extract well-aligned fine-grained visual and textual features. In the proposed VGSG, we develop a Semantic-Group Textual Learning (SGTL) module and a Vision-guided Knowledge Transfer (VGKT) module to extract textual local features under the guidance of visual local clues. In SGTL, in order to obtain the local textual representation, we group textual features from the channel dimension based on the semantic cues of language expression, which encourages similar semantic patterns to be grouped implicitly without external tools. In VGKT, a vision-guided attention is employed to extract visual-related textual features, which are inherently aligned with visual cues and termed vision-guided textual features. Furthermore, we design a relational knowledge transfer, including a vision-language similarity transfer and a class probability transfer, to adaptively propagate information of the vision-guided textual features to semantic-group textual features. With the help of relational knowledge transfer, VGKT is capable of aligning semantic-group textual features with corresponding visual features without external tools and complex pairwise interaction. Experimental results on two challenging benchmarks demonstrate its superiority over state-of-the-art methods. Shuting He, Hao Luo 0004, Wei Jiang 0009, Xudong Jiang 0001, Henghui Ding |
IEEE Trans. Image Process. | 2 |
| 2024 | Dual-View Curricular Optimal Transport for Cross-Lingual Cross-Modal RetrievalabstractCurrent research on cross-modal retrieval is mostly English-oriented, as the availability of a large number of English-oriented human-labeled vision-language corpora. In order to break the limit of non-English labeled data, cross-lingual cross-modal retrieval (CCR) has attracted increasing attention. Most CCR methods construct pseudo-parallel vision-language corpora via Machine Translation (MT) to achieve cross-lingual transfer. However, the translated sentences from MT are generally imperfect in describing the corresponding visual contents. Improperly assuming the pseudo-parallel data are correctly correlated will make the networks overfit to the noisy correspondence. Therefore, we propose Dual-view Curricular Optimal Transport (DCOT) to learn with noisy correspondence in CCR. In particular, we quantify the confidence of the sample pair correlation with optimal transport theory from both the cross-lingual and cross-modal views, and design dual-view curriculum learning to dynamically model the transportation costs according to the learning stage of the two views. Extensive experiments are conducted on two multilingual image-text datasets and one video-text dataset, and the results demonstrate the effectiveness and robustness of the proposed method. Besides, our proposed method also shows a good expansibility to cross-lingual image-text baselines and a decent generalization on out-of-domain data. Shuhui Wang, Hao Luo 0004, Jianfeng Dong, Fan Wang 0019, Xun Wang 0007, Meng Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Frequency Domain Disentanglement for Arbitrary Neural Style TransferabstractArbitrary neural style transfer has been a popular research topic due to its rich application scenarios. Effective disentanglement of content and style is the critical factor for synthesizing an image with arbitrary style. The existing methods focus on disentangling feature representations of content and style in the spatial domain where the content and style components are innately entangled and difficult to be disentangled clearly. Therefore, these methods always suffer from low-quality results because of the sub-optimal disentanglement. To address such a challenge, this paper proposes the frequency mixer (FreMixer) module that disentangles and re-entangles the frequency spectrum of content and style components in the frequency domain. Since content and style components have different frequency-domain characteristics (frequency bands and frequency patterns), the FreMixer could well disentangle these two components. Based on the FreMixer module, we design a novel Frequency Domain Disentanglement (FDD) framework for arbitrary neural style transfer. Qualitative and quantitative experiments verify that the proposed method can render better stylized results compared to the state-of-the-art methods. Hao Luo 0004, Pichao Wang, Zhibin Wang 0004, Shang Liu 0002, Fan Wang 0019 |
AAAI | 2 |
| 2023 | Good Helper Is around You: Attention-Driven Masked Image ModelingabstractIt has been witnessed that masked image modeling (MIM) has shown a huge potential in self-supervised learning in the past year. Benefiting from the universal backbone vision transformer, MIM learns self-supervised visual representations through masking a part of patches of the image while attempting to recover the missing pixels. Most previous works mask patches of the image randomly, which underutilizes the semantic information that is beneficial to visual representation learning. On the other hand, due to the large size of the backbone, most previous works have to spend much time on pre-training. In this paper, we propose Attention-driven Masking and Throwing Strategy (AMT), which could solve both problems above. We first leverage the self-attention mechanism to obtain the semantic information of the image during the training process automatically without using any supervised methods. Masking strategy can be guided by that information to mask areas selectively, which is helpful for representation learning. Moreover, a redundant patch throwing strategy is proposed, which makes learning more efficient. As a plug-and-play module for masked image modeling, AMT improves the linear probing accuracy of MAE by 2.9% ~ 5.9% on CIFAR-10/100, STL-10, Tiny ImageNet, and ImageNet-1K, and obtains an improved performance with respect to fine-tuning accuracy of MAE and SimMIM. Moreover, this design also achieves superior performance on downstream detection and segmentation tasks. Zhengqi Liu, Jie Gui, Hao Luo 0004 |
AAAI | 3 |
| 2023 | Beyond Appearance: A Semantic Controllable Self-Supervised Learning Framework for Human-Centric Visual TasksabstractHuman-centric visual tasks have attracted increasing research attention due to their widespread applications. In this paper, we aim to learn a general human representation from massive unlabeled human images which can benefit downstream human-centric tasks to the maximum extent. We call this method SOLIDER, a Semantic cOntrollable seLf-supervIseD lEaRning framework. Unlike the existing self-supervised learning methods, prior knowledge from human images is utilized in SOLIDER to build pseudo semantic labels and import more semantic information into the learned representation. Meanwhile, we note that different downstream tasks always require different ratios of semantic information and appearance information. For example, human parsing requires more semantic information, while person re-identification needs more appearance information for identification purpose. So a single learned representation cannot fit for all requirements. To solve this problem, SOLIDER introduces a conditional network with a semantic controller. After the model is trained, users can send values to the controller to produce representations with different ratios of semantic information, which can fit different needs of downstream tasks. Finally, SOLIDER is verified on six downstream human-centric visual tasks. It outperforms state of the arts and builds new baselines for these tasks. The code is released in https://github.com/tinyvision/SOLIDER. Jian Jia, Hao Luo 0004, Fan Wang 0019, Rong Jin 0001, Xiuyu Sun |
CVPR | 4 |
| 2023 | MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReIDabstractNeural Architecture Search (NAS) has been increasingly appealing to the society of object Re-Identification (ReID), for that task-specific architectures significantly improve the retrieval performance. Previous works explore new optimizing targets and search spaces for NAS ReID, yet they neglect the difference of training schemes between image classification and ReID. In this work, we propose a novel Twins Contrastive Mechanism (TCM) to provide more appropriate supervision for ReID architecture search. TCM reduces the category overlaps between the training and validation data, and assists NAS in simulating real-world ReID training schemes. We then design a Multi-Scale Interaction (MSI) search space to search for rational interaction operations between multi-scale features. In addition, we introduce a Spatial Alignment Module (SAM) to further enhance the attention consistency confronted with images from different sources. Under the proposed NAS scheme, a specific architecture is automatically searched, named as MSINet. Extensive experiments demonstrate that our method surpasses state-of-the-art ReID methods on both indomain and cross-domain scenarios. Source code available in https://github.com/vimar-gu/MSINet. Jianyang Gu, Kai Wang 0036, Hao Luo 0004, Chen Chen 0114, Wei Jiang 0009, Yuqiang Fang, Shanghang Zhang, Yang You 0001, Jian Zhao 0006 |
CVPR | 3 |
| 2023 | Revisiting Vision Transformer from the View of Path EnsembleabstractVision Transformers (ViTs) are normally regarded as a stack of transformer layers. In this work, we propose a novel view of ViTs showing that they can be seen as ensemble networks containing multiple parallel paths with different lengths. Specifically, we equivalently transform the traditional cascade of multi-head self-attention (MSA) and feed-forward network (FFN) into three parallel paths in each transformer layer. Then, we utilize the identity connection in our new transformer form and further transform the ViT into an explicit multi-path ensemble network. From the new perspective, these paths perform two functions: the first is to provide the feature for the classifier directly, and the second is to provide the lower-level feature representation for subsequent longer paths. We investigate the influence of each path for the final prediction and discover that some paths even pull down the performance. Therefore, we propose the path pruning and EnsembleScale skills for improvement, which cut out the underperforming paths and reweight the ensemble components, respectively, to optimize the path combination and make the short paths focus on providing high-quality representation for subsequent paths. We also demonstrate that our path combination strategies can help ViTs go deeper and act as high-pass filters to filter out partial low-frequency signals. To further enhance the representation of paths served for subsequent paths, self-distillation is applied to transfer knowledge from the long paths to the short paths. This work calls for more future research to explain and design ViTs from new perspectives. Shuning Chang, Pichao Wang, Hao Luo 0004, Fan Wang 0019, Zheng Shou 0001 |
ICCV | 3 |
| 2022 | Scaled ReLU Matters for Training Vision TransformersabstractVision transformers (ViTs) have been an alternative design paradigm to convolutional neural networks (CNNs). However, the training of ViTs is much harder than CNNs, as it is sensitive to the training parameters, such as learning rate, optimizer and warmup epoch. The reasons for training difficulty are empirically analysed in the paper Early Convolutions Help Transformers See Better, and the authors conjecture that the issue lies with the patchify-stem of ViT models. In this paper, we further investigate this problem and extend the above conclusion: only early convolutions do not help for stable training, but the scaled ReLU operation in the convolutional stem (conv-stem) matters. We verify, both theoretically and empirically, that scaled ReLU in conv-stem not only improves training stabilization, but also increases the diversity of patch tokens, thus boosting peak performance with a large margin via adding few parameters and flops. In addition, extensive experiments are conducted to demonstrate that previous ViTs are far from being well trained, further showing that ViTs have great potential to be a better substitute of CNNs. Pichao Wang, Xue Wang 0010, Hao Luo 0004, Jingkai Zhou, Fan Wang 0019, Hao Li 0030, Rong Jin 0001 |
AAAI | 3 |
| 2022 | Unstructured Feature Decoupling for Vehicle Re-identification
Hao Luo 0004, Silong Peng, Fan Wang 0019, Chen Chen 0036, Hao Li 0030 |
ECCV (14) | 2 |
| 2022 | MimCo: Masked Image Modeling Pre-training with Contrastive TeacherabstractRecent masked image modeling (MIM) has received much attention in self-supervised learning (SSL), which requires the target model to recover the masked part of the input image. Although MIM-based pre-training methods achieve new state-of-the-art performance when transferred to many downstream tasks, the visualizations show that the learned representations are less separable, especially compared to those based on contrastive learning pre-training. This inspires us to think whether the linear separability of MIM pre-trained representation can be further improved, thereby improving the pre-training performance. Since MIM and contrastive learning tend to utilize different data augmentations and training strategies, combining these two pretext tasks is not trivial. In this work, we propose a novel and flexible pre-training framework, named MimCo, which combines MIM and contrastive learning through two-stage pre-training. Specifically, MimCo takes a pre-trained contrastive learning model as the teacher model and is pre-trained with two types of learning targets: patch-level and image-level reconstruction losses. Qiang Zhou 0001, Chaohui Yu, Hao Luo 0004, Zhibin Wang 0004, Hao Li 0030 |
ACM Multimedia | 3 |
| 2022 | VTC-LFC: Vision Transformer Compression with Low-Frequency ComponentsabstractAlthough Vision transformers (ViTs) have recently dominated many vision tasks, deploying ViT models on resource-limited devices remains a challenging problem. To address such a challenge, several methods have been proposed to compress ViTs. Most of them borrow experience in convolutional neural networks (CNNs) and mainly focus on the spatial domain. However, the compression only in the spatial domain suffers from a dramatic performance drop without fine-tuning and is not robust to noise, as the noise in the spatial domain can easily confuse the pruning criteria, leading to some parameters/channels being pruned incorrectly. Inspired by recent findings that self-attention is a low-pass filter and low-frequency signals/components are more informative to ViTs, this paper proposes compressing ViTs with low-frequency components. Two metrics named low-frequency sensitivity (LFS) and low-frequency energy (LFE) are proposed for better channel pruning and token pruning. Additionally, a bottom-up cascade pruning scheme is applied to compress different dimensions jointly. Extensive experiments demonstrate that the proposed method could save 40% ~ 60% of the FLOPs in ViTs, thus significantly increasing the throughput on practical devices with less than 1% performance drop on ImageNet-1K. Zhenyu Wang 0008, Hao Luo 0004, Pichao Wang, Fan Wang 0019, Hao Li 0030 |
NeurIPS | 2 |
| 2022 | SFGN: Representing the sequence with one super frame for video person re-identification
Xiao Pan 0001, Hao Luo 0004, Wei Jiang 0009, Jianming Zhang 0005, Jianyang Gu, Peike Li |
Knowl. Based Syst. | 2 |
| 2022 | Multi-View Evolutionary Training for Unsupervised Domain Adaptive Person Re-IdentificationabstractClustering-based approaches have been successfully applied to unsupervised domain adaptation (UDA) tasks for person re-identification (Re-ID), where no annotations are provided in target domain. However, the clustering process is sensitive to noises, leading to imperfect pseudo labels that could damage the training performance. In this work, we propose a Multi-view Evolutionary Training (MET) method to effectively reduce noises in clustering results from two dimensions. First, to improve the clustering accuracy at each time frame (i.e. snapshot quality), a Multi-view Diffusion (MvD) module is proposed. Through capturing data relationships from multiple viewpoints and aggregating their information, noises and bias from each individual viewpoint can be eliminated, and more reliable similarity matrix can be produced for clustering. Second, to improve the temporal consistency between clustering at different iterations, i.e. temporal consistency, we propose an Evolutionary Local Refinement (ELR) module, which utilizes the previous clustering results to guide and improve current results, and further make the training process more stable and robust. Extensive experiments demonstrate that our method can provide clustering results with high quality, and achieve state-of-the-art performance on UDA Re-ID. Jianyang Gu, Hao Luo 0004, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Modality-transfer generative adversarial network and dual-level unified latent representation for visible thermal Person re-identification
Wei Jiang 0009, Hao Luo 0004 |
Vis. Comput. | 3 |
| 2021 | TransReID: Transformer-based Object Re-IdentificationabstractExtracting robust feature representation is one of the key challenges in object re-identification (ReID). Although convolution neural network (CNN)-based methods have achieved great success, they only process one local neighborhood at a time and suffer from information loss on details caused by convolution and downsampling operators (e.g. pooling and strided convolution). To overcome these limitations, we propose a pure transformer-based object ReID framework named TransReID. Specifically, we first encode an image as a sequence of patches and build a transformer-based strong baseline with a few critical improvements, which achieves competitive results on several ReID benchmarks with CNN-based methods. To further enhance the robust feature learning in the context of transformers, two novel modules are carefully designed. (i) The jigsaw patch module (JPM) is proposed to rearrange the patch embeddings via shift and patch shuffle operations which generates robust features with improved discrimination ability and more diversified coverage. (ii) The side information embeddings (SIE) is introduced to mitigate feature bias towards camera/view variations by plugging in learnable embeddings to incorporate these non-visual clues. To the best of our knowledge, this is the first work to adopt a pure transformer for ReID research. Experimental results of TransReID are superior promising, which achieve state-of-the-art performance on both person and vehicle ReID benchmarks. Code is available at https://github.com/heshuting555/TransReID. Shuting He, Hao Luo 0004, Pichao Wang, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009 |
ICCV | 2 |
| 2021 | An efficient global representation constrained by Angular Triplet loss for vehicle re-identification
Jianyang Gu, Wei Jiang 0009, Hao Luo 0004, Hongyan Yu |
Pattern Anal. Appl. | 3 |
| 2020 | STNReID: Deep Convolutional Networks With Pairwise Spatial Transformer Networks for Partial Person Re-IdentificationabstractPartial person re-identification (ReID) is a challenging task because only partial information of person images is available for matching target persons. Few studies, especially on deep learning, have focused on matching partial person images with holistic person images. This study presents a novel deep partial ReID framework based on pairwise spatial transformer networks (STNReID), which can be trained on existing holistic person datasets. STNReID includes a spatial transformer network (STN) module and a ReID module. The STN module samples an affined image (a semantically corresponding patch) from the holistic image to match the partial image. The ReID module extracts the features of the holistic, partial, and affined images. Competition (or confrontation) is observed between the STN module and the ReID module, and two-stage training is applied to acquire a strong STNReID for partial ReID. Experimental results show that our STNReID obtains 66.7% and 54.6% rank-1 accuracies on Partial-ReID and Partial-iLIDS datasets, respectively. These values are at par with those obtained with state-of-the-art methods. Hao Luo 0004, Wei Jiang 0009, Chi Zhang 0026 |
IEEE Trans. Multim. | 1 |
| 2020 | A Strong Baseline and Batch Normalization Neck for Deep Person Re-IdentificationabstractThis study proposes a simple but strong baseline for deep person re-identification (ReID). Deep person ReID has achieved great progress and high performance in recent years. However, many state-of-the-art methods design complex network structures and concatenate multi-branch features. In the literature, some effective training tricks briefly appear in several papers or source codes. The present study collects and evaluates these effective training tricks in person ReID. By combining these tricks, the model achieves 94.5% rank-1 and 85.9% mean average precision on Market1501 with only using the global features of ResNet50. The performance surpasses all existing global- and part-based baselines in person ReID. We propose a novel neck structure named as batch normalization neck (BNNeck). BNNeck adds a batch normalization layer after global pooling layer to separate metric and classification losses into two different feature spaces because we observe they are inconsistent in one embedding space. Extended experiments show that BNNeck can boost the baseline, and our baseline can improve the performance of existing state-of-the-art methods. Our codes and models are available at: https://github.com/michuanhaohao/reid-strong-baseline. Hao Luo 0004, Wei Jiang 0009, Youzhi Gu, Fuxu Liu, Xingyu Liao, Shenqi Lai, Jianyang Gu |
IEEE Trans. Multim. | 1 |
| 2019 | SphereReID: Deep hypersphere manifold embedding for person re-identification
Wei Jiang 0009, Hao Luo 0004, Mengjuan Fei |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | AlignedReID++: Dynamically matching local information for person re-identification
Hao Luo 0004, Wei Jiang 0009, Jingjing Qian, Chi Zhang 0026 |
Pattern Recognit. | 1 |
| 2018 | SCPNet: Spatial-Channel Parallelism Network for Joint Holistic and Partial Person Re-identification
Hao Luo 0004, Lingxiao He, Chi Zhang 0026, Wei Jiang 0009 |
ACCV (2) | 2 |
| 2018 | Optical plasma boundary reconstruction based on least squares for EAST TokamakabstractReconstructing the shape and position of plasma is an important issue in Tokamaks. Equilibrium and fitting (EFIT) code is generally used for plasma boundary reconstruction in some Tokamaks. However, this magnetic method still has some inevitable disadvantages. In this paper, we present an optical plasma boundary reconstruction algorithm. This method uses EFIT reconstruction results as the standard to create the optimally optical reconstruction. Traditional edge detection methods cannot extract a clear plasma boundary for reconstruction. Based on global contrast, we propose an edge detection algorithm to extract the plasma boundary in the image plane. Illumination in this method is robust. The extracted boundary and the boundary reconstructed by EFIT are fitted by same-order polynomials and the transformation matrix exists. To acquire this matrix without camera calibration, the extracted plasma boundary is transformed from the image plane to the Tokamak poloidal plane by a mathematical model, which is optimally resolved by using least squares to minimize the error between the optically reconstructed result and the EFIT result. Once the transform matrix is acquired, we can optically reconstruct the plasma boundary with only an arbitrary image captured. The error between the method and EFIT is presented and the experimental results of different polynomial orders are discussed. Hao Luo 0004, Zhengping Luo 0002, Chao Xu 0001, Wei Jiang 0009 |
Frontiers Inf. Technol. Electron. Eng. | 1 |