EDBT 2026 Demo / reviewers in the wild / expert
Hao Li 0030
dblp:17/5705-30
· DBLP profile ↗
74ranked-venue papers
1as first author
59since 2021 · last 2026
0000-0001-9305-8524ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 35 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Stroke-Aware CycleGAN: Improving Low-Field MRI Image Quality for Accurate Stroke AssessmentabstractLow-field portable magnetic resonance imaging (pMRI) devices address a crucial requirement in the realm of healthcare by offering the capability for on-demand and timely access to MRI, especially in the context of routine stroke emergency. Nevertheless, images acquired by these devices often exhibit poor clarity and low resolution, resulting in their reduced potential to support precise diagnostic evaluations and lesion quantification. In this paper, we propose a 3D deep learning based model, named Stroke-Aware CycleGAN (SA-CycleGAN), to enhance the quality of low-field images for further improving diagnosis of routine stroke. Firstly, based on traditional CycleGAN, SA-CycleGAN incorporates a prior of stroke lesions by applying a novel spatial feature transform mechanism. Secondly, gradient difference losses are combined to deal with the problem that the synthesized images tend to be overly smooth. We present a dataset comprising 101 paired high-field and low-field diffusion-weighted imaging (DWI), which were acquired through dual scans of the same patient in close temporal proximity. Our experiments demonstrate that SA-CycleGAN is capable of generating images with higher quality and greater clarity compared to the original low-field DWI. Additionally, in terms of quantifying stroke lesions, SA-CycleGAN outperforms existing methods. The lesion volume exhibits a strong correlation between the generated images and the high-field images, with R=0.852. In contrast, the lesion volume correlation between the low-field images and the high-field images is notably lower, with R=0.462. Furthermore, the mean absolute difference in lesion volumes between the generated images and high-field images ( $1.73\pm 2.03$ mL) was significantly smaller than the difference between the low-field images and high-field images ( $2.53\pm 4.24$ mL). It shows that the synthesized images not only exhibit superior visual clarity compared to the low-field acquired images, but also possess a high degree of consistency with high-field images. In routine clinical practice, the proposed SA-CycleGAN offers an accessible and cost-effective means of rapidly obtaining higher-quality images, holding the potential to enhance the efficiency and accuracy of stroke diagnosis in routine clinical settings. The code and trained models will be released on GitHub: SA-CycleGAN. Ziyang Liu 0002, Xuewei Xie, Hao Li 0030, Wanlin Zhu, Yue Suo, Xia Meng, Jian Cheng 0002, Ning Wang 0018, Yihuai Wang, Bingshan Xue, Jing Jing 0002, Tao Liu 0067 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | GRPose: Learning Graph Relations for Human Image Generation with Pose PriorsabstractRecent methods using diffusion models have made significant progress in human image generation with various control signals such as pose priors. However, existing efforts are still struggling to generate high-quality images with consistent pose alignment, resulting in unsatisfactory output. In this paper, we propose a framework that delves into the graph relations of pose priors to provide control information for human image generation. The main idea is to establish a graph topological structure between the pose priors and latent representation of diffusion models to capture the intrinsic associations between different pose parts. A Progressive Graph Integrator (PGI) is designed to learn the spatial relationships of the pose priors with the graph structure, adopting a hierarchical strategy within an Adapter to gradually propagate information across different pose parts. Besides, a pose perception loss is introduced based on a pretrained pose estimation network to minimize the pose differences. Extensive qualitative and quantitative experiments conducted on the Human-Art and LAION-Human datasets clearly demonstrate that our model can achieve significant performance improvement over the latest benchmark models. Xiangchen Yin, Donglin Di, Lei Fan 0007, Hao Li 0030, Wei Chen 0089, Gouxiao Fei, Yang Song 0001, Xiao Sun 0003, Xun Yang 0001 |
AAAI | 4 |
| 2025 | DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
Donglin Di, Wenzhang Sun, Yongjia Ma, Hao Li 0030, Wei Chen 0089, Lei Fan 0007, Tonghua Su, Xun Yang 0001 |
ICCV | 5 |
| 2025 | QR-LoRA: Efficient and Disentangled Fine-Tuning via QR Decomposition for Customized GenerationabstractExisting text-to-image models often rely on parameter fine-tuning techniques such as Low-Rank Adaptation (LoRA) to customize visual attributes. However, when combining multiple LoRA models for content-style fusion tasks, unstructured modifications of weight matrices often lead to undesired feature entanglement between content and style attributes. We propose QR-LoRA, a novel fine-tuning framework leveraging QR decomposition for structured parameter updates that effectively separate visual attributes. Our key insight is that the orthogonal Q matrix naturally minimizes interference between different visual features, while the upper triangular R matrix efficiently encodes attribute-specific transformations. Our approach fixes both Q and R matrices while only training an additional task-specific $ΔR$ matrix. This structured design reduces trainable parameters to half of conventional LoRA methods and supports effective merging of multiple adaptations without cross-contamination due to the strong disentanglement properties between $ΔR$ matrices. Experiments demonstrate that QR-LoRA achieves superior disentanglement in content-style fusion tasks, establishing a new paradigm for parameter-efficient, disentangled fine-tuning in generative models. The project page is available at: https://luna-ai-lab.github.io/QR-LoRA/. Yongjia Ma, Donglin Di, Jianxun Cui, Hao Li 0030, Wei Chen 0089, Xun Yang 0001, Wangmeng Zuo |
ICCV | 5 |
| 2025 | EverybodyDance: Bipartite Graph-Based Identity Correspondence for Multi-Character AnimationabstractConsistent pose‐driven character animation has achieved remarkable progress in single‐character scenarios. However, extending these advances to multi‐character settings is non‐trivial, especially when position swap is involved. Beyond mere scaling, the core challenge lies in enforcing correct Identity Correspondence (IC) between characters in reference and generated frames. To address this, we introduce EverybodyDance, a systematic solution targeting IC correctness in multi-character animation. EverybodyDance is built around the **Identity Matching Graph (IMG)**, which models characters in the generated and reference frames as two node sets in a weighted complete bipartite graph. Edge weights, computed via our proposed Mask–Query Attention (MQA), quantify the affinity between each pair of characters. Our key insight is to formalize IC correctness as a graph structural metric and to optimize it during training. We also propose a series of targeted strategies tailored for multi-character animation, including identity-embedded guidance, a multi-scale matching strategy, and pre-classified sampling, which work synergistically. Finally, to evaluate IC performance, we curate the **Identity Correspondence Evaluation** benchmark, dedicated to multi‐character IC correctness. Extensive experiments demonstrate that EverybodyDance substantially outperforms state‐of‐the‐art baselines in both IC and visual fidelity. Haotian Ling, Zequn Chen, Qiuying Chen, Donglin Di, Yongjia Ma, Hao Li 0030, Zhulin Tao, Xun Yang 0001 |
NeurIPS | 6 |
| 2025 | EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose EstimationabstractLocating 3D objects from a single RGB image via Perspective-n-Point (PnP) is a long-standing problem in computer vision. Driven by end-to-end deep learning, recent studies suggest interpreting PnP as a differentiable layer, allowing for partial learning of 2D-3D point correspondences by backpropagating the gradients of pose loss. Yet, learning the entire correspondences from scratch is highly challenging, particularly for ambiguous pose solutions, where the globally optimal pose is theoretically non-differentiable w.r.t. the points. In this paper, we propose the EPro-PnP, a probabilistic PnP layer for general end-to-end pose estimation, which outputs a distribution of pose with differentiable probability density on the SE(3) manifold. The 2D-3D coordinates and corresponding weights are treated as intermediate variables learned by minimizing the KL divergence between the predicted and target pose distribution. The underlying principle generalizes previous approaches, and resembles the attention mechanism. EPro-PnP can enhance existing correspondence networks, closing the gap between PnP-based method and the task-specific leaders on the LineMOD 6DoF pose estimation benchmark. Furthermore, EPro-PnP helps to explore new possibilities of network design, as we demonstrate a novel deformable correspondence network with the state-of-the-art pose accuracy on the nuScenes 3D object detection benchmark. Hansheng Chen 0001, Wei Tian 0001, Pichao Wang, Fan Wang 0019, Lu Xiong 0001, Hao Li 0030 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | EasyOutPainter: One Step Image Outpainting With Both Continuous Multiple and ResolutionabstractImage outpainting aims to generate the content of an input sub-image outside its boundaries, which remains open for existing generative models. This paper explores image outpainting in three directions that have not been achieved in literature to our knowledge: outpainting 1) with continuous multiples (in contrast to the discrete ones by existing methods); 2) with arbitrary resolutions; and 3) in a single step (for any multiples and resolutions). The arbitrary multiple outpainting is achieved by utilizing randomly cropped views from the same image during training to capture arbitrary relative positional information. Specifically, by feeding one view and relative positional embeddings as queries, we can reconstruct another view. At inference, we generate images with arbitrary expansion multiples by inputting an anchor image and its corresponding positional embeddings. The continuous-resolution outpainting is achieved by introducing the multi-scale training strategy into generative models. Specifically, by disentangling the image resolution and the number of patches, it can generate images with arbitrary resolutions without post-processing. Meanwhile, we propose a query-based contrastive objective to make our method not rely on a pre-trained backbone network which is otherwise often required in peer methods. The comprehensive experimental results on public benchmarks show its superior performance over state-of-the-art approaches. Shaofeng Zhang, Qiang Zhou 0001, Zhibin Wang 0004, Hao Li 0030, Junchi Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image GenerationabstractVanilla text-to-image diffusion models struggle with generating accurate human images, commonly resulting in imperfect anatomies such as unnatural postures or disproportionate limbs. Existing methods address this issue mostly by fine-tuning the model with extra images or adding additional control - human-centric priors such as pose or depth maps - during the image generation phase. This paper explores the integration of these human-centric priors directly into the model fine-tuning stage, essentially eliminating the need for extra conditions at the inference stage. We realize this idea by proposing a human-centric alignment loss to strengthen human-related information from the textual prompts within the cross-attention maps. To ensure semantic detail richness and human structural accuracy during fine-tuning, we introduce scale-aware and step-wise constraints within the diffusion process, according to an indepth analysis of the cross-attention layer. Extensive experiments show that our method largely improves over state-of-the-art text-to-image models to synthesize high-quality human images based on user-written prompts. Project page: https://hcplayercvpr2024.github.io. Junyan Wang 0001, Zhenhong Sun, Zhiyu Tan, Xuanbai Chen, Hao Li 0030, Cheng Zhang 0014, Yang Song 0001 |
CVPR | 6 |
| 2024 | An Empirical Study and Analysis of Text-to-Image Generation Using Large Language Model-Powered Textual Representation
Zhiyu Tan, Mengping Yang, Luozheng Qin, Ye Qian, Qiang Zhou 0001, Cheng Zhang 0014, Hao Li 0030 |
ECCV (80) | 8 |
| 2024 | EGGen: Image Generation with Multi-entity Prior Learning through Entity GuidanceabstractDiffusion models have shown remarkable prowess in text-to-image synthesis and editing, yet they often stumble when tasked with interpreting complex prompts that describe multiple entities with specific attributes and interrelations. The generated images often contain inconsistent multi-entity representation (IMR), reflected as inaccurate presentations of the multiple entities and their attributes. Although providing spatial layout guidance improves the multi-entity generation quality in existing works, it is still challenging to handle the leakage attributes and avoid unnatural characteristics. To address the IMR challenge, we first conduct in-depth analyses of the diffusion process and attention operation, revealing that the IMR challenges largely stem from the process of cross-attention mechanisms. According to the analyses, we introduce the entity guidance generation mechanism, which maintains the integrity of the original diffusion model parameters by integrating plug-in networks. Our work advances the stable diffusion model by segmenting comprehensive prompts into distinct entity-specific prompts with bounding boxes, enabling a transition from multi-entity to single-entity generation in cross-attention layers. More importantly, we introduce entity-centric cross-attention layers that focus on individual entities to preserve their uniqueness and accuracy, alongside global entity alignment layers that refine cross-attention maps using multi-entity priors for precise positioning and attribute accuracy. Additionally, a linear attenuation module is integrated to progressively reduce the influence of these layers during inference, preventing oversaturation and preserving generation fidelity. Our comprehensive experiments demonstrate that this entity guidance generation enhances existing text-to-image models in generating detailed, multi-entity images. Zhenhong Sun, Junyan Wang 0001, Zhiyu Tan, Daoyi Dong, Hailan Ma, Hao Li 0030, Dong Gong |
ACM Multimedia | 6 |
| 2024 | Multimodal LLM Enhanced Cross-lingual Cross-modal RetrievalabstractCross-lingual cross-modal retrieval (CCR) aims to retrieve visually relevant content based on non-English queries, without relying on human-labeled cross-modal data pairs during training. One popular approach involves utilizing machine translation (MT) to create pseudo-parallel data pairs, establishing correspondence between visual and non-English textual data. However, aligning their representations poses challenges due to the significant semantic gap between vision and text, as well as the lower quality of non-English representations caused by pre-trained encoders and data noise. To overcome these challenges, we propose LECCR, a novel solution that incorporates the multi-modal large language model (MLLM) to improve the alignment between visual and non-English representations. Specifically, we first employ MLLM to generate detailed visual content descriptions and aggregate them into multi-view semantic slots that encapsulate different semantics. Then, we take these semantic slots as internal features and leverage them to interact with the visual features. By doing so, we enhance the semantic information within the visual features, narrowing the semantic gap between modalities and generating local visual semantics for subsequent multi-level matching. Additionally, to further enhance the alignment between visual and non-English features, we introduce softened matching under English guidance. This approach provides more comprehensive and reliable inter-modal correspondences between visual and non-English features. Extensive experiments on four CCR benchmarks, i.e., Multi30K, MSCOCO, VATEX, and MSR-VTT-CN, demonstrate the effectiveness of our proposed method. Code: https://github.com/LiJiaBei-7/leccr. Le Wang 0003, Qiang Zhou 0001, Zhibin Wang 0004, Hao Li 0030, Gang Hua 0001, Wei Tang 0016 |
ACM Multimedia | 5 |
| 2024 | A clinically actionable and explainable real-time risk assessment framework for stroke-associated pneumonia
Lutao Dai, Hao Li 0030, Xingquan Zhao, Zixiao Li, Haipeng Shen |
Artif. Intell. Medicine | 3 |
| 2024 | Predicting functional outcome in ischemic stroke patients using genetic, environmental, and clinical factors: a machine learning analysis of population-based prospective cohort studyabstractIschemic stroke (IS) is a leading cause of adult disability that can severely compromise the quality of life for patients. Accurately predicting the IS functional outcome is crucial for precise risk stratification and effective therapeutic interventions. We developed a predictive model integrating genetic, environmental, and clinical factors using data from 7819 IS patients in the Third China National Stroke Registry. Employing an 80:20 split, we randomly divided the dataset into development and internal validation cohorts. The discrimination and calibration performance of models were evaluated using the area under the receiver operating characteristic curves (AUC) for discrimination and Brier score with calibration curve in the internal validation cohort. We conducted genome-wide association studies (GWAS) in the development cohort, identifying rs11109607 (ANKS1B) as the most significant variant associated with IS functional outcome. We employed principal component analysis to reduce dimensionality on the top 100 significant variants identified by the GWAS, incorporating them as genetic factors in the predictive model. We employed a machine learning algorithm capable of identifying nonlinear relationships to establish predictive models for IS patient functional outcome. The optimal model was the XGBoost model, which outperformed the logistic regression model (AUC 0.818 versus 0.756, P < .05) and significantly improved reclassification efficiency. Our study innovatively incorporated genetic, environmental, and clinical factors for predicting the IS functional outcome in East Asian populations, thereby offering novel insights into IS functional outcome. Siding Chen, Jinfeng Yin, Hongqiu Gu, Yanfeng Shi, Cang Guo, Xia Meng, Hao Li 0030, Xinying Huang |
Briefings Bioinform. | 8 |
| 2024 | Dynamic gradient reactivation for backward compatible person re-identification
Xiao Pan 0001, Hao Luo 0004, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009, Jianming Zhang 0005, Jianyang Gu, Peike Li |
Pattern Recognit. | 5 |
| 2023 | Point-Teaching: Weakly Semi-supervised Object Detection with Point AnnotationsabstractPoint annotations are considerably more time-efficient than bounding box annotations. However, how to use cheap point annotations to boost the performance of semi-supervised object detection is still an open question. In this work, we present Point-Teaching, a weakly- and semi-supervised object detection framework to fully utilize the point annotations. Specifically, we propose a Hungarian-based point-matching method to generate pseudo labels for point-annotated images. We further propose multiple instance learning (MIL) approaches at the level of images and points to supervise the object detector with point annotations. Finally, we propose a simple data augmentation, named Point-Guided Copy-Paste, to reduce the impact of those unmatched points. Experiments demonstrate the effectiveness of our method on a few datasets and various data regimes. In particular, Point-Teaching outperforms the previous best method Group R-CNN by 3.1 AP with 5% fully labeled data and 2.3 AP with 30% fully labeled data on the MS COCO dataset. We believe that our proposed framework can largely lower the bar of learning accurate object detectors and pave the way for its broader applications. The code is available at https://github.com/YongtaoGe/Point-Teaching. Yongtao Ge, Qiang Zhou 0001, Chunhua Shen, Zhibin Wang 0004, Hao Li 0030 |
AAAI | 6 |
| 2023 | What Limits the Performance of Local Self-attention?
Jingkai Zhou, Pichao Wang, Jiasheng Tang, Fan Wang 0019, Hao Li 0030, Rong Jin 0001 |
Int. J. Comput. Vis. | 6 |
| 2022 | TransZero: Attribute-Guided Transformer for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen ones. Semantic knowledge is learned from attribute descriptions shared between different classes, which are strong prior for localization of object attribute for representing discriminative region features enabling significant visual-semantic interaction. Although few attention-based models have attempted to learn such region features in a single image, the transferability and discriminative attribute localization of visual features are typically neglected. In this paper, we propose an attribute-guided Transformer network to learn the attribute localization for discriminative visual-semantic embedding representations in ZSL, termed TransZero. Specifically, TransZero takes a feature augmentation encoder to alleviate the cross-dataset bias between ImageNet and ZSL benchmarks and improve the transferability of visual features by reducing the entangled relative geometry relationships among region features. To learn locality-augmented visual features, TransZero employs a visual-semantic decoder to localize the most relevant image regions to each attributes from a given image under the guidance of attribute semantic information. Then, the locality-augmented visual features and semantic vectors are used for conducting effective visual-semantic interaction in a visual-semantic embedding network. Extensive experiments show that TransZero achieves a new state-of-the-art on three ZSL benchmarks. The codes are available at: https://github.com/shiming-chen/TransZero. Shiming Chen 0002, Ziming Hong, Yang Liu 0069, Guosen Xie, Baigui Sun, Hao Li 0030, Qinmu Peng, Ke Lu 0002, Xinge You |
AAAI | 6 |
| 2022 | Scaled ReLU Matters for Training Vision TransformersabstractVision transformers (ViTs) have been an alternative design paradigm to convolutional neural networks (CNNs). However, the training of ViTs is much harder than CNNs, as it is sensitive to the training parameters, such as learning rate, optimizer and warmup epoch. The reasons for training difficulty are empirically analysed in the paper Early Convolutions Help Transformers See Better, and the authors conjecture that the issue lies with the patchify-stem of ViT models. In this paper, we further investigate this problem and extend the above conclusion: only early convolutions do not help for stable training, but the scaled ReLU operation in the convolutional stem (conv-stem) matters. We verify, both theoretically and empirically, that scaled ReLU in conv-stem not only improves training stabilization, but also increases the diversity of patch tokens, thus boosting peak performance with a large margin via adding few parameters and flops. In addition, extensive experiments are conducted to demonstrate that previous ViTs are far from being well trained, further showing that ViTs have great potential to be a better substitute of CNNs. Pichao Wang, Xue Wang 0010, Hao Luo 0004, Jingkai Zhou, Fan Wang 0019, Hao Li 0030, Rong Jin 0001 |
AAAI | 7 |
| 2022 | Unsupervised Visual Representation Learning by Online Constrained K-MeansabstractCluster discrimination is an effective pretext task for unsupervised representation learning, which often consists of two phases: clustering and discrimination. Clustering is to assign each instance a pseudo label that will be used to learn representations in discrimination. The main challenge resides in clustering since prevalent clustering methods (e.g., k-means) have to run in a batch mode. Besides, there can be a trivial solution consisting of a dominating cluster. To address these challenges, we first investigate the objective of clustering-based representation learning. Based on this, we propose a novel clustering-based pretext task with online Constrained K-means (CoKe). Compared with the balanced clustering that each cluster has exactly the same size, we only constrain the minimal size of each cluster to flexibly capture the inherent data structure. More importantly, our online assignment method has a theoretical guarantee to approach the global optimum. By decoupling clustering and discrimination, CoKe can achieve competitive performance when optimizing with only a single view from each instance. Extensive experiments on ImageNet and other benchmark data sets verify both the efficacy and efficiency of our proposal. Qi Qian 0001, Yuanhong Xu, Juhua Hu, Hao Li 0030, Rong Jin 0001 |
CVPR | 4 |
| 2022 | EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose EstimationabstractLocating 3D objects from a single RGB image via Perspective-n-Points (PnP) is a long-standing problem in computer vision. Driven by end-to-end deep learning, recent studies suggest interpreting PnP as a differentiable layer, so that 2D-3D point correspondences can be partly learned by backpropagating the gradient w.r.t. object pose. Yet, learning the entire set of unrestricted 2D-3D points from scratch fails to converge with existing approaches, since the deterministic pose is inherently non-differentiable. In this paper, we propose the EPro-PnP a probabilistic PnP layer for general end-to-end pose estimation, which outputs a distribution of pose on the SE(3) manifold, essentially bringing categorical Softmax to the continuous domain. The 2D-3D coordinates and corresponding weights are treated as intermediate variables learned by minimizing the KL divergence between the predicted and target pose distribution. The underlying principle unifies the existing approaches and resembles the attention mechanism. EPro-PnP significantly outperforms competitive baselines, closing the gap between PnP-based method and the task-specific leaders on the LineMOD 6DoF pose estimation and nuScenes 3D object detection benchmarks.3 Hansheng Chen 0001, Pichao Wang, Fan Wang 0019, Wei Tian 0001, Lu Xiong 0001, Hao Li 0030 |
CVPR | 6 |
| 2022 | MogFace: Towards a Deeper Appreciation on Face DetectionabstractBenefiting from the pioneering design of generic object detectors, significant achievements have been made in the field of face detection. Typically, the architectures of the backbone, feature pyramid layer, and detection head module within the face detector all assimilate the excellent experience from general object detectors. However, several effective methods, including label assignment and scale-level data augmentation strategy, fail to maintain consistent superiority when applying on the face detector directly. Concretely, the former strategy involves a vast body of hyperparameters and the latter one suffers from the challenge of scale distribution bias between different detection tasks, which both limit their generalization abilities. Furthermore, in order to provide accurate face bounding boxes for facial down-stream tasks, the face detector imperatively requires the elimination of false alarms. As a result, practical solutions on label assignment, scale-level data augmentation, and reducing false alarms are necessary for advancing face detectors. In this paper, we focus on resolving three aforementioned challenges that exiting methods are difficult to finish off and present a novel face detector, termed MogFace. In our Mogface, three key components, Adaptive Online Incremental Anchor Mining Strategy, Selective Scale Enhancement Strategy and Hierarchical Context-Aware Module, are separately proposed to boost the performance of face detectors. Finally, to the best of our knowledge, our MogFace is the best face detector on the Wider Face leader-board, achieving all champions across different testing scenarios. The code is available at https://github.com/damo-cv/MogFace. Yang Liu 0069, Fei Wang 0032, Jiankang Deng, Baigui Sun, Hao Li 0030 |
CVPR | 6 |
| 2022 | An Efficient Training Approach for Very Large Scale Face RecognitionabstractFace recognition has achieved significant progress in deep learning era due to the ultra-large-scale and well- labeled datasets. However, training on the outsize datasets is time-consuming and takes up a lot of hardware resource. Therefore, designing an efficient training approach is in- dispensable. The heavy computational and memory costs mainly result from the million-level dimensionality of the fully connected (FC) layer. To this end, we propose a novel training approach, termed Faster Face Classification (F2C), to alleviate time and cost without sacrificing the performance. This method adopts Dynamic Class Pool (DCP) for storing and updating the identities' features dy-namically, which could be regarded as a substitute for the FC layer. DCP is efficiently time-saving and cost-saving, as its smaller size with the independence from the whole face identities together. We further validate the proposed F2C method across several face benchmarks and private datasets, and display comparable results, meanwhile the speed is faster than state-of-the-art FC-based methods in terms of recognition accuracy and hardware costs. More-over, our method is further improved by a well-designed dual data loader including indentity-based and instance- based loaders, which makes it more efficient for updating DCP parameters. Kai Wang 0036, Shuo Wang 0001, Xiaojiang Peng, Baigui Sun, Hao Li 0030, Yang You 0001 |
CVPR | 9 |
| 2022 | Decoupling and Recoupling Spatiotemporal Representation for RGB-D-based Motion RecognitionabstractDecoupling spatiotemporal representation refers to decomposing the spatial and temporal features into dimension-independent factors. Although previous RGB-D-based motion recognition methods have achieved promising performance through the tightly coupled multi-modal spatiotemporal representation, they still suffer from (i) optimization difficulty under small data setting due to the tightly spatiotemporal-entangled modeling; (ii) information redundancy as it usually contains lots of marginal information that is weakly relevant to classification; and (iii) low interaction between multi-modal spatiotemporal information caused by insufficient late fusion. To alleviate these drawbacks, we propose to decouple and recouple spatiotemporal representation for RGB-D-based motion recognition. Specifically, we disentangle the task of learning spatiotemporal representation into 3 sub-tasks: (1) Learning high-quality and dimension independent features through a decoupled spatial and temporal modeling network. (2) Recoupling the decoupled representation to establish stronger space-time dependency. (3) Introducing a Cross-modal Adaptive Posterior Fusion (CAPF) mechanism to capture cross-modal spatiotemporal information from RGB-D data. Seamless combination of these novel designs forms a robust spatiotemporal representation and achieves better performance than state-of-the-art methods on four public motion datasets. Our code is available at https://github.com/damo-cv/MotionRGBD. Benjia Zhou, Pichao Wang, Jun Wan 0001, Yanyan Liang 0001, Fan Wang 0019, Zhen Lei 0001, Hao Li 0030, Rong Jin 0001 |
CVPR | 8 |
| 2022 | Unstructured Feature Decoupling for Vehicle Re-identification
Hao Luo 0004, Silong Peng, Fan Wang 0019, Chen Chen 0036, Hao Li 0030 |
ECCV (14) | 6 |
| 2022 | KVT: k-NN Attention for Boosting Vision Transformers
Pichao Wang, Xue Wang 0010, Fan Wang 0019, Ming Lin 0002, Shuning Chang, Hao Li 0030, Rong Jin 0001 |
ECCV (24) | 6 |
| 2022 | TransFGU: A Top-Down Approach to Fine-Grained Unsupervised Semantic Segmentation
Zhaoyuan Yin, Pichao Wang, Fan Wang 0019, Hanling Zhang, Hao Li 0030, Rong Jin 0001 |
ECCV (29) | 6 |
| 2022 | DLME: Deep Local-Flatness Manifold Embedding
Zelin Zang, Siyuan Li 0002, Di Wu 0057, Kai Wang 0036, Baigui Sun, Hao Li 0030, Stan Z. Li |
ECCV (21) | 8 |
| 2022 | Jmpnet: Joint Motion Prediction for Learning-Based Video CompressionabstractIn recent years, more attention is attracted by learning-based approaches in the field of video compression. Recent methods of this kind normally consist of three major components: intra-frame network, motion prediction network, and residual network, among which the motion prediction part is particularly critical for video compression. Benefiting from the optical flow which enables dense motion prediction, recent methods have shown competitive performance compared with traditional codecs. However, problems such as tail shadow and background distortion in the predicted frame remain unsolved. To tackle these problems, JMPNet is introduced in this paper to provide more accurate motion information by using both optical flow and dynamic local filter as well as an attention map to further fuse these motion information in a smarter way. Experimental results show that the proposed method surpasses state-of-the-art (SOTA) rate-distortion (RD) performance in the most data-sets. Zhenhong Sun, Zhiyu Tan, Xiuyu Sun, Fangyi Zhang, Yichen Qian, Hao Li 0030 |
ICASSP | 7 |
| 2022 | Adaptive Matching Strategy for Multi-Target Multi-Camera TrackingabstractMulti-Target Multi-Camera Tracking has a wide range of applications and is the basis for many high-level inference and prediction tasks. How to make the system perform efficiently on a large number of cameras is a crucial research issue. Previous works have proposed many matching strategies to reduce the matching range and improve the matching accuracy. However, these works require human participation when formulating matching strategies, which becomes infeasible as the scale of the camera system increases. To tackle this problem, we propose an adaptive matching strategy to replace manual rules when guiding the matching between cameras. Specifically, we use the Markov decision process to model the tracklets matching problem between cameras. Reinforcement learning and imitation learning are combined to predict a set of cameras where the tracking target might be located. The predicted candidate camera set can be used for intercamera matching and association between tracklets. Moreover, our method can be trained with or without ground truth inter-camera trajectories, making it more practical in real scenarios. We evaluate our method on the city-scale tracking dataset Cityflow, and the proposed method is sufficient to replace manual rules, and finally improve the performance of the overall MTMCT system. Chong Liu 0002, Yuqi Zhang 0001, Fan Wang 0019, Hao Li 0030, Yidong Shen |
ICASSP | 5 |
| 2022 | Image-to-Video Re-Identification via Mutual Discriminative Knowledge TransferabstractThe gap in representations between image and video makes Image-to-Video Re-identification (I2V Re-ID) challenging, and recent works formulate this problem as a knowledge distillation (KD) process. In this paper, we propose a mutual discriminative knowledge distillation framework to transfer a video-based richer representation to an image based representation more effectively. Specifically, we propose the triplet contrast loss (TCL), a novel loss designed for KD. During the KD process, the TCL loss transfers the local structure, exploits the higher order information, and mitigates the misalignment of the heterogeneous output of teacher and student networks. Compared with other losses for KD, the proposed TCL loss selectively transfers the local discriminative features from teacher to student, making it effective in the ReID. Besides the TCL loss, we adopt mutual learning to regularize both the teacher and student networks training. Extensive experiments demonstrate the effectiveness of our method on the MARS, DukeMTMC-VideoReID and VeRi-776 benchmarks. Pichao Wang, Fan Wang 0019, Hao Li 0030 |
ICASSP | 3 |
| 2022 | Graph Convolution for Re-Ranking in Person Re-IdentificationabstractNowadays, deep learning is widely applied to extract features for similarity computation in person re-identification (re-ID). However, the difference between the training data and testing data makes the performance of learned feature degraded during testing. Hence, re-ranking is proposed to mitigate this issue and various algorithms have been developed. However, most of existing re-ranking methods focus on replacing the Euclidean distance with sophisticated distance metrics, which are not friendly to downstream tasks and hard to be used for fast retrieval of massive data in real applications. In this work, we propose a graph-based re-ranking method to improve learned features while still keeping Euclidean distance as the similarity metric. Inspired by graph convolution networks, we develop an operator to propagate features over an appropriate graph. Since graph is the essential key for the propagation, two important criteria are considered for designing the graph, and different graphs are explored accordingly. Furthermore, a simple yet effective method is proposed to generate a profile vector for each tracklet in videos, which helps extend our method to video re-ID. Extensive experiments on three benchmark data sets, e.g., Market-1501, Duke, and MARS, demonstrate the effectiveness of our proposed approach. Yuqi Zhang 0001, Qi Qian 0001, Chong Liu 0002, Fan Wang 0019, Hao Li 0030, Rong Jin 0001 |
ICASSP | 6 |
| 2022 | GiraffeDet: A Heavy-Neck Paradigm for Object Detection
Yiqi Jiang, Zhiyu Tan, Junyan Wang 0001, Xiuyu Sun, Ming Lin 0002, Hao Li 0030 |
ICLR | 6 |
| 2022 | CDTrans: Cross-domain Transformer for Unsupervised Domain Adaptation
Tongkun Xu, Pichao Wang, Fan Wang 0019, Hao Li 0030, Rong Jin 0001 |
ICLR | 5 |
| 2022 | MAE-DET: Revisiting Maximum Entropy Principle in Zero-Shot NAS for Efficient Object DetectionabstractIn object detection, the detection backbone consumes more than half of the overall inference cost. Recent researches attempt to reduce this cost by optimizing the backbone architecture with the help of Neural Architecture Search (NAS). However, existing NAS methods for object detection require hundreds to thousands of GPU hours of searching, making them impractical in fast-paced research and development. In this work, we propose a novel zero-shot NAS method to address this issue. The proposed method, named MAE-DET, automatically designs efficient detection backbones via the Maximum Entropy Principle without training network parameters, reducing the architecture design cost to nearly zero yet delivering the state-of-the-art (SOTA) performance. Under the hood, MAE-DET maximizes the differential entropy of detection backbones, leading to a better feature extractor for object detection under the same computational budgets. After merely one GPU day of fully automatic design, MAE-DET innovates SOTA detection backbones on multiple detection benchmark datasets with little human intervention. Comparing to ResNet-50 backbone, MAE-DET is $+2.0%$ better in mAP when using the same amount of FLOPs/parameters, and is $1.54$ times faster on NVIDIA V100 at the same mAP. Code and pre-trained models are available here (https://github.com/alibaba/lightweight-neural-architecture-search). Zhenhong Sun, Ming Lin 0002, Xiuyu Sun, Zhiyu Tan, Hao Li 0030, Rong Jin 0001 |
ICML | 5 |
| 2022 | Dependency-Aware Traffic Management for Configuring On-demand in Service MeshesabstractService mesh is a promising micro-services architecture due to its excellent governance capabilities. Unlike traditional service invocation, configurations for governance need to be issued in the service mesh. However, we find that the control-plane traffic of governance is distributed in full by default, i.e., each service in the data plane receives all configurations. The vast majority of the configurations are redundant for a specific service. Hence, it is important and challenging to make the control plane aware of the calling relationships between services. In this paper, we propose a traffic management mechanism named DATM. Using this mechanism, the entire cluster can be dynamically controlled and services can be configured on demand. It is implemented through a dependency-aware controller and monitors. The controller first processes the information listened to by the monitors and then analyzes the connection between the metrics and the service requests through intelligent algorithms. Finally, the control traffic for regulating the control plane is generated. Our proposed mechanism is experimentally compared with the default strategy and existing work across a wide set of load scenarios in a testbed based on Istio service mesh and Kubernetes. Experimental results demonstrate that our mechanism can save the storage resources of a single agent by 40% to 60%, and the number of cluster updates can be greatly reduced. From the perspective of the whole cluster, the optimization results are even better. Lin Wang 0015, Xin Li 0005, Ning Wang 0018, Hao Li 0030, Xiaolin Qin, Jie Wu 0001 |
ICPADS | 4 |
| 2022 | Semantic Data Augmentation based Distance Metric Learning for Domain GeneralizationabstractDomain generalization (DG) aims to learn a model on one or more different but related source domains that could be generalized into an unseen target domain. Existing DG methods try to prompt the diversity of source domains for the model's generalization ability, while they may have to introduce auxiliary networks or striking computational costs. On the contrary, this work applies the implicit semantic augmentation in feature space to capture the diversity of source domains. Concretely, an additional loss function of distance metric learning (DML) is included to optimize the local geometry of data distribution. Besides, the logits from cross entropy loss with infinite augmentations is adopted as input features for the DML loss in lieu of the deep features. We also provide a theoretical analysis to show that the logits can approximate the distances defined on original features well. Further, we provide an in-depth analysis of the mechanism and rational behind our approach, which gives us a better understanding of why leverage logits in lieu of features can help domain generalization. The proposed DML loss with the implicit augmentation is incorporated into a recent DG method, that is, Fourier Augmented Co-Teacher framework (FACT). Meanwhile, our method also can be easily plugged into various DG methods. Extensive experiments on three benchmarks (Digits-DG, PACS and Office-Home) have demonstrated that the proposed method is able to achieve the state-of-the-art performance. Mengzhu Wang, Jianlong Yuan, Qi Qian 0001, Zhibin Wang 0004, Hao Li 0030 |
ACM Multimedia | 5 |
| 2022 | MimCo: Masked Image Modeling Pre-training with Contrastive TeacherabstractRecent masked image modeling (MIM) has received much attention in self-supervised learning (SSL), which requires the target model to recover the masked part of the input image. Although MIM-based pre-training methods achieve new state-of-the-art performance when transferred to many downstream tasks, the visualizations show that the learned representations are less separable, especially compared to those based on contrastive learning pre-training. This inspires us to think whether the linear separability of MIM pre-trained representation can be further improved, thereby improving the pre-training performance. Since MIM and contrastive learning tend to utilize different data augmentations and training strategies, combining these two pretext tasks is not trivial. In this work, we propose a novel and flexible pre-training framework, named MimCo, which combines MIM and contrastive learning through two-stage pre-training. Specifically, MimCo takes a pre-trained contrastive learning model as the teacher model and is pre-trained with two types of learning targets: patch-level and image-level reconstruction losses. Qiang Zhou 0001, Chaohui Yu, Hao Luo 0004, Zhibin Wang 0004, Hao Li 0030 |
ACM Multimedia | 5 |
| 2022 | Improved Fine-Tuning by Better Leveraging Pre-Training DataabstractAs a dominant paradigm, fine-tuning a pre-trained model on the target data is widely used in many deep learning applications, especially for small data sets. However, recent studies have empirically shown that training from scratch has the final performance that is no worse than this pre-training strategy once the number of training samples is increased in some vision tasks. In this work, we revisit this phenomenon from the perspective of generalization analysis by using excess risk bound which is popular in learning theory. The result reveals that the excess risk bound may have a weak dependency on the pre-trained model. The observation inspires us to leverage pre-training data for fine-tuning, since this data is also available for fine-tuning. The generalization result of using pre-training data shows that the excess risk bound on a target task can be improved when the appropriate pre-training data is included in fine-tuning. With the theoretical motivation, we propose a novel selection strategy to select a subset from pre-training data to help improve the generalization on the target task. Extensive experimental results for image classification tasks on 8 benchmark data sets verify the effectiveness of the proposed data selection based fine-tuning pipeline. Our code is available at https://github.com/ziquanliu/NeurIPS2022UOTfine_tuning. Ziquan Liu, Yi Xu 0008, Yuanhong Xu, Qi Qian 0001, Hao Li 0030, Xiangyang Ji, Antoni B. Chan, Rong Jin 0001 |
NeurIPS | 5 |
| 2022 | Entropy-Driven Mixed-Precision Quantization for Deep Network DesignabstractDeploying deep convolutional neural networks on Internet-of-Things (IoT) devices is challenging due to the limited computational resources, such as limited SRAM memory and Flash storage. Previous works re-design a small network for IoT devices, and then compress the network size by mixed-precision quantization. This two-stage procedure cannot optimize the architecture and the corresponding quantization jointly, leading to sub-optimal tiny deep models. In this work, we propose a one-stage solution that optimizes both jointly and automatically. The key idea of our approach is to cast the joint architecture design and quantization as an Entropy Maximization process. Particularly, our algorithm automatically designs a tiny deep model such that: 1) Its representation capacity measured by entropy is maximized under the given computational budget; 2) Each layer is assigned with a proper quantization precision; 3) The overall design loop can be done on CPU, and no GPU is required. More impressively, our method can directly search high-expressiveness architecture for IoT devices within less than half a CPU hour. Extensive experiments on three widely adopted benchmarks, ImageNet, VWW and WIDER FACE, demonstrate that our method can achieve the state-of-the-art performance in the tiny deep model regime. Code and pre-trained models are available at https://github.com/alibaba/lightweight-neural-architecture-search. Zhenhong Sun, Ce Ge, Junyan Wang 0001, Ming Lin 0002, Hesen Chen, Hao Li 0030, Xiuyu Sun |
NeurIPS | 6 |
| 2022 | VTC-LFC: Vision Transformer Compression with Low-Frequency ComponentsabstractAlthough Vision transformers (ViTs) have recently dominated many vision tasks, deploying ViT models on resource-limited devices remains a challenging problem. To address such a challenge, several methods have been proposed to compress ViTs. Most of them borrow experience in convolutional neural networks (CNNs) and mainly focus on the spatial domain. However, the compression only in the spatial domain suffers from a dramatic performance drop without fine-tuning and is not robust to noise, as the noise in the spatial domain can easily confuse the pruning criteria, leading to some parameters/channels being pruned incorrectly. Inspired by recent findings that self-attention is a low-pass filter and low-frequency signals/components are more informative to ViTs, this paper proposes compressing ViTs with low-frequency components. Two metrics named low-frequency sensitivity (LFS) and low-frequency energy (LFE) are proposed for better channel pruning and token pruning. Additionally, a bottom-up cascade pruning scheme is applied to compress different dimensions jointly. Extensive experiments demonstrate that the proposed method could save 40% ~ 60% of the FLOPs in ViTs, thus significantly increasing the throughput on practical devices with less than 1% performance drop on ImageNet-1K. Zhenyu Wang 0008, Hao Luo 0004, Pichao Wang, Fan Wang 0019, Hao Li 0030 |
NeurIPS | 6 |
| 2022 | Improved Knowledge Distillation via Full Kernel Matrix TransferabstractKnowledge distillation is an effective way for model compression in deep learning. Given a large model (i.e., teacher model), it aims to improve the performance of a compact model (i.e., student model) by transferring the information from the teacher. Various information for distillation has been studied. Recently, a number of works propose to transfer the pairwise similarity between examples to distill relative information. However, most of efforts are devoted to developing different similarity measurements, while only a small matrix consisting of examples within a mini-batch is transferred at each iteration that can be inefficient for optimizing the pairwise similarity over the whole data set. In this work, we aim to transfer the full similarity matrix effectively. The main challenge is from the size of the full matrix that is quadratic to the number of examples. To address the challenge, we decompose the original full matrix with Nyström method. By selecting appropriate landmark points, our theoretical analysis indicates that the loss for transfer can be further simplified. Concretely, we find that the difference between the original full kernel matrices between teacher and student can be well bounded by that of the corresponding partial matrices, which only consists of similarities between original examples and landmark points. Compared with the full matrix, the size of the partial matrix is linear in the number of examples, which improves the efficiency of optimization significantly. The empirical study on benchmark data sets demonstrates the effectiveness of the proposed algorithm. Qi Qian 0001, Hao Li 0030, Juhua Hu |
SDM | 2 |
| 2022 | Revisiting instance search: A new benchmark using cycle self-training
Yuqi Zhang 0001, Chong Liu 0002, Fan Wang 0019, Hao Li 0030, Xin Zhao 0012 |
Neurocomputing | 6 |
| 2022 | Refining pseudo labels for unsupervised Domain Adaptive Re-Identification
Shanshan Wang 0008, Lei Zhang 0038, Fan Wang 0019, Hao Li 0030 |
Knowl. Based Syst. | 5 |
| 2022 | Class-Aware Feature Aggregation Network for Video Object DetectionabstractRecent progress in video object detection (VOD) has shown that aggregating features from other frames to capture long-range contextual information is very important to deal with the challenges in VOD, such as partial occlusion, motion blur, etc. To exploit more effective feature aggregation, we propose several improvements over previous works in this paper: (1) a class-aware pixel-level feature aggregation module, which characterizes a pixel by exploiting the context information lying in the instances from both the current frame and other frames. Different from the previous non-local operation, the proposed class-aware pixel-level feature aggregation filters out most of the noisy information from the large scope of background and objects in different classes, and only enhances representation of a foreground pixel with the same class instances with limited ambiguous information; (2) a class-aware instance-level feature aggregation module, which aggregates features for object proposals by learning two kinds of relations: the temporal dependencies among the same class object proposals from support frames sampled in a long time range or even the whole sequence, and spatial topology relation among proposals of different objects in the target frame. The homogeneity constraint in instance-level feature aggregation filters out many defective proposals, making the feature aggregation more accurate; and (3) a correlation-based feature alignment module embedded in the instance-level feature aggregation, which aligns the feature maps of the support and target proposals. Without bells or whistles, the proposed method achieves state-of-the-art performance on the ImageNet VID dataset without any post-processing methods. This project is publicly availablehttps://github.com/LiangHann/Class-aware-Feature-Aggregation-Network-for-Video-Object-Detection. Pichao Wang, Zhaozheng Yin, Fan Wang 0019, Hao Li 0030 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Contrastive Haze-Aware Learning for Dynamic Remote Sensing Image DehazingabstractImage dehazing methods aim to recover a clear image from its hazy counterpart. While various dehazing methods have been proposed, their performance on real-world remote sensing (RS) images remains unsatisfying. A key reason is that the complex weather and imaging conditions (e.g., large fields of view) cause the haze condition to dramatically change in different images, while most existing methods fail to flexibly adapt their dehazing model to the specific haze condition in each image. To mitigate this problem, we present a contrastive haze-aware learning based dynamic dehazing method which demonstrates two aspects of advantage. On one hand, a contrastive clustering scheme is utilized to learn the image-wise haze representation using a set of real-world hazy images in an unsupervised manner, which enables identifying and discriminating the specific haze condition in each given hazy image. On the other hand, with the learned haze representation, a parameter generator can produce haze-aware parameters to dynamically construct a dehazing model for the given hazy image, which empowers us to adaptively dehaze the image based on its specific haze condition and thus improves the generalization ability. In addition, a new contrastive loss defined based on the learned haze representation is further utilized for model training and leads to better performance. To demonstrate the effectiveness of the proposed method, we evaluate it on two benchmark RS image datasets including various real-world hazy images, and observe obviously superiority over other state-of-the-art competitors. Jiangtao Nie, Wei Wei 0008, Lei Zhang 0054, Jianlong Yuan, Zhibin Wang 0004, Hao Li 0030 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Multi-View Evolutionary Training for Unsupervised Domain Adaptive Person Re-IdentificationabstractClustering-based approaches have been successfully applied to unsupervised domain adaptation (UDA) tasks for person re-identification (Re-ID), where no annotations are provided in target domain. However, the clustering process is sensitive to noises, leading to imperfect pseudo labels that could damage the training performance. In this work, we propose a Multi-view Evolutionary Training (MET) method to effectively reduce noises in clustering results from two dimensions. First, to improve the clustering accuracy at each time frame (i.e. snapshot quality), a Multi-view Diffusion (MvD) module is proposed. Through capturing data relationships from multiple viewpoints and aggregating their information, noises and bias from each individual viewpoint can be eliminated, and more reliable similarity matrix can be produced for clustering. Second, to improve the temporal consistency between clustering at different iterations, i.e. temporal consistency, we propose an Evolutionary Local Refinement (ELR) module, which utilizes the previous clustering results to guide and improve current results, and further make the training process more stable and robust. Extensive experiments demonstrate that our method can provide clustering results with high quality, and achieve state-of-the-art performance on UDA Re-ID. Jianyang Gu, Hao Luo 0004, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Cross-Batch Hard Example Mining With Pseudo Large Batch for ID vs. Spot Face RecognitionabstractIn our daily life, a large number of activities require identity verification, e.g., ePassport gates. Most of those verification systems recognize who you are by matching the ID document photo (ID face) to your live face image (spot face). The ID vs. Spot (IvS) face recognition is different from general face recognition where each dataset usually contains a small number of subjects and sufficient images for each subject. In IvS face recognition, the datasets usually contain massive class numbers (million or more) while each class only has two image samples (one ID face and one spot face), which makes it very challenging to train an effective model (e.g., excessive demand on GPU memory if conducting the classification on such massive classes, hardly capture the effective features for bisample data of each identity, etc.). To avoid the excessive demand on GPU memory, a two-stage training method is developed, where we first train the model on the dataset in general face recognition (e.g., MS-Celeb-1M) and then employ the metric learning losses (e.g., triplet and quadruplet losses) to learn the features on IvS data with million classes. To extract more effective features for IvS face recognition, we propose two novel algorithms to enhance the network by selecting harder samples for training. Firstly, a Cross-Batch Hard Example Mining (CB-HEM) is proposed to select the hard triplets from not only the current mini-batch but also past dozens of mini-batches (for convenience, we use batch to denote a mini-batch in the following), which can significantly expand the space of sample selection. Secondly, a Pseudo Large Batch (PLB) is proposed to virtually increase the batch size with a fixed GPU memory. The proposed PLB and CB-HEM can be employed simultaneously to train the network, which dramatically expands the selecting space by hundreds of times, where the very hard sample pairs especially the hard negative pairs can be selected for training to enhance the discriminative capability. Extensive comparative evaluations conducted on multiple IvS benchmarks demonstrate the effectiveness of the proposed method. Zichang Tan, Ajian Liu 0001, Jun Wan 0001, Hao Li 0030, Zhen Lei 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Image Process. | 4 |
| 2021 | Instant-Teaching: An End-to-End Semi-Supervised Object Detection FrameworkabstractSupervised learning based object detection frameworks demand plenty of laborious manual annotations, which may not be practical in real applications. Semi-supervised object detection (SSOD) can effectively leverage unlabeled data to improve the model performance, which is of great significance for the application of object detection models. In this paper, we revisit SSOD and propose Instant-Teaching, a completely end-to-end and effective SSOD framework, which uses instant pseudo labeling with extended weak-strong data augmentations for teaching during each training iteration. To alleviate the confirmation bias problem and improve the quality of pseudo annotations, we further propose a co-rectify scheme based on Instant-Teaching, denoted as Instant-Teaching∗. Extensive experiments on both MS-COCO and PASCAL VOC datasets substantiate the superiority of our framework. Specifically, our method surpasses state-of-the-art methods by 4.2 mAP on MS-COCO when using 2% labeled data. Even with full supervised information of MS-COCO, the proposed method still outperforms state-of-the-art methods by about 1.0 mAP. On PASCAL VOC, we can achieve more than 5 mAP improvement by applying VOC07 as labeled data and VOC12 as unlabeled data. Qiang Zhou 0001, Chaohui Yu, Zhibin Wang 0004, Qi Qian 0001, Hao Li 0030 |
CVPR | 5 |
| 2021 | TransReID: Transformer-based Object Re-IdentificationabstractExtracting robust feature representation is one of the key challenges in object re-identification (ReID). Although convolution neural network (CNN)-based methods have achieved great success, they only process one local neighborhood at a time and suffer from information loss on details caused by convolution and downsampling operators (e.g. pooling and strided convolution). To overcome these limitations, we propose a pure transformer-based object ReID framework named TransReID. Specifically, we first encode an image as a sequence of patches and build a transformer-based strong baseline with a few critical improvements, which achieves competitive results on several ReID benchmarks with CNN-based methods. To further enhance the robust feature learning in the context of transformers, two novel modules are carefully designed. (i) The jigsaw patch module (JPM) is proposed to rearrange the patch embeddings via shift and patch shuffle operations which generates robust features with improved discrimination ability and more diversified coverage. (ii) The side information embeddings (SIE) is introduced to mitigate feature bias towards camera/view variations by plugging in learnable embeddings to incorporate these non-visual clues. To the best of our knowledge, this is the first work to adopt a pure transformer for ReID research. Experimental results of TransReID are superior promising, which achieve state-of-the-art performance on both person and vehicle ReID benchmarks. Code is available at https://github.com/heshuting555/TransReID. Shuting He, Hao Luo 0004, Pichao Wang, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009 |
ICCV | 5 |
| 2021 | Zen-NAS: A Zero-Shot NAS for High-Performance Image RecognitionabstractAccuracy predictor is a key component in Neural Architecture Search (NAS) for ranking architectures. Building a high-quality accuracy predictor usually costs enormous computation. To address this issue, instead of using an accuracy predictor, we propose a novel zero-shot index dubbed Zen-Score to rank the architectures. The Zen-Score represents the network expressivity and positively correlates with the model accuracy. The calculation of Zen-Score only takes a few forward inferences through a randomly initialized network, without training network parameters. Built upon the Zen-Score, we further propose a new NAS algorithm, termed as Zen-NAS, by maximizing the Zen-Score of the target network under given inference budgets. Within less than half GPU day, Zen-NAS is able to directly search high performance architectures in a data-free style. Comparing with previous NAS methods, the proposed Zen-NAS is magnitude times faster on multiple server-side and mobile-side GPU platforms with state-of-the-art accuracy on ImageNet. Searching and training code as well as pre-trained models are available from https://github.com/idstcv/ZenNAS. Ming Lin 0002, Pichao Wang, Zhenhong Sun, Hesen Chen, Xiuyu Sun, Qi Qian 0001, Hao Li 0030, Rong Jin 0001 |
ICCV | 7 |
| 2021 | Weakly Supervised Representation Learning with Coarse LabelsabstractWith the development of computational power and techniques for data collection, deep learning demonstrates a superior performance over most existing algorithms on visual benchmark data sets. Many efforts have been devoted to studying the mechanism of deep learning. One important observation is that deep learning can learn the discriminative patterns from raw materials directly in a task-dependent manner. Therefore, the representations obtained by deep learning outperform hand-crafted features significantly. However, for some real-world applications, it is too expensive to collect the task-specific labels, such as visual search in online shopping. Compared to the limited availability of these task-specific labels, their coarse-class labels are much more affordable, but representations learned from them can be suboptimal for the target task. To mitigate this challenge, we propose an algorithm to learn the fine-grained patterns for the target task, when only its coarse-class labels are available. More importantly, we provide a theoretical guarantee for this. Extensive experiments on real-world data sets demonstrate that the proposed method can significantly improve the performance of learned representations on the target task, when only coarse-class information is available for training. Yuanhong Xu, Qi Qian 0001, Hao Li 0030, Rong Jin 0001, Juhua Hu |
ICCV | 3 |
| 2021 | Digging into Uncertainty in Self-supervised Multi-view StereoabstractSelf-supervised Multi-view stereo (MVS) with a pretext task of image reconstruction has achieved significant progress recently. However, previous methods are built upon intuitions, lacking comprehensive explanations about the effectiveness of the pretext task in self-supervised MVS. To this end, we propose to estimate epistemic uncertainty in self-supervised MVS, accounting for what the model ignores. Specially, the limitations can be categorized into two types: ambiguious supervision in foreground and invalid supervision in background. To address these issues, we propose a novel Uncertainty reduction Multi-view Stereo (U-MVS) framework for self-supervised learning. To alleviate ambiguous supervision in foreground, we involve extra correspondence prior with a flow-depth consistency loss. The dense 2D correspondence of optical flows is used to regularize the 3D stereo correspondence in MVS. To handle the invalid supervision in background, we use Monte-Carlo Dropout to acquire the uncertainty map and further filter the unreliable supervision signals on invalid regions. Extensive experiments on DTU and Tank&Temples benchmark show that our U-MVS framework1achieves the best performance among unsupervised MVS methods, with competitive performance with its supervised opponents. Yali Wang 0001, Wenxiong Kang, Baigui Sun, Hao Li 0030, Yu Qiao 0001 |
ICCV | 6 |
| 2021 | A Simple Baseline for Semi-supervised Semantic Segmentation with Strong Data Augmentation*abstractRecently, significant progress has been made on semantic segmentation. However, the success of supervised semantic segmentation typically relies on a large amount of labeled data, which is time-consuming and costly to obtain. Inspired by the success of semi-supervised learning methods for image classification, here we propose a simple yet effective semi-supervised learning framework for semantic segmentation. We demonstrate that the devil is in the details: a set of simple designs and training techniques can collectively improve the performance of semi-supervised semantic segmentation significantly. Previous works [3], [25] fail to effectively employ strong augmentation in pseudo-label learning, as the large distribution disparity caused by strong augmentation harms the batch nor-malization statistics. We design a new batch normalization, namely distribution-specific batch normalization (DSBN) to address this problem and show the importance of strong augmentation for semantic segmentation. Moreover, we design a self-correction loss, which is effective in terms of noise resistance. We conduct a series of ablation studies to show the effectiveness of each component. Our method achieves state-of-the-art results in the semi-supervised settings on the Cityscapes and Pascal VOC datasets. Jianlong Yuan, Yifan Liu 0001, Chunhua Shen, Zhibin Wang 0004, Hao Li 0030 |
ICCV | 5 |
| 2021 | Learning Accurate Entropy Model with Global Reference for Image Compression
Yichen Qian, Zhiyu Tan, Xiuyu Sun, Ming Lin 0002, Zhenhong Sun, Hao Li 0030, Rong Jin 0001 |
ICLR | 7 |
| 2021 | Dash: Semi-Supervised Learning with Dynamic ThresholdingabstractWhile semi-supervised learning (SSL) has received tremendous attentions in many machine learning tasks due to its successful use of unlabeled data, existing SSL algorithms use either all unlabeled examples or the unlabeled examples with a fixed high-confidence prediction during the training progress. However, it is possible that too many correct/wrong pseudo labeled examples are eliminated/selected. In this work we develop a simple yet powerful framework, whose key idea is to select a subset of training examples from the unlabeled data when performing existing SSL methods so that only the unlabeled examples with pseudo labels related to the labeled data will be used to train models. The selection is performed at each updating iteration by only keeping the examples whose losses are smaller than a given threshold that is dynamically adjusted through the iteration. Our proposed approach, Dash, enjoys its adaptivity in terms of unlabeled data selection and its theoretical guarantee. Specifically, we theoretically establish the convergence rate of Dash from the view of non-convex optimization. Finally, we empirically demonstrate the effectiveness of the proposed method in comparison with state-of-the-art over benchmarks. Yi Xu 0008, Jinxing Ye, Qi Qian 0001, Yufeng Li 0008, Baigui Sun, Hao Li 0030, Rong Jin 0001 |
ICML | 7 |
| 2021 | Exploring the Quality of GAN Generated Images for Person Re-IdentificationabstractRecently, GAN based method has demonstrated strong effectiveness in generating augmentation data for person re-identification (ReID), on account of its ability to bridge the gap between domains and enrich the data variety in feature space. However, most of the ReID works pick all the GAN generated data as additional training samples or evaluate the quality of GAN generation at the entire data set level, ignoring the image-level essential feature of data in ReID task. In this paper, we analyze the in-depth characteristics of ReID sample and solve the problem of "What makes a GAN-generated image good for ReID''. Specifically, we propose to examine each data sample with id-consistency and diversity constraints by mapping image onto different spaces. With a metric-based sampling method, we demonstrate that not every GAN-generated data is beneficial for augmentation. Models trained with data filtered by our quality evaluation outperform those trained with the full augmentation set by a large margin. Extensive experiments show the effectiveness of our method on both supervised ReID task and unsupervised domain adaptation ReID task. Yiqi Jiang, Xiuyu Sun, Fan Wang 0019, Hao Li 0030 |
ACM Multimedia | 6 |
| 2021 | Interpolation Variable Rate Image CompressionabstractCompression standards have been used to reduce the cost of image storage and transmission for decades. In recent years, learned image compression methods have been proposed and achieved compelling performance to the traditional standards. However, in these methods, a set of different networks are used for various compression rates, resulting in a high cost in model storage and training. Although some variable-rate approaches have been proposed to reduce the cost by using a single network, most of them brought some performance degradation when applying fine rate control. To enable variable-rate control without sacrificing the performance, we propose an efficient Interpolation Variable-Rate (IVR) network, by introducing a handy Interpolation Channel Attention (InterpCA) module in the compression network. With the use of two hyperparameters for rate control and linear interpolation, the InterpCA achieves a fine PSNR interval of 0.001 dB and a fine rate interval of 0.0001 Bits-Per-Pixel (BPP) with 9000 rates in the IVR network. Experimental results demonstrate that the IVR network is the first variable-rate learned method that outperforms VTM 9.0 (intra) in PSNR and Multiscale Structural Similarity (MS-SSIM). Zhenhong Sun, Zhiyu Tan, Xiuyu Sun, Fangyi Zhang, Yichen Qian, Hao Li 0030 |
ACM Multimedia | 7 |
| 2021 | HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot LearningabstractZero-shot learning (ZSL) tackles the unseen class recognition problem, transferring semantic knowledge from seen classes to unseen ones. Typically, to guarantee desirable knowledge transfer, a common (latent) space is adopted for associating the visual and semantic domains in ZSL. However, existing common space learning methods align the semantic and visual domains by merely mitigating distribution disagreement through one-step adaptation. This strategy is usually ineffective due to the heterogeneous nature of the feature representations in the two domains, which intrinsically contain both distribution and structure variations. To address this and advance ZSL, we propose a novel hierarchical semantic-visual adaptation (HSVA) framework. Specifically, HSVA aligns the semantic and visual domains by adopting a hierarchical two-step adaptation, i.e., structure adaptation and distribution adaptation. In the structure adaptation step, we take two task-specific encoders to encode the source data (visual domain) and the target data (semantic domain) into a structure-aligned common space. To this end, a supervised adversarial discrepancy (SAD) module is proposed to adversarially minimize the discrepancy between the predictions of two task-specific classifiers, thus making the visual and semantic feature manifolds more closely aligned. In the distribution adaptation step, we directly minimize the Wasserstein distance between the latent multivariate Gaussian distributions to align the visual and semantic distributions using a common encoder. Finally, the structure and distribution adaptation are derived in a unified framework under two partially-aligned variational autoencoders. Extensive experiments on four benchmark datasets demonstrate that HSVA achieves superior performance on both conventional and generalized ZSL. The code is available at \url{https://github.com/shiming-chen/HSVA}. Shiming Chen 0002, Guosen Xie, Yang Liu 0069, Qinmu Peng, Baigui Sun, Hao Li 0030, Xinge You, Ling Shao 0001 |
NeurIPS | 6 |
| 2021 | Context and Structure Mining Network for Video Object Detection
Pichao Wang, Zhaozheng Yin, Fan Wang 0019, Hao Li 0030 |
Int. J. Comput. Vis. | 5 |
| 2020 | DR Loss: Improving Object Detection by Distributional RankingabstractMost of object detection algorithms can be categorized into two classes: two-stage detectors and one-stage detectors. Recently, many efforts have been devoted to one-stage detectors for the simple yet effective architecture. Different from two-stage detectors, one-stage detectors aim to identify foreground objects from all candidates in a single stage. This architecture is efficient but can suffer from the imbalance issue with respect to two aspects: the inter-class imbalance between the number of candidates from foreground and background classes and the intra-class imbalance in the hardness of background candidates, where only a few candidates are hard to be identified. In this work, we propose a novel distributional ranking (DR) loss to handle the challenge. For each image, we convert the classification problem to a ranking problem, which considers pairs of candidates within the image, to address the inter-class imbalance problem. Then, we push the distributions of confidence scores for foreground and background towards the decision boundary. After that, we optimize the rank of the expectations of derived distributions in lieu of original pairs. Our method not only mitigates the intra-class imbalance issue in background candidates but also improves the efficiency for the ranking algorithm. By merely replacing the focal loss in RetinaNet with the developed DR loss and applying ResNet-101 as the backbone, mAP of the single-scale test on COCO can be improved from 39.1% to 41.7% without bells and whistles, which demonstrates the effectiveness of the proposed loss function. Qi Qian 0001, Hao Li 0030, Rong Jin 0001 |
CVPR | 3 |
| 2020 | Hierarchically Robust Representation LearningabstractWith the tremendous success of deep learning in visual tasks, the representations extracted from intermediate layers of learned models, that is, deep features, attract much attention of researchers. Previous empirical analysis shows that those features can contain appropriate semantic information. Therefore, with a model trained on a large-scale benchmark data set (e.g., ImageNet), the extracted features can work well on other tasks. In this work, we investigate this phenomenon and demonstrate that deep features can be suboptimal due to the fact that they are learned by minimizing the empirical risk. When the data distribution of the target task is different from that of the benchmark data set, the performance of deep features can degrade. Hence, we propose a hierarchically robust optimization method to learn more generic features. Considering the example-level and concept-level robustness simultaneously, we formulate the problem as a distributionally robust optimization problem with Wasserstein ambiguity set constraints, and an efficient algorithm with the conventional training pipeline is proposed. Experiments on benchmark data sets demonstrate the effectiveness of the robust deep representations. Qi Qian 0001, Juhua Hu, Hao Li 0030 |
CVPR | 3 |
| 2020 | Unsupervised Style Transfer via Dualgan for Cross-Domain Aerial Image ClassificationabstractDue to its wide applications, aerial image classification, which is also called semantic segmentation of aerial imagery, attracts increasing research interest in recent years. Until now, deep semantic segmentation network (DSSN) has been widely adopted to address aerial image classification and achieves tremendous success. However, the superior performance of DSSN highly depends on massive targeted data with labels. When DSSN is trained on data from the source domain but tested on data from the target domain, the performance of DSSN is often very limited due to the data shift between source and target domains. To alleviate the disadvantage influence of data shift, this paper proposes a domain adaptation approach via unsupervised style transfer to cope with cross-domain aerial image classification. More specifically, this paper innovatively recommends DualGAN to conduct unsupervised style transfer for mapping aerial images in the source domain to the target domain. The mapped aerial imagery with labels is adopted to train DSSN, which is further used to classify aerial imagery in the target domain. To verify the validity of the presented approach, we give two cross-domain experimental settings including: (I) variation of geographic location; (II) variation of both geographic location and imaging mode. Extensive experiments under two typical cross-domain settings show that our proposed method can obviously outperform the state-of-the-art methods. Yansheng Li 0001, Te Shi 0001, Wei Chen 0089, Yongjun Zhang 0002, Zhibin Wang 0004, Hao Li 0030 |
IGARSS | 6 |
| 2020 | Exploiting Better Feature Aggregation for Video Object DetectionabstractVideo object detection (VOD) has been a rising topic in recent years due to the challenges such as occlusion, motion blur, etc. To deal with these challenges, feature aggregation from local or global support frames is verified effective. To exploit better feature aggregation, in this paper, we propose two improvements over previous works: a class-constrained spatial-temporal relation network and a correlation-based feature alignment module. For the class constrained spatial-temporal relation network, it operates on object region proposals, and learns two kinds of relations: (1) the dependencies among region proposals of the same object class from support frames sampled in a long time range or even the whole sequence, and (2) spatial relations among proposals of different objects in the target frame. The homogeneity constraint in spatial-temporal relation network not only filters out many defective proposals but also implicitly embeds the traditional post-processing strategies (e.g., Seq-NMS), leading to a unified end-to-end training networks. In the feature alignment module, we propose a correlation based feature alignment method to align the support and target frames for feature aggregation in the temporal domain. Our experiments show that the proposed method improves the accuracy of single-frame detectors significantly, and outperforms previous temporal or spatial relation networks. Without bells or whistles, the proposed method achieves state-of-the-art performance on the ImageNet VID dataset (84.80% with ResNet-101) without any post-processing methods. Pichao Wang, Zhaozheng Yin, Fan Wang 0019, Hao Li 0030 |
ACM Multimedia | 5 |
| 2019 | Robust Optimization over Multiple DomainsabstractIn this work, we study the problem of learning a single model for multiple domains. Unlike the conventional machine learning scenario where each domain can have the corresponding model, multiple domains (i.e., applications/users) may share the same machine learning model due to maintenance loads in cloud computing services. For example, a digit-recognition model should be applicable to hand-written digits, house numbers, car plates, etc. Therefore, an ideal model for cloud computing has to perform well at each applicable domain. To address this new challenge from cloud computing, we develop a framework of robust optimization over multiple domains. In lieu of minimizing the empirical risk, we aim to learn a model optimized to the adversarial distribution over multiple domains. Hence, we propose to learn the model and the adversarial distribution simultaneously with the stochastic algorithm for efficiency. Theoretically, we analyze the convergence rate for convex and non-convex models. To our best knowledge, we first study the convergence rate of learning a robust non-convex model with a practical algorithm. Furthermore, we demonstrate that the robustness of the framework and the convergence rate can be further enhanced by appropriate regularizers over the adversarial distribution. The empirical study on real-world fine-grained visual categorization and digits recognition tasks verifies the effectiveness and efficiency of the proposed framework. Qi Qian 0001, Shenghuo Zhu, Jiasheng Tang, Rong Jin 0001, Baigui Sun, Hao Li 0030 |
AAAI | 6 |
| 2019 | SoftTriple Loss: Deep Metric Learning Without Triplet SamplingabstractDistance metric learning (DML) is to learn the embeddings where examples from the same class are closer than examples from different classes. It can be cast as an optimization problem with triplet constraints. Due to the vast number of triplet constraints, a sampling strategy is essential for DML. With the tremendous success of deep learning in classifications, it has been applied for DML. When learning embeddings with deep neural networks (DNNs), only a mini-batch of data is available at each iteration. The set of triplet constraints has to be sampled within the mini-batch. Since a mini-batch cannot capture the neighbors in the original set well, it makes the learned embeddings sub-optimal. On the contrary, optimizing SoftMax loss, which is a classification loss, with DNN shows a superior performance in certain DML tasks. It inspires us to investigate the formulation of SoftMax. Our analysis shows that SoftMax loss is equivalent to a smoothed triplet loss where each class has a single center. In real-world data, one class can contain several local clusters rather than a single one, e.g., birds of different poses. Therefore, we propose the SoftTriple loss to extend the SoftMax loss with multiple centers for each class. Compared with conventional deep metric learning algorithms, optimizing SoftTriple loss can learn the embeddings without the sampling phase by mildly increasing the size of the last fully connected layer. Experiments on the benchmark fine-grained data sets demonstrate the effectiveness of the proposed loss function. Qi Qian 0001, Baigui Sun, Juhua Hu, Tacoma Tacoma, Hao Li 0030, Rong Jin 0001 |
ICCV | 6 |
| 2019 | Learning to Rank Proposals for Object DetectionabstractNon-Maximum Suppression (NMS) is an essential step of modern object detection models for removing duplicated candidates. The efficacy of NMS heavily affects the final detection results. Prior works exploit suppression criterions relying on either the objectiveness derived from classification or the overlapness produced by regression, both of which are heuristically designed and fail to explicitly link with the suppression rank. To address this issue, in this paper, we propose a novel Learning-to-Rank (LTR) model to produce the suppression rank via a learning procedure, thus facilitating the candidate generation and lifting the detection performance. In particular, we define a ranking score based on IoU to indicate the ranks of candidates during the NMS step, where candidates with high ranking score will be reserved and the ones with low ranking score will be eliminated. We design a lightweight network to predict the ranking score. We introduce a ranking loss to supervise the generation of these ranking scores, which encourages candidates with IoU to the ground-truth to rank higher. To facilitate the training procedure, we design a novel sampling strategy via dividing candidates into different levels and select hard pairs to adopt in the training. During the inference phase, this module can be exploited as a plugin to the current object detector. The training and inference of the overall framework is end-to-end. Comprehensive experiments on benchmarks PASCAL VOC and MS COCO demonstrate the generality and effectiveness of our model for facilitating existing object detectors to state-of-the-art accuracy. Zhiyu Tan, Xuecheng Nie, Qi Qian 0001, Nan Li 0019, Hao Li 0030 |
ICCV | 5 |
| 2019 | Robust Gaussian Process Regression for Real-Time High Precision GPS Signal EnhancementabstractSatellite-based positioning system such as GPS often suffers from large amount of noise that degrades the positioning accuracy dramatically especially in real-time applications. In this work, we consider a data-mining approach to enhance the GPS signal. We build a large-scale high precision GPS receiver grid system to collect real-time GPS signals for training. The Gaussian Process (GP) regression is chosen to model the vertical Total Electron Content (vTEC) distribution of the ionosphere of the Earth. Our experiments show that the noise in the real-time GPS signals often exceeds the breakdown point of the conventional robust regression methods resulting in sub-optimal system performance. We propose a three-step approach to address this challenge. In the first step we perform a set of signal validity tests to separate the signals into clean and dirty groups. In the second step, we train an initial model on the clean signals and then reweigting the dirty signals based on the residual error. A final model is retrained on both the clean signals and the reweighted dirty signals. In the theoretical analysis, we prove that the proposed three-step approach is able to tolerate much higher noise level than the vanilla robust regression methods if two reweighting rules are followed. We validate the superiority of the proposed method in our real-time high precision positioning system against several popular state-of-the-art robust regression methods. Our method achieves centimeter positioning accuracy in the benchmark region with probability $78.4%$ , outperforming the second best baseline method by a margin of $8.3%$. The benchmark takes 6 hours on 20,000 CPU cores or 14 years on a single CPU. Ming Lin 0002, Xiaomin Song, Qi Qian 0001, Hao Li 0030, Liang Sun 0001, Shenghuo Zhu, Rong Jin 0001 |
KDD | 4 |
| 2018 | Extremely Low Bit Neural Network: Squeeze the Last Bit Out With ADMMabstractAlthough deep learning models are highly effective for various learning tasks, their high computational costs prohibit the deployment to scenarios where either memory or computational resources are limited. In this paper, we focus on compressing and accelerating deep models with network weights represented by very small numbers of bits, referred to as extremely low bit neural network. We model this problem as a discretely constrained optimization problem. Borrowing the idea from Alternating Direction Method of Multipliers (ADMM), we decouple the continuous parameters from the discrete constraints of network, and cast the original hard problem into several subproblems. We propose to solve these subproblems using extragradient and iterative quantization algorithms that lead to considerably faster convergency compared to conventional optimization methods. Extensive experiments on image recognition and object detection verify that the proposed algorithm is more effective than state-of-the-art approaches when coming to extremely low bit neural network. Cong Leng, Zesheng Dou, Hao Li 0030, Shenghuo Zhu, Rong Jin 0001 |
AAAI | 3 |
| 2018 | Large-Scale Distance Metric Learning With UncertaintyabstractDistance metric learning (DML) has been studied extensively in the past decades for its superior performance with distance-based algorithms. Most of the existing methods propose to learn a distance metric with pairwise or triplet constraints. However, the number of constraints is quadratic or even cubic in the number of the original examples, which makes it challenging for DML to handle the large-scale data set. Besides, the real-world data may contain various uncertainty, especially for the image data. The uncertainty can mislead the learning procedure and cause the performance degradation. By investigating the image data, we find that the original data can be observed from a small set of clean latent examples with different distortions. In this work, we propose the margin preserving metric learning framework to learn the distance metric and latent examples simultaneously. By leveraging the ideal properties of latent examples, the training efficiency can be improved significantly while the learned metric also becomes robust to the uncertainty in the original data. Furthermore, we can show that the metric is learned from latent examples only, but it can preserve the large margin property even for the original data. The empirical study on the benchmark image data sets demonstrates the efficacy and efficiency of the proposed method. Qi Qian 0001, Jiasheng Tang, Hao Li 0030, Shenghuo Zhu, Rong Jin 0001 |
CVPR | 3 |
| 2017 | The Opensesame NIST 2016 Speaker Recognition Evaluation System
Qi Qian 0001, Zhibin Wang 0004, Qingen Zhao, Tianzhou Wang, Hao Li 0030, Shenghuo Zhu, Rong Jin 0001, Tuo Zhao |
INTERSPEECH | 6 |
| 2016 | Deep CTR Prediction in Display AdvertisingabstractClick through rate (CTR) prediction of image ads is the core task of online display advertising systems, and logistic regression (LR) has been frequently applied as the prediction model. However, LR model lacks the ability of extracting complex and intrinsic nonlinear features from handcrafted high-dimensional image features, which limits its effectiveness. To solve this issue, in this paper, we introduce a novel deep neural network (DNN) based model that directly predicts the CTR of an image ad based on raw image pixels and other basic features in one step. The DNN model employs convolution layers to automatically extract representative visual features from images, and nonlinear CTR features are then learned from visual features and other contextual features by using fully-connected layers. Empirical evaluations on a real world dataset with over 50 million records demonstrate the effectiveness and efficiency of this method. Junxuan Chen, Baigui Sun, Hao Li 0030, Hongtao Lu 0001, Xian-Sheng Hua 0001 |
ACM Multimedia | 3 |
| 2012 | Multimodal Graph-Based Reranking for Web Image SearchabstractThis paper introduces a web image search reranking approach that explores multiple modalities in a graph-based learning scheme. Different from the conventional methods that usually adopt a single modality or integrate multiple modalities into a long feature vector, our approach can effectively integrate the learning of relevance scores, weights of modalities, and the distance metric and its scaling for each modality into a unified scheme. In this way, the effects of different modalities can be adaptively modulated and better reranking performance can be achieved. We conduct experiments on a large dataset that contains more than 1000 queries and 1 million images to evaluate our approach. Experimental results demonstrate that the proposed reranking approach is more robust than using each individual modality, and it also performs better than many existing methods. Meng Wang 0001, Hao Li 0030, Dacheng Tao, Ke Lu 0002, Xindong Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2011 | Optimizing multimodal reranking for web image searchabstractIn this poster, we introduce a web image search reranking approach with exploring multiple modalities. Diff erent from the conventional methods that build graph with one feature set for reranking, our approach integrates multiple feature sets that describe visual content from different aspects. We simultaneously integrate the learning of relevance scores, the weighting of different feature sets, the distance metric and the scaling for each feature set into a unified scheme. Experimental results on a large data set that contains more than 1,100 queries and 1 million images demonstrate the effectiveness of our approach. Hao Li 0030, Meng Wang 0001, Zhisheng Li, Zhengjun Zha, Jialie Shen 0001 |
SIGIR | 1 |
| 2007 | New timing and routability driven placement algorithms for FPGA synthesisabstractWe present new timing and congestion driven FPGA placement algorithms with minimal runtime overhead. By predicting the post-routing critical edges and estimating congestion accurately, our algorithms simultaneously reduce the critical path delay and the minimum number of routing tracks. The core of our algorithm consists of a criticality history record of connection edges and a congestion map. This approach is applied to the 20 largest MCNC benchmark circuits. Experimental results show that compared with VPR [1], our algorithms yield an average of 8.1% reduction (maximum 30.5%) in the critical path delay and 5% reduction in channel width. Meanwhile, the average runtime of our algorithms is only 2.3X as of VPR's. Hao Li 0030, Qiang Zhou 0001, Yici Cai, Xianlong Hong |
ACM Great Lakes Symposium on VLSI | 2 |