Hanjiang Lai

dblp:31/9937 · DBLP profile ↗
← Back
79ranked-venue papers
11as first author
42since 2021 · last 2026
0000-0001-8057-6744ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 7 first-author · 21 since 2021Artificial intelligence and machine learning · 40 · 4 first-author · 18 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 Inductive Controlled Generation Based on Adaptive Templates for Answering Subjective Product Questions
Yian Yao, Jianxing Yu, Huaijie Zhu, Hanjiang Lai, Wei Liu 0061, Yanghui Rao, Jian Yin 0001
DASFAA (6)4
2026 GAFD-CC: Global-aware feature decoupling with confidence calibration for out-of-distribution detection
Yongheng Xu, Jianxing Yu, Yan Pan 0002, Jian Yin 0001, Hanjiang Lai
Neural Networks6
2026 Flow-Matching Posterior Sampling: A Training-Free Conditional Generation for Flow Matching
abstract
Training-free conditional generation based on flow matching aims to leverage pre-trained unconditional flow matching models to perform conditional generation without retraining. Recently, a successful training-free conditional generation approach incorporates conditions via posterior sampling, which relies on the availability of a score function in the unconditional diffusion model. However, flow matching models lack an explicit score function, rendering this strategy inapplicable. Approximate posterior sampling for flow matching has been explored, but it is limited to linear inverse problems. In this paper, we propose Flow Matching-based Posterior Sampling (FMPS) to broaden its scope of application. We introduce a correction term by steering the velocity field. This correction term can be reformulated to incorporate a surrogate score function, thereby bridging the gap between flow matching models and score-based posterior sampling. Hence, FMPS enables posterior sampling to be adjusted within the flow-matching framework. Furthermore, we propose two practical implementations of the correction mechanism: one to improve generation quality and the other to enhance computational efficiency. Experimental results on diverse conditional generation tasks demonstrate that our method achieves superior generation quality compared to existing state-of-the-art approaches, validating the effectiveness and generality of FMPS.
Kaiyu Song, Hanjiang Lai, Yan Pan 0002, Kun Yue, Jian Yin 0001
IEEE Trans. Image Process.2
2026 Enhancing Zero-Shot Adversarial Robustness of Vision-Language Models With Training-Free Adaptive Feature Movement
abstract
Pre-trained Vision-Language Models (VLMs) have demonstrated strong zero-shot generalization capabilities. Despite their effectiveness on various downstream tasks, they remain vulnerable to adversarial samples. Existing methods fine-tune VLMs to improve their robust performance by performing adversarial training on a certain dataset. However, this can lead to model overfitting and is not a true zero-shot scenario. In this paper, we propose a truly zero-shot and training-free approach that can improve the zero-shot adversarial robustness of VLMs on the evaluated benchmarks. Specifically, we first discover that simply adding Gaussian noise can enhance the VLM's zero-shot robustness. Then, we treat the adversarial examples with added Gaussian noise as anchors and strive to find a path in the embedding space that leads from the adversarial examples to the cleaner samples. Furthermore, to avoid the overfitting issue caused by fixed hyperparameters, we propose an adaptive parameter adjustment method based on the distance between the anchors and adversarial samples in the embedding space. We largely preserve the original VLMs' zero-shot generalization abilities in a truly zero-shot and training-free manner on the evaluated benchmarks compared to previous methods. Extensive experiments on 16 datasets demonstrate that our method can achieve stronger zero-shot robust performance, improving the top-1 robust accuracy by an average of 10.83%.
Baoshun Tong, Hanjiang Lai, Yan Pan 0002, Jian Yin 0001, Liang Lin 0004
IEEE Trans. Image Process.2
2025 Answering Complex Geographic Questions by Adaptive Reasoning with Visual Context and External Commonsense Knowledge
abstract
Fan Li, Jianxing Yu, Jielong Tang, Wenqing Chen, Hanjiang Lai, Yanghui Rao, Jian Yin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jianxing Yu, Jielong Tang, Wenqing Chen, Hanjiang Lai, Yanghui Rao, Jian Yin 0001
ACL (1)5
2025 Detecting Emotional Incongruity of Sarcasm by Commonsense Reasoning
abstract
This paper focuses on sarcasm detection, which aims to identify whether given statements convey criticism, mockery, or other negative sentiment opposite to the literal meaning. To detect sarcasm, humans often require a comprehensive understanding of the semantics in the statement and even resort to external commonsense to infer the fine-grained incongruity. However, existing methods lack commonsense inferential ability when they face complex real-world scenarios, leading to unsatisfactory performance. To address this problem, we propose a novel framework for sarcasm detection, which conducts incongruity reasoning based on commonsense augmentation, called EICR. Concretely, we first employ retrieval-augmented large language models to supplement the missing but indispensable commonsense background knowledge. To capture complex contextual associations, we construct a dependency graph and obtain the optimized topology via graph refinement. We further introduce an adaptive reasoning skeleton that integrates prior rules to extract sentiment-inconsistent subgraphs explicitly. To eliminate the possible spurious relations between words and labels, we employ adversarial contrastive learning to enhance the robustness of the detector. Experiments conducted on five datasets demonstrate the effectiveness of EICR.
Ziqi Qiu, Jianxing Yu, Hanjiang Lai, Yanghui Rao, Qinliang Su, Jian Yin 0001
COLING4
2025 Generating Commonsense Reasoning Questions with Controllable Complexity through Multi-step Structural Composition
abstract
This paper studies the task of generating commonsense reasoning questions (QG) with desired difficulty levels. Compared to traditional shallow questions that can be solved by simple term matching, ours are more challenging. Our answering process requires reasoning over multiple contextual and commonsense clues. That involves advanced comprehension skills, such as abstract semantics learning and missing knowledge inference. Existing work mostly learns to map the given text into questions, lacking a mechanism to control results with the desired complexity. To address this problem, we propose a novel controllable framework. We first derive contextual and commonsense clues involved in reasoning questions from the text. These clues are used to create simple sub-questions. We then aggregate multiple sub-questions to compose complex ones under the guidance of prior reasoning structures. By iterating this process, we can compose a complex QG task based on a series of smaller and simpler QG subtasks. Each subtask serves as a building block for a larger one. Each composition corresponds to an increase in the reasoning step. Moreover, we design a voting verifier to ensure results’ validity from multiple views, including answer consistency, reasoning difficulty, and context correlation. Finally, we can learn the optimal QG model to yield thought-provoking results. Evaluations on two typical datasets validate our method.
Jianxing Yu, Shiqi Wang 0016, Hanjiang Lai, Wenqing Chen, Yanghui Rao, Qinliang Su, Jian Yin 0001
COLING3
2025 CISTF: A Causal Inference Method for Stock Trend Forecasting
abstract
Stock trend forecasting, which aims to predict future fluctuations of stock price, has garnered significant attention in recent years. However, it is quite hard to train a model that can consistently and accurately forecasts the future movements of stocks. Since stock data is influenced by various unobservable factors, such as market sentiment and industry conditions, its distribution is unstable and can change over time. As a result, models trained on stock data often exhibit poor performance in prediction. Causal inference is an effective method that has been widely used to address distribution shift. However, how to conduct causal inference for stock trend forecasting still remains under-explored. To this end, we propose CISTF, a method based on causal inference to explore the causal dependencies between the input historical data and the future trend of stocks. We first construct a causal graph for the stock trend forecasting task, in which we use confounders to represent unobservable factors that affect both the input data and the future trend of stocks. Such time-varying confounders contribute to the distribution drift in the stock data. Then we use front-door adjustment, an intervention strategie that allow we to intervene the input data, to eliminate the spurious correlations brought by the unobserved confounders. Additionally, we develop a deep architecture to implement front-door adjustment, which consists of a feature extractor, a mediator estimation module and a conditional probability estimation module. We evaluate the effectiveness of CISTF on real-world stock data, and experiments on three public stock datasets demonstrate that our model achieves state-of-the-art performance.
Rongbang Qiu, Qinkang Gong, Yan Pan 0002, Hanjiang Lai
CSCWD4
2025 On the Zero-shot Adversarial Robustness of Vision-Language Models: A Truly Zero-shot and Training-free Approach
abstract
Pre-trained Vision-Language Models (VLMs) like CLIP, have demonstrated strong zero-shot generalization capabilities. Despite their effectiveness on various downstream tasks, they remain vulnerable to adversarial samples. Existing methods fine-tune VLMs to improve their performance via performing adversarial training on a certain dataset. However, this can lead to model overfitting and is not a true zero-shot scenario. In this paper, we propose a truly zero-shot and training-free approach that can significantly improve the VLM’s zero-shot adversarial robustness. Specifically, we first discover that simply adding Gaussian noise greatly enhances the VLM’s zero-shot performance. Then, we treat the adversarial examples with added Gaussian noise as anchors and strive to find a path in the embedding space that leads from the adversarial examples to the cleaner samples. We improve the VLMs’ generalization abilities in a truly zero-shot and training-free manner compared to previous methods. Extensive experiments on 16 datasets demonstrate that our method can achieve state-of-the-art zero-shot robust performance, improving the top-1 robust accuracy by an average of 9.77%. The code will be publicly available.
Baoshun Tong, Hanjiang Lai, Yan Pan 0002, Jian Yin 0001
CVPR2
2025 Test-time Alignment-Enhanced Adapter for Vision-Language Models
abstract
Test-time adaptation with pre-trained vision-language models (VLMs) has attracted increasing attention for tackling the issue of distribution shift during the test phase. While prior methods have shown effectiveness in addressing distribution shift by adjusting classification logits, they are not optimal due to keeping text features unchanged. To address this issue, we introduce a new approach called Test-time Alignment-Enhanced Adapter (TAEA), which trains an adapter with test samples to adjust text features during the test phase. We can enhance the text-to-image alignment prediction by utilizing an adapter to adapt text features. Furthermore, we also propose to adopt the negative cache from TDA as enhancement module, which further improves the performance of TAEA. Our approach outperforms the state-of-the-art TTA method of pre-trained VLMs by an average of 0.75% on the out-of-distribution benchmark and 2.5% on the cross-domain benchmark, with an acceptable training time. Code will be available at https://github.com/BaoshunWq/clip-TAEA.
Baoshun Tong, Kaiyu Song, Hanjiang Lai
ICASSP3
2025 Enhancing Few-Shot Out-of-Distribution Detection with Gradient Aligned Context Optimization
abstract
Few-shot out-of-distribution (OOD) detection aims to detect OOD images from unseen classes with only a few labeled in-distribution (ID) images. To detect OOD images and classify ID samples, prior methods have been proposed by regarding the background regions of ID samples as the OOD knowledge and performing OOD regularization and ID classification optimization. However, the gradient conflict still exists between ID classification optimization and OOD regularization caused by biased recognition. To address this issue, we present Gradient Aligned Context Optimization (GaCoOp) to mitigate this gradient conflict. Specifically, we decompose the optimization gradient to identify the scenario when the conflict occurs. Then we alleviate the conflict in inner ID samples and optimize the prompts via leveraging gradient projection. Extensive experiments over the large-scale ImageNet OOD detection benchmark demonstrate that our GaCoOp can effectively mitigate the conflict and achieve great performance. Code will be available at https://github.com/BaoshunWq/ood-GaCoOp.
Baoshun Tong, Kaiyu Song, Hanjiang Lai
ICASSP3
2025 HBRW: A Hardness-Based Re-Weighting Approach for Long-tailed Medical Image Classification
abstract
Traditional image classification methods struggle with the skewed distribution of medical images, where some diseases are overrepresented while others are underrepresented. In light of the diverse learning challenges presented by different medical image samples, we introduce a novel approach to long-tailed medical image classification through a re-weighting strategy based on sample hardness. Our method differs from existing strategies by determining a sample’s weight in the loss function based solely on its hardness - a measure of the sample’s learning difficulty. This approach simplifies the training process, avoids heavy reliance on hyper-parameters, and enhances the model’s generalizability and robustness, especially in the context of long-tailed distributions common in medical image datasets. We validate the effectiveness of our method with comprehensive experiments on the long-tailed medical datasets, demonstrating significant improvements in classification performance compared to existing methods.
Yongheng Xu, Hanjiang Lai
ICASSP2
2025 Ranking-Aware Uncertainty for Text-Guided Image Retrieval
Hanjiang Lai
ICIC (16)2
2025 Pretrain like Your Inference: Masked Tuning Improves Zero-Shot Composed Image Retrieval
abstract
Zero-shot composed image retrieval (ZS-CIR), which takes a textual modification and a reference image as a query to retrieve a target image without triplet labeling, has gained more and more attention. Current ZS-CIR research relies mainly on the generalization ability of pre-trained vision language models (VLMs). However, the pre-trained VLMs and CIR tasks have substantial discrepancies, where the VLMs focus on learning the similarities but CIR aims to learn the modifications of the image guided by text. In this paper, we introduce a novel unlabeled and pre-trained masked tuning approach, which reduces the gap between the pre-trained VLMs and the downstream CIR task. First, to reduce the gap, we reformulate the contrastive learning of the VLMs as the CIR task, where we randomly mask input image patches to generate ⟨masked image, text, image⟩ triplet from an image-text pair. Then, we propose a simple but novel pre-trained masked tuning method, which uses the text and the masked image to learn the modifications of the original image. With such a simple design, the proposed masked tuning can learn to better capture fine-grained text-guided modifications. Extensive experimental results demonstrate the significant superiority of our approach over the baseline models on four ZS-CIR datasets, including FashionIQ, CIRR, CIRCO, and GeneCIS.
Hanjiang Lai
ICME2
2025 CPMDiff: Classifier Probability Measurement for Out-of-Distribution Detection via Diffusion Models
abstract
Recent research has explored using diffusion models for out-of-distribution (OOD) detection, leveraging their ability to distinguish OOD samples by measuring reconstruction errors. Existing methods use visual feature metrics to measure reconstruction errors. However, the inherent stochasticity in diffusion models introduces variability, distorting reconstruction error measurements with visual feature metrics. To address this issue, we propose Classifier Probability Measurement via Diffusion Models (CPMDiff), a novel OOD detection method that measures the classifier probabilities’ discrepancy between input and reconstructed data. By leveraging classifier space measurement, CPMDiff alleviates the distortion in reconstruction error metrics caused by the stochasticity of diffusion models. Experiments on benchmark datasets demonstrate that CPMDiff significantly improves the performance of diffusion model-based methods in OOD detection, especially on challenging datasets.
Yongheng Xu, Kaiyu Song, Hanjiang Lai
ICME3
2025 A Causal Intervention Method for Domain Generalization with a Self-Supervised Auxiliary Task
Qinkang Gong, Yan Pan 0002, Hanjiang Lai, Jian Yin 0001
Int. J. Comput. Vis.3
2025 Efficient Multimodal Selection for Retrieval in Knowledge-Based Visual Question Answering
abstract
Retrieval plays an important role in knowledge-based visual question answering (KB-VQA), which relies on external knowledge to answer questions related to an image. However, not all information in the external knowledge is beneficial in retrieval, e.g., the knowledge that is only semantically similar to the query but is not useful for question answering. To improve the effectiveness and efficiency of retrieval, in this paper, we propose efficient multimodal selection to filter out irrelevant information and increase the retriever performance for KB-VQA. First, to exclude most irrelevant knowledge from the large external knowledge, multimodal selection uses a query-aware sample selection method, which uses the pretrained answer generator’s prediction to obtain better positive and negative training samples to help retrievers distinguish knowledge that is semantically relevant to the multimodal query. Then, question-aware visual feature selection is proposed to select the distinguishable visual information related to the question: where cross-attention to questions and images is proposed to obtain question-aware visual features. These visual features are used to perform fine-grained multimodal retrieval within the small set to obtain the final top-related knowledge. The experimental results show that the proposed approach achieves state-of-the-art retrieval performance on the OK-VQA and FVQA datasets, indicating the effectiveness of our selection strategy for retrieval.
Linyin Luo, Hanjiang Lai, Yan Pan 0002, Jian Yin 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Causal-TSF: A Causal Intervention Approach to Mitigate Confounding Bias in Time Series Forecasting
abstract
Time series forecasting, aiming to learn models from historical data and predict future values in time series, is a fundamental research topic in machine learning. However, few efforts have been devoted to addressing the confounding effects in time series data, e.g., the historical data are affected by some hidden surrounding factors (i.e., confounders), leading to biased forecasting models for future data. This paper presents a causal intervention approach to eliminate the bias that is raised by some hidden confounders. By using a causal graph, we illustrate why hidden confounders can bring bias in time series forecasting and how to tackle it. We implement causal intervention by a deep architecture that consists of two modules, a Confounders Estimation module to estimate the hidden confounders and a Debiasing module to eliminate the confounding bias in the forecasting model via sampling on confounders. We conduct comprehensive evaluations on various time series datasets. The experiment results indicate that the proposed method can reduce the negative confounding effects in time series data, and it achieves superior gains over state-of-the-art baselines for time series forecasting.
Qinkang Gong, Yan Pan 0002, Hanjiang Lai, Rongbang Qiu, Jian Yin 0001
IEEE Trans. Knowl. Data Eng.3
2024 Hierarchical Topic Modeling via Contrastive Learning and Hyperbolic Embedding
abstract
Hierarchical topic modeling, which can mine implicit semantics in the corpus and automatically construct topic hierarchical relationships, has received considerable attention recently. However, the current hierarchical topic models are mainly based on Euclidean space, which cannot well retain the implicit hierarchical semantic information in the corpus, leading to irrational structure of the generated topics. On the other hand, the existing Generative Adversarial Network (GAN) based neural topic models perform satisfactorily, but they remain constrained by pattern collapse due to the discontinuity of latent space. To solve the above problems, with the hypothesis of hyperbolic space, we propose a novel GAN-based hierarchical topic model to mine high-quality topics by introducing contrastive learning to capture information from documents. Furthermore, the distinct tree-like property of hyperbolic space preserves the implicit hierarchical semantics of documents in topic embeddings, which are projected into the hyperbolic space. Finally, we use a multi-head self-attention mechanism to learn implicit hierarchical semantics of topics and mine topic structure information. Experiments on real-world corpora demonstrate the remarkable performance of our model on topic coherence and topic diversity, as well as the rationality of the topic hierarchy.
Zhicheng Lin, Hegang Chen, Yuyin Lu, Yanghui Rao, Hanjiang Lai
LREC/COLING6
2024 MimicDiffusion: Purifying Adversarial Perturbation via Mimicking Clean Diffusion Model
abstract
Deep neural networks (DNNs) are vulnerable to adversarial perturbation, where an imperceptible perturbation is added to the image that can fool the DNNs. Diffusion-based adversarial purification uses the diffusion model to generate a clean image against such adversarial attacks. Un-fortunately, the generative process of the diffusion model is also inevitably affected by adversarial perturbation since the diffusion model is also a deep neural network where its input has adversarial perturbation. In this work, we propose MimicDiffusion, a new diffusion-based adversarial purification technique that directly approximates the generative process of the diffusion model with the clean image as input. Concretely, we analyze the differences between the guided terms using the clean image and the ad-versarial sample. After that, we first implement MimicD-iffusion based on Manhattan distance. Then, we propose two guidance to purify the adversarial perturbation and ap-proximate the clean diffusion model. Extensive experiments on three image datasets, including CIFAR-10, CIFAR-100, and ImageNet, with three classifier backbones including WideResNet-70-16, WideResNet-28-10, and ResNet-50 demonstrate that MimicDiffusion significantly performs better than the state-of-the-art baselines. On CIFAR -10, CIFAR-100, and ImageNet, it achieves 92.67%, 61.35%, and 61.53% average robust accuracy, which are 18.49%, 13.23%, and 17.64% higher, respectively. The code is available at https://github.com/psky1111/MimicDiffusion.
Kaiyu Song, Hanjiang Lai, Yan Pan 0002, Jian Yin 0001
CVPR2
2024 MALIP: Improving Few-Shot Image Classification with Multimodal Fusion Enhancement
abstract
With the significant progress in pre-trained visionlanguage models like CLIP, recent CLIP-based methods have shown impressive performance in few-shot tasks. However, CLIPbased representations have a natural gap in downstream few-shot tasks due to the label-related multimodal information scarcity caused by limited data. We then question, whether the generative model trained in downstream tasks could be used to enhance label-related multimodal fusion. In this paper, we propose a generative model-based multimodal fusion enhancement method, MALIP, to improve the few-shot performance of CLIP via a Multimodal Adapter module. Specifically, we first leverage a variational autoencoder (VAE) that could be trained in the few-shot scenario to extend the data. Then we create adapter weights by a key-value cache model constructed from the image and text information based on the expanded data. In the end, through extensive experiments on 11 datasets, we demonstrate the effectiveness of MALIP to perform state-of-the-art few-shot image classification.
Kaifen Cai, Kaiyu Song, Yan Pan 0002, Hanjiang Lai
ICME4
2024 DNAF: Diffusion with Noise-Aware Feature for Pose-Guided Person Image Synthesis
abstract
Pose-guided person image synthesis aims at generating images based on the related pose skeleton and the appearance of a source image. As a popular generative model, the diffusion model shows its potential. However, there are two gaps to hinder the fusion between pose information and appearance: 1) Directly injecting pixel-level pose information into semantic features leads to the representation gap. 2) The timestep-dependent nature of the diffusion model introduces the noise-induced gap. To alleviate these, we propose Diffusion with Noise-Aware Feature(DNAF). Concretely, we leverage the T2I-Adapter-based pose adapter to achieve the mapping from the pixel level to the feature level. Then, we propose a lightweight trainable layer to infuse the multi-scale constant feature adaptively. In the end, we construct noise-aware features to more effectively guide the diffusion process. Experimental results show that DNAF achieves competitive results on DeepFashion and Market-1501 datasets.
Liyan Guo, Kaiyu Song, Mengying Xu, Hanjiang Lai
ICME4
2024 Long-Tailed Hashing with Wasserstein Quantization
Zujun Fu, Hanjiang Lai, Yan Pan 0002
ICPR (21)2
2024 MAMixer: Multivariate Time Series Forecasting via Multi-axis Mixing
Yongyu Liu, Guoliang Lin, Hanjiang Lai, Yan Pan 0002
MMM (1)3
2024 Cross-Modal Hash Retrieval with Category Semantics
Mengying Xu, Hanjiang Lai, Jian Yin 0001
MMM (1)2
2024 Category-Level Contrastive Learning for Unsupervised Hashing in Cross-Modal Retrieval
abstract
Abstract Unsupervised hashing for cross-modal retrieval has received much attention in the data mining area. Recent methods rely on image-text paired data to conduct unsupervised cross-modal hashing in batch samples. There are two main limitations for existing models: (1) learning of cross-modal representations is restricted to batches; (2) semantically similar samples may be wrongly treated as negative. In this paper, we propose a novel category-level contrastive learning for unsupervised cross-modal hashing, which alleviates the above problems and improves cross-modal query accuracy. To break the limitation of learning in small batches, a selected memory module is first proposed to take global relations into account. Then, we obtain pseudo labels through clustering and combine the labels with the Hadamard Matrix for category-centered learning. To reduce wrong negatives, we further propose a memory bank to store clusters of samples and construct negatives by selecting samples from different categories for contrastive learning. Extensive experiments show the significant superiority of our approach over the state-of-the-art models on MIRFLICKR-25K and NUS-WIDE datasets.
Mengying Xu, Linyin Luo, Hanjiang Lai, Jian Yin 0001
Data Sci. Eng.3
2024 Revisiting Few-Shot Learning From a Causal Perspective
abstract
Few-shot learning with$N$-way$K$-shot scheme is an open challenge in machine learning. Many metric-based approaches have been proposed to tackle this problem, e.g., the Matching Networks and CLIP-Adapter. Despite that these approaches have shown significant progress, the mechanism of why these methods succeed has not been well explored. In this paper, we try to interpret these metric-based few-shot learning methods via causal mechanism. We show that the existing approaches can be viewed as specific forms of front-door adjustment, which can alleviate the effect of spurious correlations and thus learn the causality. This causal interpretation could provide us a new perspective to better understand these existing metric-based methods. Further, based on this causal interpretation, we simply introduce two causal methods for metric-based few-shot learning, which considers not only the relationship between examples but also the diversity of representations. Experimental results demonstrate the superiority of our proposed methods in few-shot classification on various benchmark datasets. Code is available inhttps://github.com/lingl1024/causalFewShot.
Guoliang Lin, Yongheng Xu, Hanjiang Lai, Jian Yin 0001
IEEE Trans. Knowl. Data Eng.3
2023 Counterfactual Multihop QA: A Cause-Effect Approach for Reducing Disconnected Reasoning
abstract
Multi-hop QA requires reasoning over multiple supporting facts to answer the question.However, the existing QA models always rely on shortcuts, e.g., providing the true answer by only one fact, rather than multi-hop reasoning, which is referred to as disconnected reasoning problem.To alleviate this issue, we propose a novel counterfactual multihop QA, a causaleffect approach that enables to reduce the disconnected reasoning.It builds upon explicitly modeling of causality: 1) the direct causal effects of disconnected reasoning and 2) the causal effect of true multi-hop reasoning from the total causal effect.With the causal graph, a counterfactual inference is proposed to disentangle the disconnected reasoning from the total causal effect, which provides us a new perspective and technology to learn a QA model that exploits the true multi-hop reasoning instead of shortcuts.Extensive experiments have been conducted on the benchmark HotpotQA dataset, which demonstrate that the proposed method can achieve notable improvement on reducing disconnected reasoning.For example, our method achieves 5.8% higher points of its Supp s score on HotpotQA through true multihop reasoning.The code is available at https://github.com/guowzh/CFMQA.
Wangzhen Guo, Qinkang Gong, Yanghui Rao, Hanjiang Lai
ACL (1)4
2023 Spatial Commonsense Reasoning for Machine Reading Comprehension
Miaopei Lin, Mengxiang Wang, Jianxing Yu, Shiqi Wang 0016, Hanjiang Lai, Wei Liu 0061, Jian Yin 0001
ADMA (2)5
2023 CLIP-HASH: A Lightweight Hashing Network for Cross-Modal Retrieval
abstract
Cross-modal hashing has attracted great interest in the past decades. Due to the traditional hashing retrieval model requiring hand-crafted feature extraction, much research on the end-to-end deep hashing models has been explored. However, this training process of deep models is complex and needs a large number of samples. To address the above issues, we propose a lightweight hashing retrieval network with the pre-trained CLIP model, which is called CLIP-Hash. It can obtain better hash features via the pre-training CLIP model and provide a lightweight fine-tune model which is easily trained and only few training samples are required. First, we obtain the powerful feature representations by the visual and textual encodes. Then, we propose a simple hash encoding module, which is a lightweight network and only few parameters are needed to be trained. With the proposed hash encoding module, we can obtain efficient binary hash codes that can align the images and texts. We conduct experiments on MIRFlickr and NUS-WIDE two benchmark datasets. The results show that the proposed simple method can outperform the SOTA hashing methods.
Mengying Xu, Hanjiang Lai, Jian Yin 0001
CSCWD2
2023 Deep Hashing with Minimal-Distance-Separated Hash Centers
abstract
Deep hashing is an appealing approach for large-scale image retrieval. Most existing supervised deep hashing methods learn hash functions using pairwise or triple image similarities in randomly sampled mini-batches. They suffer from low training efficiency, insufficient coverage of data distribution, and pair imbalance problems. Recently, central similarity quantization (CSQ) attacks the above problems by using “hash centers” as a global similarity metric, which encourages the hash codes of similar images to approach their common hash center and distance themselves from other hash centers. Although achieving SOTA retrieval performance, CSQ falls short of a worst-case guarantee on the minimal distance between its constructed hash centers, i.e. the hash centers can be arbitrarily close. This paper presents an optimization method that finds hash centers with a constraint on the minimal distance between any pair of hash centers, which is non-trivial due to the non-convex nature of the problem. More importantly, we adopt the Gilbert-Varshamov bound from coding theory, which helps us to obtain a large minimal distance while ensuring the empirical feasibility of our optimization approach. With these clearly-separated hash centers, each is assigned to one image class, we propose several effective loss functions to train deep hashing networks. Extensive experiments on three datasets for image retrieval demonstrate that the proposed method achieves superior retrieval performance over the state-of-the-art deep hashing methods.
Liangdao Wang, Yan Pan 0002, Cong Liu 0001, Hanjiang Lai, Jian Yin 0001
CVPR4
2023 From Parse-Execute to Parse-Execute-Refine: Improving Semantic Parser for Complex Question Answering over Knowledge Base
abstract
Parsing questions into executable logical forms has showed impressive results for knowledgebase question answering (KBQA).However, complex KBQA is a more challenging task that requires to perform complex multi-step reasoning.Recently, a new semantic parser called KoPL (Cao et al., 2022a) has been proposed to explicitly model the reasoning processes, which achieved the state-of-the-art on complex KBQA.In this paper, we further explore how to unlock the reasoning ability of semantic parsers by a simple proposed parseexecute-refine paradigm.We refine and improve the KoPL parser by demonstrating the executed intermediate reasoning steps to the KBQA model.We show that such simple strategy can significantly improve the ability of complex reasoning.Specifically, we propose three components: a parsing stage, an execution stage and a refinement stage, to enhance the ability of complex reasoning.The parser uses the KoPL to generate the transparent logical forms.Then, the execution stage aligns and executes the logical forms over knowledge base to obtain intermediate reasoning processes.Finally, the intermediate step-by-step reasoning processes are demonstrated to the KBQA model in the refinement stage.With the explicit reasoning processes, it is much easier to answer the complex questions.Experiments on benchmark dataset shows that the proposed PER-KBQA performs significantly better than the stage-of-the-art baselines on the complex KBQA.
Wangzhen Guo, Linyin Luo, Hanjiang Lai, Jian Yin 0001
EMNLP3
2023 Multimodal Reconstruct and Align Net for Missing Modality Problem in Sentiment Analysis
Mengying Xu, Hanjiang Lai
MMM (2)3
2023 TADC: A Topic-Aware Dynamic Convolutional Neural Network for Aspect Extraction
abstract
Aspect extraction is one of the key tasks in fine-grained sentiment analysis. This task aims to identify explicit opinion targets from user-generated documents. Currently, the mainstream methods for aspect extraction are built on recurrent neural networks (RNNs), which are difficult to parallelize. To accelerate the training/testing process, convolutional neural network (CNN)-based methods are introduced. However, such models usually utilize the same set of filters to convolve all input documents, and hence, the unique information inherent in each document may not be fully captured. To alleviate this issue, we propose a CNN-based model that employs a set of dynamic filters. Specifically, the proposed model extracts the aspects in a document using the filters generated from the aspect information intrinsic in the document. With the dynamically generated filters, our model is capable of learning more important features concerning aspects, thus promoting the effectiveness of aspect extraction. Furthermore, considering that aspects can be grouped into certain topics that conversely indicate the target words that need to be extracted, we naturally introduce a neural topic model (NTM) and integrate latent topics into the CNN-based module to help identify aspects. Experiments on two benchmark datasets demonstrate that the joint model is able to effectively identify aspects and produce interpretable topics.
Zusheng Zhang 0003, Yanghui Rao, Hanjiang Lai, Jiahai Wang, Jian Yin 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Image Retrieval with Well-Separated Semantic Hash Centers
Liangdao Wang, Yan Pan 0002, Hanjiang Lai, Jian Yin 0001
ACCV (6)3
2022 Towards Better Plasticity-Stability Trade-off in Incremental Learning: A Simple Linear Connector
abstract
Plasticity-stability dilemma is a main problem for incremental learning, where plasticity is referring to the ability to learn new knowledge, and stability retains the knowledge of previous tasks. Many methods tackle this problem by storing previous samples, while in some applications, training data from previous tasks cannot be legally stored. In this work, we propose to employ mode connectivity in loss landscapes to achieve better plasticity-stability trade-off without any previous samples. We give an analysis of why and how to connect two independently optimized optima of networks, null-space projection for previous tasks and simple SGD for the current task, can attain a meaningful balance between preserving already learned knowledge and granting sufficient flexibility for learning a new task. This analysis of mode connectivity also provides us a new perspective and technology to control the trade-off between plasticity and stability. We evaluate the proposed method on several benchmark datasets. The results indicate our simple method can achieve notable improvement, and perform well on both the past and current tasks. On 10-split-CIFAR-100 task, our method achieves 79.79% accuracy, which is 6.02% higher. Our method also achieves 6.33% higher accuracy on TinyImageNet. Code is available at https://github.com/lingl1024/Connector.
Guoliang Lin, Hanlu Chu, Hanjiang Lai
CVPR3
2022 You Never Stop Dancing: Non-freezing Dance Generation via Bank-constrained Manifold Projection
abstract
One of the most overlooked challenges in dance generation is that the auto-regressive frameworks are prone to freezing motions due to noise accumulation. In this paper, we present two modules that can be plugged into the existing models to enable them to generate non-freezing and high fidelity dances. Since the high-dimensional motion data are easily swamped by noise, we propose to learn a low-dimensional manifold representation by an auto-encoder with a bank of latent codes, which can be used to reduce the noise in the predicted motions, thus preventing from freezing. We further extend the bank to provide explicit priors about the future motions to disambiguate motion prediction, which helps the predictors to generate motions with larger magnitude and higher fidelity than possible before. Extensive experiments on AIST++, a public large-scale 3D dance motion benchmark, demonstrate that our method notably outperforms the baselines in terms of quality, diversity and time length.
Jiangxin Sun, Huang Hu, Hanjiang Lai, Zhi Jin 0002, Jianfang Hu
NeurIPS4
2022 Efficient modal-aware feature learning with application in multimodal hashing
abstract
Many retrieval applications can benefit from multiple modalities, for which how to represent multimodal data is the critical component. Most deep multimodal learning methods typically involve two steps to construct the joint representations: 1) learning of multiple intermediate features, with each intermediate feature corresponding to a modality, using separate and independent deep models; 2) merging the intermediate features into a joint representation using a fusion strategy. However, in the first step, these intermediate features do not have previous knowledge of each other and cannot fully exploit the information contained in the other modalities. In this paper, we present a modal-aware operation as a generic building block to capture the non-linear dependencies among the heterogeneous intermediate features, which can learn the underlying correlation structures in other multimodal data as soon as possible. The modal-aware operation consists of a kernel network and an attention network. The kernel network is utilized to learn the non-linear relationships with other modalities. The attention network finds the informative regions of these modal-aware features that are favorable for retrieval. We verify the proposed modal-aware feature learning in the multimodal hashing task. The experiments conducted on three public benchmark datasets demonstrate significant improvements in the performance of our method relative to state-of-the-art methods.
Hanlu Chu, Haien Zeng, Hanjiang Lai, Yong Tang 0001
Intell. Data Anal.3
2022 Deep Listwise Triplet Hashing for Fine-Grained Image Retrieval
abstract
Hashing is a practical approach for the approximate nearest neighbor search. Deep hashing methods, which train deep networks to generate compact and similarity-preserving binary codes for entities (e.g. images), have received lots of attention in the information retrieval community. A representative stream of deep hashing methods is triplet-based hashing that learns hashing models from triplets of data. The existing triplet-based hashing methods only consider triplets that are in the form of$(q,q^{+},q^{-})$, where$q$and$q^{+}$are in the same class and$q$and$q^{-}$are in different classes. However, the number of possible triplets is approximately the cube of training examples, triplets used in the existing methods are only a small fraction of all possible triplets. This motivates us to develop a new triplet-based hashing method that adopts many more triplets in training phase. We propose Deep Listwise Triplet Hashing (DLTH) that introduces more triplets into batch-based training and a novel listwise triplet loss to capture the relative similarity in new triplets. This method has a pipeline of two steps. In Step 1, we propose a novel way to generate triplets from the soft class labels obtained by knowledge distillation module, where the triplets in the form of$(q,q^{+},q^{-})$are a subset of the newly obtained triplets. In Step 2, we develop a novel listwise triplet loss to train the hashing network, which seeks to capture the relative similarity between images in triplets according to soft labels. We conduct comprehensive image retrieval experiments on four benchmark datasets. The experimental results show that the proposed method has superior performances over state-of-the-art baselines.
Yan Pan 0002, Hanjiang Lai, Wei Liu 0061, Jian Yin 0001
IEEE Trans. Image Process.3
2021 Variable-Length Metric Learning for Fast Image Retrieval
abstract
Learning a powerful distance metric is the key component of image retrieval. Recently, deep metric learning has been an active research topic for image retrieval. However, most existing metric learning approaches treat all the input images equally and learn all image embeddings at equal lengths. These methods ignore those easy examples that can be encoded as the shorter features, which is search-inefficient. We propose a simple but efficient variable-length metric learning method for search efficiency, in which the different query samples have different feature-length. First, we propose to learn the ranked (prioritized) list of features. The more distinguishing features, the higher the rank. We show that the proposed prioritized features can be used to perform fast retrieval in different feature-length configurations. Further, we propose an adaptive feature-length selection policy to determine the amount of feature-length for each query sample. Extensive experiments are conducted on three benchmark datasets. The results demonstrate that the proposed method can reduce computational costs without incurring a decrease in accuracy.
Hanjiang Lai, Yan Pan 0002
CSCWD2
2021 Incorporating Domain Knowledge and Semantic Information into Language Models for Commonsense Question Answering
abstract
Commonsense question answering (CSQA) aims to answer questions which require the system to understand related commonsense knowledge that is not explicitly expressed in the given context. Recent advance in neural language models (e.g., BERT) that are pre-trained on a large-scale text corpus and fine-tuned on downstream tasks has boosted the performance on CSQA. However, due to the lack of domain knowledge (e.g., in social situations), these models fail to reason about specific tasks. In this work, we propose an approach to incorporate domain knowledge and semantic information into language model based approaches for better understanding the related commonsense knowledge. Firstly, we extract the knowledge from existing resources by jointly learning to ask and answer as well as semantic role labeling based answering. These two tasks are correlated and can reinforce each other to discover the domain knowledge. Then, we utilize Semantic Role Labeling to enable the system to gain a better understanding of relations among relevant entities. Experimental results on several CSQA benchmarks demonstrate the effectiveness of the proposed approach.
Ruiying Zhou, Keke Tian, Hanjiang Lai, Jian Yin 0001
CSCWD3
2021 Know Yourself and Know Others: Efficient Common Representation Learning for Few-shot Cross-modal Retrieval
abstract
Learning the common representations for various modalities of data is the key component in cross-modal retrieval. Most existing deep approaches learn multiple networks to independently project each sample into a common representation. However, each representation is only extracted from the corresponding data, which totally ignores the relationships between other data. Thus it is challenging to learn efficient common representations when lacking sufficient supervised multi-modal data for training, e.g., few-shot cross-modal retrieval. How to efficiently exploit the information contained in other examples is underexplored. In this work, we present the Self-Others Net, a few-shot cross-modal retrieval model that fully exploits information contained both in its own and other samples. First, we propose a self-network to fully exploit the correlations that lurk in the data itself. It integrates the features at different layers and extracts the multi-level information in the self-network. Second, an others-network is further proposed to model the relationships among all samples, which learns the Mahalanobis tensor and mixes the prototypes of all data to capture the non-linear dependencies for common representation learning. Extensive experiments are conducted on three benchmark datasets, which demonstrate clear improvements of the proposed method over the state-of-the-arts.
Shaoying Wang, Hanjiang Lai
ICMR2
2020 Learning Expensive Coordination: An Event-Based Deep RL Approach
Runsheng Yu, Xinrun Wang, Youzhi Zhang 0001, Hanjiang Lai, Bo An 0001
ICLR6
2020 Controllable Face Aging
abstract
Motivated by the following two observations: 1) people are aging differently under different conditions for changeable facial attributes, e.g., skin color may become darker when working outside, and 2) it needs to keep some unchanged facial attributes during the aging process, e.g., race and gender, we propose a controllable face aging method via attribute disentanglement generative adversarial network. To offer fine control over the synthesized face images, first, an individual embedding of the face is directly learned from an image that contains the desired facial attribute. Second, since the image may contain other unwanted attributes, an attribute disentanglement network is used to separate the individual embedding and learn the common embedding that contains information about the face attribute (e.g., race). With the common embedding, we can manipulate the generated face image with the desired attribute in an explicit manner. Experimental results on two common benchmarks demonstrate that our proposed generator achieves comparable performance on the aging effect with state-of-the-art baselines while gaining more flexibility for attribute control. Code is available at supplementary material.
Haien Zeng, Hanjiang Lai
ICPR2
2020 Robust multi-view clustering via inter-and-intra-view low rank fusion
Yan Pan 0002, Hanjiang Lai, Jian Yin 0001
Neurocomputing3
2020 Improving Deep Binary Embedding Networks by Order-Aware Reweighting of Triplets
abstract
In this paper, we focus on triplet-based deep binary embedding networks for image retrieval task. The triplet loss has been shown to be effective for hashing retrieval. However, most of the triplet-based deep networks treat the triplets equally or select the hard triplets based on the loss. Such strategies do not consider the order relations of the binary codes and ignore the hash encoding when learning the feature representations. To this end, we propose an order-aware reweighting method to effectively train the triplet-based deep networks, which up-weights the important triplets and down-weights the uninformative triplets via the rank lists of the binary codes. First, we present the order-aware weighting factors to indicate the importance of the triplets, which depend on the rank order of binary codes. Then, we reshape the triplet loss to the squared triplet loss such that the loss function will put more weights on the important triplets. The extensive evaluations on several benchmark datasets show that the proposed method achieves significant performance compared with the state-of-the-art baselines.
Hanjiang Lai, Jikai Chen, Libing Geng, Yan Pan 0002, Xiaodan Liang, Jian Yin 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Mix geographical information into local collaborative ranking for POI recommendation
Wei Liu 0061, Hanjiang Lai, Jing Wang 0030, Ge-Yang Ke, Jian Yin 0001
World Wide Web2
2019 Towards Multi-Pose Guided Virtual Try-On Network
abstract
Virtual try-on systems under arbitrary human poses have significant application potential, yet also raise extensive challenges, such as self-occlusions, heavy misalignment among different poses, and complex clothes textures. Existing virtual try-on methods can only transfer clothes given a fixed human pose, and still show unsatisfactory performances, often failing to preserve person identity or texture details, and with limited pose diversity. This paper makes the first attempt towards a multi-pose guided virtual try-on system, which enables clothes to transfer onto a person with diverse poses. Given an input person image, a desired clothes image, and a desired pose, the proposed Multi-pose Guided Virtual Try-On Network (MG-VTON) generates a new person image after fitting the desired clothes into the person and manipulating the pose. MG-VTON is constructed with three stages: 1) a conditional human parsing network is proposed that matches both the desired pose and the desired clothes shape; 2) a deep Warping Generative Adversarial Network (Warp-GAN) that warps the desired clothes appearance into the synthesized human parsing map and alleviates the misalignment problem between the input human pose and the desired one; 3) a refinement render network recovers the texture details of clothes and removes artifacts, based on multi-pose composition masks. Extensive experiments on commonly-used datasets and our newly-collected largest virtual try-on benchmark demonstrate that our MG-VTON significantly outperforms all state-of-the-art methods both qualitatively and quantitatively, showing promising virtual try-on performances.
Haoye Dong, Xiaodan Liang, Xiaohui Shen, Bochao Wang, Hanjiang Lai, Jia Zhu 0003, Zhiting Hu, Jian Yin 0001
ICCV5
2019 Part-Preserving Pose Manipulation for Person Image Synthesis
abstract
Manipulating person images under diverse poses, which transfers a person from one pose to another desired pose, is an interesting yet challenging task due to large non-rigid spatial deformation. Most existing works fail to preserve the fine-grained appearance consistency along with the pose changes due to the lack of explicit constraints and spatial modeling, leading to unrealistic results with severe artifacts. In this paper, we propose a novel Part-Preserving Generative Adversarial Network (PP-GAN) to achieve good manipulation quality by explicitly enforcing rich structure constraints over generative modeling. PP-GAN is proposed to decompose the challenging spatial transformation of the whole body into fine-grained part-level transformations, which are then integrated via human joint structure constraint. Given arbitrary poses, PP-GAN integrates human joint structure and region-level part cues as inputs to perform explicit generative modeling. Besides, we introduce a parsing-consistent loss to enforce semantic consistency among images with diverse poses, which guides the image synthesis from a semantic perspective. Extensive qualitative and quantitative evaluations on two benchmarks show that our PP-GAN significantly outperforms the state-of-the-art baselines in generating more realistic and plausible image synthesis results. PP-GAN successfully preserves part-level characteristics even for most challenging pose changes while prior works are easy to fail.
Haoye Dong, Xiaodan Liang, Chenxing Zhou, Hanjiang Lai, Jia Zhu 0003, Jian Yin 0001
ICME4
2019 Deep Policy Hashing Network with Listwise Supervision
abstract
Deep-networks-based hashing has become a leading approach for large-scale image retrieval, which learns a similarity-preserving network to map similar images to nearby hash codes. The pairwise and triplet losses are two widely used similarity preserving manners for deep hashing. These manners ignore the fact that hashing is a prediction task on the list of binary codes. However, learning deep hashing with listwise supervision is challenging in 1) how to obtain the rank list of whole training set when the batch size of the deep network is always small and 2) how to utilize the listwise supervision. In this paper, we present a novel deep policy hashing architecture with two systems are learned in parallel: aquery network and a shared and slowly changingdatabase network. The following three steps are repeated until convergence: 1) the database network encodes all training samples into binary codes to obtain whole rank list, 2) the query network is trained based on policy learning to maximize a reward that indicates the performance of the whole ranking list of binary codes, e.g., mean average precision (MAP), and 3) the database network is updated as the query network. Extensive evaluations on several benchmark datasets show that the proposed method brings substantial improvements over state-of-the-art hashing methods.
Shaoying Wang, Hanjiang Lai, Jian Yin 0001
ICMR2
2019 Feature Pyramid Hashing
abstract
In recent years, deep-networks-based hashing has become a leading approach for large-scale image retrieval. Most deep hashing approaches use the high layer to extract the powerful semantic representations. However, these methods have limited ability for fine-grained image retrieval because the semantic features extracted from the high layer are difficult in capturing the subtle differences. To this end, we propose a novel two-pyramid hashing architecture to learn both the semantic information and the subtle appearance details for fine-grained image search. Inspired by the feature pyramids of convolutional neural network, avertical pyramid is proposed to capture the high-layer features and ahorizontal pyramid combines multiple low-layer features with structural information to capture the subtle differences. To fuse the low-level features, a novel combination strategy, called consensus fusion, is proposed to capture all subtle information from several low-layers for finer retrieval. Extensive evaluation on two fine-grained datasets CUB-200-2011 and Stanford Dogs demonstrate that the proposed method achieves significant performance compared with the state-of-art baselines.
Libing Geng, Hanjiang Lai, Yan Pan 0002, Jian Yin 0001
ICMR3
2019 UP-CNN: Un-pooling augmented convolutional neural network
Chunyan Xu, Jian Yang 0003, Hanjiang Lai, Junbin Gao, LinLin Shen, Shuicheng Yan
Pattern Recognit. Lett.3
2019 Improved Search in Hamming Space Using Deep Multi-Index Hashing
abstract
Similarity-preserving hashing is a widely used method for nearest neighbor search in large-scale image retrieval tasks. Considerable research has been conducted on deep-network-based hashing approaches to improve the performance. However, the binary codes generated from deep networks may be not uniformly distributed over the Hamming space, which will greatly increase the retrieval time. To this end, we propose a deep-network-based multi-index hashing (MIH) for retrieval efficiency. We first introduce the MIH mechanism into the proposed deep architecture, which divides the binary codes into multiple substrings. Each substring corresponds to one hash table. Then, we add the two balanced constraints to obtain more uniformly distributed binary codes: 1) balanced substrings, where the Hamming distances of each substring are equal for any two binary codes and 2) balanced hash buckets, where the sizes of each bucket are equal. Extensive evaluations on several benchmark image retrieval data sets show that the learned balanced binary codes bring dramatic speedups and achieve comparable performance over the existing baselines.
Hanjiang Lai, Yan Pan 0002, Si Liu 0001, Zhenbin Weng, Jian Yin 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 Clothing Landmark Detection Using Deep Networks With Prior of Key Point Associations
abstract
This paper considers a problem of landmark point detection in clothes, which is important and valuable for clothing industry. A novel method for landmark localization has been proposed, which is based on a deep end-to-end architecture using prior of key point associations. With the estimated landmark points as input, a deep network has been proposed to predict clothing categories and attributes. A systematic design of the proposed detecting system is implemented by using deep learning techniques and a large-scale clothes dataset containing 145 000 upper-body clothing images with landmark annotations. Experimental results indicate that clothing categories and attributes can be well classified by using the detected landmark points, which are associated with regions of interest in clothes (e.g., the sleeves and the collars) and share robust learning representation property with respect to large variances of human poses, nonfrontal views, or occlusion. A comprehensive performance evaluation over two newly released datasets is carried out in this paper, showing that the proposed system with deep architecture for clothing landmark detection outperforms the state-of-the-art techniques.
Changqin Huang, Jikai Chen, Yan Pan 0002, Hanjiang Lai, Jian Yin 0001, Qionghao Huang
IEEE Trans. Cybern.4
2018 Attention-Aware Deep Adversarial Hashing for Cross-Modal Retrieval
Xi Zhang 0015, Hanjiang Lai, Jiashi Feng
ECCV (15)2
2018 Transductive Zero-Shot Hashing via Coarse-to-Fine Similarity Mining
abstract
Zero-shot Hashing (ZSH) is to learn hashing models for novel/target classes without training data, which is an important and challenging problem. Most existing ZSH approaches exploit transfer learning via an intermediate shared semantic representations between the seen/source classes and novel/target classes. However, the hash functions learned from the source dataset may show poor performance when directly applied to the target classes due to the dataset bias. In this paper, we study the transductive ZSH, i.e., we have unlabeled data for novel classes. We put forward a simple yet efficient joint learning approach via coarse-to-fine similarity mining which transfers knowledges from source data to target data. It mainly consists of two building blocks in the proposed deep architecture: 1) a shared two-streams network to learn the effective common image representations. The first stream operates on the source data and the second stream operates on the unlabeled data. And 2) a coarse-to-fine module to transfer the similarities of the source data to the target data in a greedy fashion. It begins with a coarse search over the unlabeled data to find the images that most dissimilar to the source data, and then detects the similarities among the found images via the fine module. Extensive evaluation results on several benchmark datasets demonstrate that the proposed hashing method achieves significant improvement over the state-of-the-art methods.
Hanjiang Lai
ICMR1
2018 Soft-Gated Warping-GAN for Pose-Guided Person Image Synthesis
abstract
Despite remarkable advances in image synthesis research, existing works often fail in manipulating images under the context of large geometric transformations. Synthesizing person images conditioned on arbitrary poses is one of the most representative examples where the generation quality largely relies on the capability of identifying and modeling arbitrary transformations on different body parts. Current generative models are often built on local convolutions and overlook the key challenges (e.g. heavy occlusions, different views or dramatic appearance changes) when distinct geometric changes happen for each part, caused by arbitrary pose manipulations. This paper aims to resolve these challenges induced by geometric variability and spatial displacements via a new Soft-Gated Warping Generative Adversarial Network (Warping-GAN), which is composed of two stages: 1) it first synthesizes a target part segmentation map given a target pose, which depicts the region-level spatial layouts for guiding image synthesis with higher-level structure constraints; 2) the Warping-GAN equipped with a soft-gated warping-block learns feature-level mapping to render textures from the original image into the generated segmentation map. Warping-GAN is capable of controlling different transformation degrees given distinct target poses. Moreover, the proposed warping-block is light-weight and flexible enough to be injected into any networks. Human perceptual studies and quantitative evaluations demonstrate the superiority of our Warping-GAN that significantly outperforms all existing methods on two large datasets.
Haoye Dong, Xiaodan Liang, Hanjiang Lai, Jia Zhu 0003, Jian Yin 0001
NeurIPS4
2018 SRNN: Self-regularized neural network
Chunyan Xu, Jian Yang 0003, Junbin Gao, Hanjiang Lai, Shuicheng Yan
Neurocomputing4
2018 EGRank: An exponentiated gradient algorithm for sparse learning-to-rank
Yan Pan 0002, Jintang Ding, Hanjiang Lai, Changqin Huang
Inf. Sci.4
2018 Personalized Age Progression with Bi-Level Aging Dictionary Learning
abstract
Age progression is defined as aesthetically re-rendering the aging face at any future age for an individual face. In this work, we aim to automatically render aging faces in a personalized way. Basically, for each age group, we learn an aging dictionary to reveal its aging characteristics (e.g., wrinkles), where the dictionary bases corresponding to the same index yet from two neighboring aging dictionaries form a particular aging pattern cross these two age groups, and a linear combination of all these patterns expresses a particular personalized aging process. Moreover, two factors are taken into consideration in the dictionary learning process. First, beyond the aging dictionaries, each person may have extra personalized facial characteristics, e.g., mole, which are invariant in the aging process. Second, it is challenging or even impossible to collect faces of all age groups for a particular person, yet much easier and more practical to get face pairs from neighboring age groups. To this end, we propose a novel Bi-level Dictionary Learning based Personalized Age Progression (BDL-PAP) method. Here, bi-level dictionary learning is formulated to learn the aging dictionaries based on face pairs from neighboring age groups. Extensive experiments well demonstrate the advantages of the proposed BDL-PAP over other state-of-the-arts in term of personalized age progression, as well as the performance gain for cross-age face verification by synthesizing aging faces.
Xiangbo Shu, Jinhui Tang 0001, Zechao Li, Hanjiang Lai, Liyan Zhang 0001, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Deep Recurrent Regression for Facial Landmark Detection
abstract
We propose a novel end-to-end deep architecture for face landmark detection, based on a deep convolutional and deconvolutional network followed by carefully designed recurrent network structures. The pipeline of this architecture consists of three parts. Through the first part, we encode an input face image to resolution-preserved deconvolutional feature maps via a deep network with stacked convolutional and deconvolutional layers. Then, in the second part, we estimate the initial coordinates of the facial key points by an additional convolutional layer on top of these deconvolutional feature maps. In the last part, by using the deconvolutional feature maps and the initial facial key points as input, we refine the coordinates of the facial key points by a recurrent network that consists of multiple long short-term memory components. Extensive evaluations on several benchmark data sets show that the proposed deep architecture has superior performance against the state-of-the-art methods.
Hanjiang Lai, Shengtao Xiao, Yan Pan 0002, Zhen Cui 0001, Jiashi Feng, Chunyan Xu, Jian Yin 0001, Shuicheng Yan
IEEE Trans. Circuits Syst. Video Technol.1
2018 Object-Location-Aware Hashing for Multi-Label Image Retrieval via Automatic Mask Learning
abstract
Learning-based hashing is a leading approach of approximate nearest neighbor search for large-scale image retrieval. In this paper, we develop a deep supervised hashing method for multi-label image retrieval, in which we propose to learn a binary "mask" map that can identify the approximate locations of objects in an image, so that we use this binary "mask" map to obtain length-limited hash codes which mainly focus on an image's objects but ignore the background. The proposed deep architecture consists of four parts: 1) a convolutional sub-network to generate effective image features; 2) a binary "mask" sub-network to identify image objects' approximate locations; 3) a weighted average pooling operation based on the binary "mask" to obtain feature representations and hash codes that pay most attention to foreground objects but ignore the background; and 4) the combination of a triplet ranking loss designed to preserve relative similarities among images and a cross entropy loss defined on image labels. We conduct comprehensive evaluations on four multi-label image data sets. The results indicate that the proposed hashing method achieves superior performance gains over the state-of-the-art supervised or unsupervised hashing baselines.
Changqin Huang, Shang-Ming Yang, Yan Pan 0002, Hanjiang Lai
IEEE Trans. Image Process.4
2017 Learning Adaptive Receptive Fields for Deep Image Parsing Network
abstract
In this paper, we introduce a novel approach to regulate receptive field in deep image parsing network automatically. Unlike previous works which have stressed much importance on obtaining better receptive fields using manually selected dilated convolutional kernels, our approach uses two affine transformation layers in the networks backbone and operates on feature maps. Feature maps will be inflated/shrinked by the new layer and therefore receptive fields in following layers are changed accordingly. By end-to-end training, the whole framework is data-driven without laborious manual intervention. The proposed method is generic across dataset and different tasks. We conduct extensive experiments on both general parsing task and face parsing task as concrete examples to demonstrate the methods superior regulation ability over manual designs.
Zhen Wei 0001, Yao Sun 0004, Jinqiao Wang, Hanjiang Lai, Si Liu 0001
CVPR4
2017 Deep Progressive Hashing for Image Retrieval
abstract
This paper proposes a novel recursive hashing scheme, in contrast to conventional "one-off" based hashing algorithms. Inspired by human's "nonsalient-to-salient" perception path, the proposed hashing scheme generates a series of binary codes based on progressively expanded salient regions. Built on a recurrent deep network, i.e., LSTM structure, the binary codes generated from later output nodes naturally inherit information aggregated from previously codes while explore novel information from the extended salient region, and therefore it possesses good scalability property. The proposed deep hashing network is trained via minimizing a triplet ranking loss, which is end-to-end trainable. Extensive experimental results on several image retrieval benchmarks demonstrate good performance gain over state-of-the-art image retrieval methods and its scalability property.
Jiale Bai, Bingbing Ni, Minsi Wang, Hanjiang Lai, Lin Mei 0001, Chuanping Hu
ACM Multimedia5
2016 Robust Facial Landmark Detection via Recurrent Attentive-Refinement Networks
Shengtao Xiao, Jiashi Feng, Junliang Xing, Hanjiang Lai, Shuicheng Yan, Ashraf A. Kassim
ECCV (1)4
2016 Margin-based two-stage supervised hashing for image retrieval
Yan Pan 0002, Hanjiang Lai, Cong Liu 0001, Jian Yin 0001
Neurocomputing3
2016 Kinship-Guided Age Progression
Xiangbo Shu, Jinhui Tang 0001, Hanjiang Lai, Zhiheng Niu, Shuicheng Yan
Pattern Recognit.3
2016 Instance-Aware Hashing for Multi-Label Image Retrieval
abstract
Similarity-preserving hashing is a commonly used method for nearest neighbor search in large-scale image retrieval. For image retrieval, deep-network-based hashing methods are appealing, since they can simultaneously learn effective image representations and compact hash codes. This paper focuses on deep-network-based hashing for multi-label images, each of which may contain objects of multiple categories. In most existing hashing methods, each image is represented by one piece of hash code, which is referred to as semantic hashing. This setting may be suboptimal for multi-label image retrieval. To solve this problem, we propose a deep architecture that learns instance-aware image representations for multi-label image data, which are organized in multiple groups, with each group containing the features for one category. The instance-aware representations not only bring advantages to semantic hashing but also can be used in category-aware hashing, in which an image is represented by multiple pieces of hash codes and each piece of code corresponds to a category. Extensive evaluations conducted on several benchmark data sets demonstrate that for both the semantic hashing and the category-aware hashing, the proposed method shows substantial improvement over the state-of-the-art supervised and unsupervised hashing methods.
Hanjiang Lai, Pan Yan, Xiangbo Shu, Yunchao Wei, Shuicheng Yan
IEEE Trans. Image Process.1
2015 Simultaneous feature learning and hash coding with deep neural networks
abstract
Similarity-preserving hashing is a widely-used method for nearest neighbour search in large-scale image retrieval tasks. For most existing hashing methods, an image is first encoded as a vector of hand-engineering visual features, followed by another separate projection or quantization step that generates binary codes. However, such visual feature vectors may not be optimally compatible with the coding process, thus producing sub-optimal hashing codes. In this paper, we propose a deep architecture for supervised hashing, in which images are mapped into binary codes via carefully designed deep neural networks. The pipeline of the proposed deep architecture consists of three building blocks: 1) a sub-network with a stack of convolution layers to produce the effective intermediate image features; 2) a divide-and-encode module to divide the intermediate image features into multiple branches, each encoded into one hash bit; and 3) a triplet ranking loss designed to characterize that one image is more similar to the second image than to the third one. Extensive evaluations on several benchmark image datasets show that the proposed simultaneous feature learning and hash coding pipeline brings substantial improvements over other state-of-the-art supervised or unsupervised hashing methods.
Hanjiang Lai, Yan Pan 0002, Shuicheng Yan
CVPR1
2015 Personalized Age Progression with Aging Dictionary
abstract
In this paper, we aim to automatically render aging faces in a personalized way. Basically, a set of age-group specific dictionaries are learned, where the dictionary bases corresponding to the same index yet from different dictionaries form a particular aging process pattern cross different age groups, and a linear combination of these patterns expresses a particular personalized aging process. Moreover, two factors are taken into consideration in the dictionary learning process. First, beyond the aging dictionaries, each subject may have extra personalized facial characteristics, e.g. mole, which are invariant in the aging process. Second, it is challenging or even impossible to collect faces of all age groups for a particular subject, yet much easier and more practical to get face pairs from neighboring age groups. Thus a personality-aware coupled reconstruction loss is utilized to learn the dictionaries based on face pairs from neighboring age groups. Extensive experiments well demonstrate the advantages of our proposed solution over other state-of-the-arts in term of personalized aging progression, as well as the performance gain for cross-age face verification by synthesizing aging faces.
Xiangbo Shu, Jinhui Tang 0001, Hanjiang Lai, Luoqi Liu, Shuicheng Yan
ICCV3
2014 Supervised Hashing for Image Retrieval via Image Representation Learning
abstract
Hashing is a popular approximate nearest neighbor search approach for large-scale image retrieval. Supervised hashing, which incorporates similarity/dissimilarity information on entity pairs to improve the quality of hashing function learning, has recently received increasing attention. However, in the existing supervised hashing methods for images, an input image is usually encoded by a vector of hand-crafted visual features. Such hand-crafted feature vectors do not necessarily preserve the accurate semantic similarities of images pairs, which may often degrade the performance of hashing function learning. In this paper, we propose a supervised hashing method for image retrieval, in which we automatically learn a good image representation tailored to hashing as well as a set of hash functions. The proposed method has two stages. In the first stage, given the pairwise similarity matrix $S$ over training images, we propose a scalable coordinate descent method to decompose $S$ into a product of $HH^T$ where $H$ is a matrix with each of its rows being the approximate hash code associated to a training image. In the second stage, we propose to simultaneously learn a good feature representation for the input images as well as a set of hash functions, via a deep convolutional network tailored to the learned hash codes in $H$ and optionally the discrete class labels of the images. Extensive empirical evaluations on three benchmark datasets with different kinds of images show that the proposed method has superior performance gains over several state-of-the-art supervised and unsupervised hashing methods.
Rongkai Xia, Yan Pan 0002, Hanjiang Lai, Cong Liu 0001, Shuicheng Yan
AAAI3
2014 A new team recommendation model with applications in social network
abstract
Nowadays, some users with similar interests or same tasks may form a cooperative team in the social network, where resources and information can be wildly shared. With more and more users joining in the social network, the number of teams is growing up at the same time. How to recommend a team in which a user is interested has become a new and important topic in the social network. The purpose of this paper is to propose a novel model for team recommendation to help users find interesting teams in social network ,which is using collaborative filtering algorithm and trust propagation theory. Formally, we construct team recommendation lists from three steps. Firstly, it computes the team similarity behavior between the user and his/her friends and get the first recommendation team list from it. Secondly, it gets the team recommendation list from the user's latent friends list. Finally, the third list is got from the hot teams by applied weighting factor theory. We obtain the data from a social network called schol@t (www.scholat.com) to make experimental analysis, which demonstrate that the team recommendation model is effective.
Shaowen Hong, Chengjie Mao, Zhenxiong Yang, Hanjiang Lai
CSCWD4
2014 Efficient k-Support Matrix Pursuit
Hanjiang Lai, Yan Pan 0002, Canyi Lu, Yong Tang 0001, Shuicheng Yan
ECCV (2)1
2013 Rank Aggregation via Low-Rank and Structured-Sparse Decomposition
abstract
Rank aggregation, which combines multiple individual rank lists toobtain a better one, is a fundamental technique in various applications such as meta-search and recommendation systems. Most existing rank aggregation methods blindly combine multiple rank lists with possibly considerable noises, which often degrades their performances. In this paper, we propose a new model for robust rank aggregation (RRA) via matrix learning, which recovers a latent rank list from the possibly incomplete and noisy input rank lists. In our model, we construct a pairwise comparison matrix to encode the order information in each input rank list. Based on our observations, each comparison matrix can be naturally decomposed into a shared low-rank matrix, combined with a deviation error matrix which is the sum of a column-sparse matrix and a row-sparse one. The latent rank list can be easily extracted from the learned low-rank matrix. The optimization formulation of RRA has an element-wise multiplication operator to handle missing values, a symmetric constraint on the noise structure, and a factorization trick to restrict the maximum rank of the low-rank matrix. To solve this challenging optimization problem, we propose a novel procedure based on the Augmented Lagrangian Multiplier scheme. We conduct extensive experiments on meta-search and collaborative filtering benchmark datasets. The results show that the proposed RRA has superior performance gain over several state-of-the-art algorithms for rank aggregation.
Yan Pan 0002, Hanjiang Lai, Cong Liu 0001, Yong Tang 0001, Shuicheng Yan
AAAI2
2013 A Divide-and-Conquer Method for Scalable Low-Rank Latent Matrix Pursuit
abstract
Data fusion, which effectively fuses multiple prediction lists from different kinds of features to obtain an accurate model, is a crucial component in various computer vision applications. Robust late fusion (RLF) is a recent proposed method that fuses multiple output score lists from different models via pursuing a shared low-rank latent matrix. Despite showing promising performance, the repeated full Singular Value Decomposition operations in RLF's optimization algorithm limits its scalability in real world vision datasets which usually have large number of test examples. To address this issue, we provide a scalable solution for large-scale low-rank latent matrix pursuit by a divide-and-conquer method. The proposed method divides the original low-rank latent matrix learning problem into two size-reduced sub problems, which may be solved via any base algorithm, and combines the results from the sub problems to obtain the final solution. Our theoretical analysis shows that with fixed probability, the proposed divide-and-conquer method has recovery guarantees comparable to those of its base algorithm. Moreover, we develop an efficient base algorithm for the corresponding sub problems by factorizing a large matrix into the product of two size-reduced matrices. We also provide high probability recovery guarantees of the base algorithm. The proposed method is evaluated on various fusion problems in object categorization and video event detection. Under comparable accuracy, the proposed method performs more than $180$ times faster than the state-of-the-art baselines on the CCV dataset with about 4,500 test examples for video event detection.
Yan Pan 0002, Hanjiang Lai, Cong Liu 0001, Shuicheng Yan
CVPR2
2013 Efficient gradient descent algorithm for sparse models with application in learning-to-rank
Hanjiang Lai, Yan Pan 0002, Yong Tang 0001
Knowl. Based Syst.1
2013 Sparse Learning-to-Rank via an Efficient Primal-Dual Algorithm
abstract
Learning-to-rank for information retrieval has gained increasing interest in recent years. Inspired by the success of sparse models, we consider the problem of sparse learning-to-rank, where the learned ranking models are constrained to be with only a few nonzero coefficients. We begin by formulating the sparse learning-to-rank problem as a convex optimization problem with a sparse-inducing ℓ1constraint. Since the ℓ1constraint is nondifferentiable, the critical issue arising here is how to efficiently solve the optimization problem. To address this issue, we propose a learning algorithm from the primal dual perspective. Furthermore, we prove that, after at most O(1/ε) iterations, the proposed algorithm can guarantee the obtainment of an ε-accurate solution. This convergence rate is better than that of the popular subgradient descent algorithm. i.e., O(1/ε2). Empirical evaluation on several public benchmark data sets demonstrates the effectiveness of the proposed algorithm: 1) Compared to the methods that learn dense models, learning a ranking model with sparsity constraints significantly improves the ranking accuracies. 2) Compared to other methods for sparse learning-to-rank, the proposed algorithm tends to obtain sparser models and has superior performance gain on both ranking accuracies and training time. 3) Compared to several state-of-the-art algorithms, the ranking accuracies of the proposed algorithm are very competitive and stable.
Hanjiang Lai, Yan Pan 0002, Cong Liu 0001, Liang Lin 0004, Jie Wu 0001
IEEE Trans. Computers1
2013 FSMRank: Feature Selection Algorithm for Learning to Rank
abstract
In recent years, there has been growing interest in learning to rank. The introduction of feature selection into different learning problems has been proven effective. These facts motivate us to investigate the problem of feature selection for learning to rank. We propose a joint convex optimization formulation which minimizes ranking errors while simultaneously conducting feature selection. This optimization formulation provides a flexible framework in which we can easily incorporate various importance measures and similarity measures of the features. To solve this optimization problem, we use the Nesterov's approach to derive an accelerated gradient algorithm with a fast convergence rate O(1/T(2)). We further develop a generalization bound for the proposed optimization problem using the Rademacher complexities. Extensive experimental evaluations are conducted on the public LETOR benchmark datasets. The results demonstrate that the proposed method shows: 1) significant ranking performance gain compared to several feature selection baselines for ranking, and 2) very competitive performance compared to several state-of-the-art learning-to-rank algorithms.
Hanjiang Lai, Yan Pan 0002, Yong Tang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2011 Greedy feature selection for ranking
abstract
This paper is concerned with a study on the feature selection for ranking. Learning to rank is a useful tool for collaborative filtering and many other collaborative systems, which many algorithms have been proposed for dealing this issue. But feature selection methods receive little attention, despite of their importance in collaborative filtering problems: First, recommender systems always have massive data. Using all these data in learning to rank is unrealistic and impossible. Second, we discuss that not all the features are useful for a user's query. So choosing the most relevant data is necessary and useful. To amend this problem, we describe an algorithm called FBPCRank to choose the most relevant features for ranking. Our method combines two measures of good subsets of features, which not only can decrease the loss objective, but also reduce total similarity scores between any two features. We adopt forward and backward methods to choose the most relative features and use Pearson correlation coefficient to measure the similarity of two features. The experiments indicate that our method can outperform other state-of-the-art algorithms by selecting just small amounts of features.
Hanjiang Lai, Yong Tang 0001, Hai-Xia Luo, Yan Pan 0002
CSCWD1