VLDB 2026 Research / reviewers in the wild / expert
Yan Pan 0002
dblp:88/3327-2
· DBLP profile ↗
66ranked-venue papers
7as first author
33since 2021 · last 2026
0000-0002-0466-3763ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 2Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Inverse-free Jacobian-estimated zeroing neural network model for mechanical arm path tracking with minimal joint motion
Jielong Chen, Yan Pan 0002, Yunong Zhang, Shuai Li 0002, Min Yang 0010 |
Expert Syst. Appl. | 2 |
| 2026 | GAFD-CC: Global-aware feature decoupling with confidence calibration for out-of-distribution detection
Yongheng Xu, Jianxing Yu, Yan Pan 0002, Jian Yin 0001, Hanjiang Lai |
Neural Networks | 4 |
| 2026 | Flow-Matching Posterior Sampling: A Training-Free Conditional Generation for Flow MatchingabstractTraining-free conditional generation based on flow matching aims to leverage pre-trained unconditional flow matching models to perform conditional generation without retraining. Recently, a successful training-free conditional generation approach incorporates conditions via posterior sampling, which relies on the availability of a score function in the unconditional diffusion model. However, flow matching models lack an explicit score function, rendering this strategy inapplicable. Approximate posterior sampling for flow matching has been explored, but it is limited to linear inverse problems. In this paper, we propose Flow Matching-based Posterior Sampling (FMPS) to broaden its scope of application. We introduce a correction term by steering the velocity field. This correction term can be reformulated to incorporate a surrogate score function, thereby bridging the gap between flow matching models and score-based posterior sampling. Hence, FMPS enables posterior sampling to be adjusted within the flow-matching framework. Furthermore, we propose two practical implementations of the correction mechanism: one to improve generation quality and the other to enhance computational efficiency. Experimental results on diverse conditional generation tasks demonstrate that our method achieves superior generation quality compared to existing state-of-the-art approaches, validating the effectiveness and generality of FMPS. Kaiyu Song, Hanjiang Lai, Yan Pan 0002, Kun Yue, Jian Yin 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Enhancing Zero-Shot Adversarial Robustness of Vision-Language Models With Training-Free Adaptive Feature MovementabstractPre-trained Vision-Language Models (VLMs) have demonstrated strong zero-shot generalization capabilities. Despite their effectiveness on various downstream tasks, they remain vulnerable to adversarial samples. Existing methods fine-tune VLMs to improve their robust performance by performing adversarial training on a certain dataset. However, this can lead to model overfitting and is not a true zero-shot scenario. In this paper, we propose a truly zero-shot and training-free approach that can improve the zero-shot adversarial robustness of VLMs on the evaluated benchmarks. Specifically, we first discover that simply adding Gaussian noise can enhance the VLM's zero-shot robustness. Then, we treat the adversarial examples with added Gaussian noise as anchors and strive to find a path in the embedding space that leads from the adversarial examples to the cleaner samples. Furthermore, to avoid the overfitting issue caused by fixed hyperparameters, we propose an adaptive parameter adjustment method based on the distance between the anchors and adversarial samples in the embedding space. We largely preserve the original VLMs' zero-shot generalization abilities in a truly zero-shot and training-free manner on the evaluated benchmarks compared to previous methods. Extensive experiments on 16 datasets demonstrate that our method can achieve stronger zero-shot robust performance, improving the top-1 robust accuracy by an average of 10.83%. Baoshun Tong, Hanjiang Lai, Yan Pan 0002, Jian Yin 0001, Liang Lin 0004 |
IEEE Trans. Image Process. | 3 |
| 2026 | Zhang Neural Network Model for Time-Variant Convex Optimization Involving Nonlinear Inequality Constraints With Robotic ApplicationabstractTime-variant convex optimization involving nonlinear inequality constraints (TVCOINICs) is a challenging problem due to the nonlinearity and time-variant nature of its objective function and multitype constraints. Differing from traditional slack variables and projection operator methods, this article proposes a novel differentiable transform function that is continuous and differentiable everywhere, while eschewing the need for additional tunable hyperparameters. On the basis of the Lagrange multiplier technique and Karush–Kuhn–Tucker (KKT) conditions, the initial time-variant optimization problem is converted into a time-variant nonlinear system comprising both equalities and inequalities. By employing the proposed differentiable transform function, this system is then further refined into an equivalent time-variant nonlinear system of equalities. Subsequently, the comprehensive design process of the differentiable transform function-based Zhang neural network (DTFZNN) model is delineated, which, through the comprehensive utilization of the time derivatives and error feedback information, is capable of adeptly addressing the TVCOINICs problem. Moreover, sliding mode control (SMC) is integrated into the proposed model to endow it with verified noise-suppression capability. The relevant theorems prove its convergence and robustness, and corresponding numerical experiments verify its effectiveness. Ultimately, to evaluate the practical efficacy of the proposed model, the proposed model is implemented to address the path tracking and repetitive motion problem associated with a robotic arm, thereby illustrating its superiority in solving practical problems. Jielong Chen, Yan Pan 0002, Yunong Zhang |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | Cool-Fusion: Fuse Large Language Models without TrainingabstractWe focus on the problem of fusing two or more heterogeneous large language models (LLMs) to leverage their complementary strengths.One of the challenges of model fusion is high computational load, specifically in fine-tuning or aligning vocabularies.To address this, we propose Cool-Fusion, a simple yet effective approach that fuses the knowledge of source LLMs, which does not require training.Unlike ensemble methods, Cool-Fusion is applicable to any set of source LLMs that have different vocabularies.To overcome the vocabulary discrepancies among LLMs, we ensemble LLMs on text level, allowing them to rerank the generated texts by each other with different granularities.Extensive experiments have been conducted across a variety of benchmark datasets.On GSM8K, Cool-Fusion increases accuracy from three strong source LLMs by a significant margin of 17.4%. Cong Liu 0001, Xiaojun Quan, Yan Pan 0002, Weigang Wu, Xu Chen 0004, Liang Lin 0004 |
ACL (1) | 3 |
| 2025 | CISTF: A Causal Inference Method for Stock Trend ForecastingabstractStock trend forecasting, which aims to predict future fluctuations of stock price, has garnered significant attention in recent years. However, it is quite hard to train a model that can consistently and accurately forecasts the future movements of stocks. Since stock data is influenced by various unobservable factors, such as market sentiment and industry conditions, its distribution is unstable and can change over time. As a result, models trained on stock data often exhibit poor performance in prediction. Causal inference is an effective method that has been widely used to address distribution shift. However, how to conduct causal inference for stock trend forecasting still remains under-explored. To this end, we propose CISTF, a method based on causal inference to explore the causal dependencies between the input historical data and the future trend of stocks. We first construct a causal graph for the stock trend forecasting task, in which we use confounders to represent unobservable factors that affect both the input data and the future trend of stocks. Such time-varying confounders contribute to the distribution drift in the stock data. Then we use front-door adjustment, an intervention strategie that allow we to intervene the input data, to eliminate the spurious correlations brought by the unobserved confounders. Additionally, we develop a deep architecture to implement front-door adjustment, which consists of a feature extractor, a mediator estimation module and a conditional probability estimation module. We evaluate the effectiveness of CISTF on real-world stock data, and experiments on three public stock datasets demonstrate that our model achieves state-of-the-art performance. Rongbang Qiu, Qinkang Gong, Yan Pan 0002, Hanjiang Lai |
CSCWD | 3 |
| 2025 | On the Zero-shot Adversarial Robustness of Vision-Language Models: A Truly Zero-shot and Training-free ApproachabstractPre-trained Vision-Language Models (VLMs) like CLIP, have demonstrated strong zero-shot generalization capabilities. Despite their effectiveness on various downstream tasks, they remain vulnerable to adversarial samples. Existing methods fine-tune VLMs to improve their performance via performing adversarial training on a certain dataset. However, this can lead to model overfitting and is not a true zero-shot scenario. In this paper, we propose a truly zero-shot and training-free approach that can significantly improve the VLM’s zero-shot adversarial robustness. Specifically, we first discover that simply adding Gaussian noise greatly enhances the VLM’s zero-shot performance. Then, we treat the adversarial examples with added Gaussian noise as anchors and strive to find a path in the embedding space that leads from the adversarial examples to the cleaner samples. We improve the VLMs’ generalization abilities in a truly zero-shot and training-free manner compared to previous methods. Extensive experiments on 16 datasets demonstrate that our method can achieve state-of-the-art zero-shot robust performance, improving the top-1 robust accuracy by an average of 9.77%. The code will be publicly available. Baoshun Tong, Hanjiang Lai, Yan Pan 0002, Jian Yin 0001 |
CVPR | 3 |
| 2025 | Deep Hashing Based on Feature Fusion Enhancing Hierarchical Transformer with Distance Separated Centers Guided Polarization Loss
Yan Pan 0002, Jianfeng Cheng, Jian Yin 0001 |
ICIC (3) | 2 |
| 2025 | Few-shot Prompt Learning with Large Vision-Language Model for Image Deep HashingabstractThe image hashing algorithms can convert high-dimensional features into low-dimensional hash codes, and have a wide range of application scenarios in information retrieval and retrieval augmented generation. The current mainstream image hashing methods leverage the deep hashing framework, typically employing a backbone network model, and adjusting it to a specific image data domain through the pre-training and fine-tuning paradigm. However, it is often necessary to reselect an appropriate backbone network for fine-tuning when generating hash codes for cross-domain data with existing methods. Concurrently, the fine-tuning process requires a substantial amount of new domain samples and labels, along with updating all parameters of the backbone network, which incurs higher costs for manual data processing and labeling, as well as server computing time. To address these challenges, we introduce the large vision-language model based on the prompt-based paradigm into the deep hashing field, and redesign the deep hashing framework. The novel proposed framework can achieve better results on cross-domain image data with only a few samples and labels, while just updating part of the network parameters. Through comparative experiments on multiple datasets, it can be verified that the proposed method achieves state-of-the-art effect in the few-shot image hashing scenario, particularly when generating hash codes with a minimal number of bits. Yan Pan 0002, Jian Yin 0001 |
ICME | 2 |
| 2025 | A Causal Intervention Method for Domain Generalization with a Self-Supervised Auxiliary Task
Qinkang Gong, Yan Pan 0002, Hanjiang Lai, Jian Yin 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Discrete Jacobian-Pseudoinverse-Free Zhang Neurodynamics Algorithm Handling Path Tracking of Robot Manipulator With Unknown ModelabstractRobot manipulator path tracking, recognized as a crucial aspect in robot manipulator control, has garnered significant attention from researchers. In this paper, to address the path tracking problem of robot manipulators with unknown models, a novel Jacobian pseudoinverse estimator is first proposed based on Zhang neurodynamics method. The estimator directly provides an efficient and accurate estimation of the Jacobian matrix pseudoinverse, avoiding the complicated operation of matrix pseudoinverse and preventing potential singularity phenomenon of the Jacobian matrix. By utilizing the Euler difference formulas, a discrete model-free and Jacobian-pseudoinverse-free Zhang neurodynamics algorithm is proposed. The proposed algorithm focuses on leveraging the available current and previous known information to predict the future unknown information. Detailed theoretical analyses and proofs ensure the convergence and stability of the proposed algorithm. Finally, comparative experiments with various effective model-free algorithms, and experimental validations on different types of robot manipulators (UR5, Franka Emika Panda, and Kinova Gen3 robot manipulators) using various experimental platforms (MATLAB, CoppeliaSim, and physical platforms) illustrate the effectiveness of the proposed algorithm. Note to Practitioners—This paper is motivated by addressing the prevalent challenge of unknown models in real-time path tracking for robot manipulators. In this paper, a novel discrete model-free and Jacobian-pseudoinverse-free Zhang neurodynamics algorithm is proposed. Different from the existing model-free algorithms, the proposed algorithm avoids the complicated operation of computing the pseudoinverse of matrix without compromising precision, significantly reducing the computational complexity and preventing potential singularity phenomenon of the Jacobian matrix. The average computation time per updating for the proposed algorithm is approximately$0.1\,{\mathrm { ms}}$, which is significantly less than the sampling gap of the operation. This allows it to effectively meet the real-time requirements for robot manipulator path tracking. In addition, the error of the proposed algorithm is approximately$2\,\mu {\mathrm { m}}$, which can meet the requirements of most practical application scenarios. Moreover, the accuracy of the algorithm is limited by the differential formula and sampling gap. Improving the accuracy and robustness of the algorithm by using more accurate difference formulas and filtering technique will be our future research direction. Jielong Chen, Yan Pan 0002, Yunong Zhang, Ning Tan 0003 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Efficient Multimodal Selection for Retrieval in Knowledge-Based Visual Question AnsweringabstractRetrieval plays an important role in knowledge-based visual question answering (KB-VQA), which relies on external knowledge to answer questions related to an image. However, not all information in the external knowledge is beneficial in retrieval, e.g., the knowledge that is only semantically similar to the query but is not useful for question answering. To improve the effectiveness and efficiency of retrieval, in this paper, we propose efficient multimodal selection to filter out irrelevant information and increase the retriever performance for KB-VQA. First, to exclude most irrelevant knowledge from the large external knowledge, multimodal selection uses a query-aware sample selection method, which uses the pretrained answer generator’s prediction to obtain better positive and negative training samples to help retrievers distinguish knowledge that is semantically relevant to the multimodal query. Then, question-aware visual feature selection is proposed to select the distinguishable visual information related to the question: where cross-attention to questions and images is proposed to obtain question-aware visual features. These visual features are used to perform fine-grained multimodal retrieval within the small set to obtain the final top-related knowledge. The experimental results show that the proposed approach achieves state-of-the-art retrieval performance on the OK-VQA and FVQA datasets, indicating the effectiveness of our selection strategy for retrieval. Linyin Luo, Hanjiang Lai, Yan Pan 0002, Jian Yin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Causal-TSF: A Causal Intervention Approach to Mitigate Confounding Bias in Time Series ForecastingabstractTime series forecasting, aiming to learn models from historical data and predict future values in time series, is a fundamental research topic in machine learning. However, few efforts have been devoted to addressing the confounding effects in time series data, e.g., the historical data are affected by some hidden surrounding factors (i.e., confounders), leading to biased forecasting models for future data. This paper presents a causal intervention approach to eliminate the bias that is raised by some hidden confounders. By using a causal graph, we illustrate why hidden confounders can bring bias in time series forecasting and how to tackle it. We implement causal intervention by a deep architecture that consists of two modules, a Confounders Estimation module to estimate the hidden confounders and a Debiasing module to eliminate the confounding bias in the forecasting model via sampling on confounders. We conduct comprehensive evaluations on various time series datasets. The experiment results indicate that the proposed method can reduce the negative confounding effects in time series data, and it achieves superior gains over state-of-the-art baselines for time series forecasting. Qinkang Gong, Yan Pan 0002, Hanjiang Lai, Rongbang Qiu, Jian Yin 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Model-Free and Pseudoinverse-Free Zhang Neurodynamics Scheme for Robotic Arms' Path Tracking ControlabstractPath tracking control of robotic arms is regarded as a fundamental problem in the field of robotics. However, obtaining an accurate model of the robotic arm in practical engineering poses significant challenges. As a result, model-free schemes have become a focus of investigation. In contrast to traditional model-free schemes used for estimating the Jacobian matrix of the robotic arm, in this work, a novel estimator directly for the pseudoinverse (PI) of the Jacobian matrix based on Zhang neurodynamics (ZN) is proposed for the first time. In addition, a novel model-free and PI-free ZN (MFPIFZN) scheme for path tracking control of robotic arms is proposed. The MFPIFZN scheme not only significantly reduces the operation complexity by eliminating the requirement to compute the PI of the Jacobian matrix but also enhances the accuracy by eliminating the potential errors that may arise from the computation of the PI. Theoretical analyses provide guarantees for the convergence and stability of the MFPIFZN scheme. Finally, experimental results conducted on planar four-link and Kinova Jaco2 robotic arms vividly illustrate the excellent performance of the MFPIFZN scheme. Comparison experiments with four other model-free schemes further confirm the superiority of the MFPIFZN scheme. Jielong Chen, Yan Pan 0002, Yunong Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Deep Hashing with Triplet Loss of Hash Centers and Dissimilar Pairs for Image RetrievalabstractImage hash coding learning based on deep neural network have a wide range of application scenarios and become a research hotspot in recent years. In the framework of most existing deep hashing methods, it is necessary to first design a loss function that can converge quickly, and then select an appropriate pre-training deep network and tune it. In order to balance global and local constraints, this paper combines the design of loss functions required by the processes of image feature learning and image hash coding, and adds the constraints of triple loss of hash centers and dissimilar pairs. The proposed novel loss function can give consideration to the characteristics of global hash centers, local dissimilarity of images and image classification labels in the training process. The pre-trained large models have been incorporated into the framework, which can extract image embeddings to assist in the generation of hash codes. The comparative experiments on several image datasets show that the DTLH method proposed in this paper can achieve better results than the traditional hashing methods and the deep hashing methods under different coding lengths. Yan Pan 0002, Jian Yin 0001 |
CSCWD | 2 |
| 2024 | MimicDiffusion: Purifying Adversarial Perturbation via Mimicking Clean Diffusion ModelabstractDeep neural networks (DNNs) are vulnerable to adversarial perturbation, where an imperceptible perturbation is added to the image that can fool the DNNs. Diffusion-based adversarial purification uses the diffusion model to generate a clean image against such adversarial attacks. Un-fortunately, the generative process of the diffusion model is also inevitably affected by adversarial perturbation since the diffusion model is also a deep neural network where its input has adversarial perturbation. In this work, we propose MimicDiffusion, a new diffusion-based adversarial purification technique that directly approximates the generative process of the diffusion model with the clean image as input. Concretely, we analyze the differences between the guided terms using the clean image and the ad-versarial sample. After that, we first implement MimicD-iffusion based on Manhattan distance. Then, we propose two guidance to purify the adversarial perturbation and ap-proximate the clean diffusion model. Extensive experiments on three image datasets, including CIFAR-10, CIFAR-100, and ImageNet, with three classifier backbones including WideResNet-70-16, WideResNet-28-10, and ResNet-50 demonstrate that MimicDiffusion significantly performs better than the state-of-the-art baselines. On CIFAR -10, CIFAR-100, and ImageNet, it achieves 92.67%, 61.35%, and 61.53% average robust accuracy, which are 18.49%, 13.23%, and 17.64% higher, respectively. The code is available at https://github.com/psky1111/MimicDiffusion. Kaiyu Song, Hanjiang Lai, Yan Pan 0002, Jian Yin 0001 |
CVPR | 3 |
| 2024 | MALIP: Improving Few-Shot Image Classification with Multimodal Fusion EnhancementabstractWith the significant progress in pre-trained visionlanguage models like CLIP, recent CLIP-based methods have shown impressive performance in few-shot tasks. However, CLIPbased representations have a natural gap in downstream few-shot tasks due to the label-related multimodal information scarcity caused by limited data. We then question, whether the generative model trained in downstream tasks could be used to enhance label-related multimodal fusion. In this paper, we propose a generative model-based multimodal fusion enhancement method, MALIP, to improve the few-shot performance of CLIP via a Multimodal Adapter module. Specifically, we first leverage a variational autoencoder (VAE) that could be trained in the few-shot scenario to extend the data. Then we create adapter weights by a key-value cache model constructed from the image and text information based on the expanded data. In the end, through extensive experiments on 11 datasets, we demonstrate the effectiveness of MALIP to perform state-of-the-art few-shot image classification. Kaifen Cai, Kaiyu Song, Yan Pan 0002, Hanjiang Lai |
ICME | 3 |
| 2024 | Long-Tailed Hashing with Wasserstein Quantization
Zujun Fu, Hanjiang Lai, Yan Pan 0002 |
ICPR (21) | 3 |
| 2024 | MAMixer: Multivariate Time Series Forecasting via Multi-axis Mixing
Yongyu Liu, Guoliang Lin, Hanjiang Lai, Yan Pan 0002 |
MMM (1) | 4 |
| 2024 | Enhancing Event Tagger with Automatic Speech Recognizer for Audio Multi-task Scenarios by Distillation with Pre-Trained Large ModelsabstractWith the continuous expansion of robotics and digital humans in practical applications, the demand for the auditory system is becoming deeper, usually requiring more efficient speech recognition framework capabilities to handle multiple tasks and use fewer resources. In previous audio processing frameworks, each type of audio processing task typically requires constructing a standalone deep network model for training, which results in more training data and higher training time when constructing models for multi-task audio scenarios simultaneously. The recent improvement of audio models based on transformers have brought about methods that can handle multiple audio tasks concurrently. However, recent related methods still require retraining multi-task targets with an amount of data, and achieve the general effect after training for multi-task scenarios than the simple combination of standalone methods processed separately. In order to better build a model that can handle multiple audio tasks, we propose a novel framework of distillation through pre-trained large models for enhancing event tagger with automatic speech recognizer. Through multiple rounds of experiments on several audio datasets, it has been verified that the proposed framework can achieve better results than the baseline for multitasking, and comparative results with less parameters compared to the baselines for single-task scenarios. Jianfeng Cheng, Jian Yin 0001, Liangdao Wang, Yan Pan 0002 |
SMC | 5 |
| 2024 | Vision Transformer Based Hash Coding for Efficient Image and Audio Retrieval with Global and Local Equilibrium Distance ConstraintsabstractWith the application of efficient retrieval in information systems and retrieval augmented generation with vector database for large language models, hash coding algorithms have made progress in recent years. The rise of transformer technology in the field of deep learning has brought the possibility to further improve the effect of hash coding algorithms. We introduce the vision transformer framework to both images and audios, and propose a novel approach for the tasks of multi-label image retrieval and audio event retrieval. In the proposed hash coding model, global and local equilibrium distance constraints are integrated, so that the hash codes for images can be better obtained through the global hash centers and local similar samples. In order to realize end-to-end training and hash code generation for audios, we adopt the adapter of mel spectrogram, thus the proposed approach can be simply converted and applied to audio hash coding. Comparative experiments verify that better results can be achieved on multiple image and audio datasets. Yan Pan 0002, Jian Yin 0001 |
SMC | 2 |
| 2024 | Inverse-free zeroing neural network for time-variant nonlinear optimization with manipulator applications
Jielong Chen, Yan Pan 0002, Yunong Zhang, Shuai Li 0002, Ning Tan 0003 |
Neural Networks | 2 |
| 2024 | Enhancing Multi-Label Deep Hashing for Image and Audio With Joint Internal Global Loss Constraints and Large Vision-Language ModelabstractDeep hashing algorithms can transform high-dimensional features into low-dimensional hash codes, which can reduce storage space and improve computational efficiency in traditional information retrieval (IR) and large model related retrieval augmented generation (RAG) scenarios. In recent years, pre-trained convolutional or transformer networks are commonly chosen as the backbone in deep hashing frameworks. This involves incorporating local loss constraints among training samples, and then fine-tuning the model to generate hash codes. Due to the relatively limited local information of constraints among training samples, we propose to design the novel anchor constraint and structural constraint as internal global loss constraints with the vision transformer network, and augment external information by integrating the large vision-language model, thereby enhancing the performance of hash code generation. Additionally, to enhance the scalability of the novel deep hashing framework, we propose to incorporate the adapter module to extend its application from the image domain to the audio domain. By conducting comparative experiments and ablation analysis on various image and audio datasets, it can be confirmed that the proposed method achieves state-of-the-art retrieval results. Yan Pan 0002, Jian Yin 0001 |
IEEE Signal Process. Lett. | 2 |
| 2024 | ZNN Continuous Model and Discrete Algorithm for Temporally Variant Optimization With Nonlinear Equation Constraints via Novel TD FormulaabstractFor dealing with the temporally variant optimization with nonlinear equation constraints (TVONECs), a novel Zhang neural net (ZNN) model is proposed in this work. Two continuous-time computer numerical simulations are constructed to testify the feasibility and correctness of the continuous-time ZNN (CZNN) model. To facilitate the implementation of numerical algorithms on computer, a novel 11-instant time discretization (TD) formula is proposed in this article, and a discrete-time ZNN (DZNN) algorithm (i.e., 11-instant DZNN algorithm) is thus obtained. Besides, theoretical analyses prove the superiority as well as the feasibility of the DZNN11I algorithm. For comparison, other three TD formulas and corresponding discrete-time algorithms (i.e., 2-instant DZNN, 3-instant DZNN, and 7-instant DZNN algorithms) are presented. Finally, numerical experiments and an application to Kinova Jaco2 manipulator control are conducted to illustrate the superiority of the proposed model and algorithm. Jielong Chen, Yan Pan 0002, Yunong Zhang |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2023 | Deep Hashing with Minimal-Distance-Separated Hash CentersabstractDeep hashing is an appealing approach for large-scale image retrieval. Most existing supervised deep hashing methods learn hash functions using pairwise or triple image similarities in randomly sampled mini-batches. They suffer from low training efficiency, insufficient coverage of data distribution, and pair imbalance problems. Recently, central similarity quantization (CSQ) attacks the above problems by using “hash centers” as a global similarity metric, which encourages the hash codes of similar images to approach their common hash center and distance themselves from other hash centers. Although achieving SOTA retrieval performance, CSQ falls short of a worst-case guarantee on the minimal distance between its constructed hash centers, i.e. the hash centers can be arbitrarily close. This paper presents an optimization method that finds hash centers with a constraint on the minimal distance between any pair of hash centers, which is non-trivial due to the non-convex nature of the problem. More importantly, we adopt the Gilbert-Varshamov bound from coding theory, which helps us to obtain a large minimal distance while ensuring the empirical feasibility of our optimization approach. With these clearly-separated hash centers, each is assigned to one image class, we propose several effective loss functions to train deep hashing networks. Extensive experiments on three datasets for image retrieval demonstrate that the proposed method achieves superior retrieval performance over the state-of-the-art deep hashing methods. Liangdao Wang, Yan Pan 0002, Cong Liu 0001, Hanjiang Lai, Jian Yin 0001 |
CVPR | 2 |
| 2023 | Multi-label Image Deep Hashing with Hybrid Loss of Global Center and Local Alignment
Yan Pan 0002, Jian Yin 0001 |
ICANN (10) | 2 |
| 2023 | Deep Hashing for Multi-label Image Retrieval with Similarity Matrix Optimization of Hash Centers and Anchor Constraint of Center Pairs
Yan Pan 0002, Jian Yin 0001 |
ICONIP (13) | 2 |
| 2023 | Efficient Hash Coding for Image Retrieval Based on Improved Center Generation and Contrastive Pre-training Knowledge Model
Yan Pan 0002, Jian Yin 0001 |
KSEM (4) | 2 |
| 2022 | Image Retrieval with Well-Separated Semantic Hash Centers
Liangdao Wang, Yan Pan 0002, Hanjiang Lai, Jian Yin 0001 |
ACCV (6) | 2 |
| 2022 | Multiview Spectral Clustering via Robust Subspace SegmentationabstractMultiview clustering refers to partition data according to its multiple views, where information from different perspectives can be jointly used in some certain complementary manner to produce more sensible clusters. It is believed that most of the existing multiview clustering methods technically suffer from possibly corrupted data, resulting in a dramatically decreased clustering performance. To overcome this challenge, we propose a multiview spectral clustering method based on robust subspace segmentation in this article. Our proposed algorithm is composed of three modules, that is: 1) the construction of multiple feature matrices from all views; 2) the formulation of a shared low-rank latent matrix by a low rank and sparse decomposition; and 3) the use of the Markov-chain-based spectral clustering method for producing the final clusters. To solve the optimization problem for a low rank and sparse decomposition, we develop an optimization procedure based on the scheme of the augmented Lagrangian method of multipliers. The experimental results on several benchmark datasets indicate that the proposed method outperforms favorably compared to several state-of-the-art multiview clustering techniques. Yan Pan 0002, Changqin Huang, Dianhui Wang 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Deep Listwise Triplet Hashing for Fine-Grained Image RetrievalabstractHashing is a practical approach for the approximate nearest neighbor search. Deep hashing methods, which train deep networks to generate compact and similarity-preserving binary codes for entities (e.g. images), have received lots of attention in the information retrieval community. A representative stream of deep hashing methods is triplet-based hashing that learns hashing models from triplets of data. The existing triplet-based hashing methods only consider triplets that are in the form of$(q,q^{+},q^{-})$, where$q$and$q^{+}$are in the same class and$q$and$q^{-}$are in different classes. However, the number of possible triplets is approximately the cube of training examples, triplets used in the existing methods are only a small fraction of all possible triplets. This motivates us to develop a new triplet-based hashing method that adopts many more triplets in training phase. We propose Deep Listwise Triplet Hashing (DLTH) that introduces more triplets into batch-based training and a novel listwise triplet loss to capture the relative similarity in new triplets. This method has a pipeline of two steps. In Step 1, we propose a novel way to generate triplets from the soft class labels obtained by knowledge distillation module, where the triplets in the form of$(q,q^{+},q^{-})$are a subset of the newly obtained triplets. In Step 2, we develop a novel listwise triplet loss to train the hashing network, which seeks to capture the relative similarity between images in triplets according to soft labels. We conduct comprehensive image retrieval experiments on four benchmark datasets. The experimental results show that the proposed method has superior performances over state-of-the-art baselines. Yan Pan 0002, Hanjiang Lai, Wei Liu 0061, Jian Yin 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Variable-Length Metric Learning for Fast Image RetrievalabstractLearning a powerful distance metric is the key component of image retrieval. Recently, deep metric learning has been an active research topic for image retrieval. However, most existing metric learning approaches treat all the input images equally and learn all image embeddings at equal lengths. These methods ignore those easy examples that can be encoded as the shorter features, which is search-inefficient. We propose a simple but efficient variable-length metric learning method for search efficiency, in which the different query samples have different feature-length. First, we propose to learn the ranked (prioritized) list of features. The more distinguishing features, the higher the rank. We show that the proposed prioritized features can be used to perform fast retrieval in different feature-length configurations. Further, we propose an adaptive feature-length selection policy to determine the amount of feature-length for each query sample. Extensive experiments are conducted on three benchmark datasets. The results demonstrate that the proposed method can reduce computational costs without incurring a decrease in accuracy. Hanjiang Lai, Yan Pan 0002 |
CSCWD | 3 |
| 2020 | Robust multi-view clustering via inter-and-intra-view low rank fusion
Yan Pan 0002, Hanjiang Lai, Jian Yin 0001 |
Neurocomputing | 2 |
| 2020 | Improving Deep Binary Embedding Networks by Order-Aware Reweighting of TripletsabstractIn this paper, we focus on triplet-based deep binary embedding networks for image retrieval task. The triplet loss has been shown to be effective for hashing retrieval. However, most of the triplet-based deep networks treat the triplets equally or select the hard triplets based on the loss. Such strategies do not consider the order relations of the binary codes and ignore the hash encoding when learning the feature representations. To this end, we propose an order-aware reweighting method to effectively train the triplet-based deep networks, which up-weights the important triplets and down-weights the uninformative triplets via the rank lists of the binary codes. First, we present the order-aware weighting factors to indicate the importance of the triplets, which depend on the rank order of binary codes. Then, we reshape the triplet loss to the squared triplet loss such that the loss function will put more weights on the important triplets. The extensive evaluations on several benchmark datasets show that the proposed method achieves significant performance compared with the state-of-the-art baselines. Hanjiang Lai, Jikai Chen, Libing Geng, Yan Pan 0002, Xiaodan Liang, Jian Yin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Feature Pyramid HashingabstractIn recent years, deep-networks-based hashing has become a leading approach for large-scale image retrieval. Most deep hashing approaches use the high layer to extract the powerful semantic representations. However, these methods have limited ability for fine-grained image retrieval because the semantic features extracted from the high layer are difficult in capturing the subtle differences. To this end, we propose a novel two-pyramid hashing architecture to learn both the semantic information and the subtle appearance details for fine-grained image search. Inspired by the feature pyramids of convolutional neural network, avertical pyramid is proposed to capture the high-layer features and ahorizontal pyramid combines multiple low-layer features with structural information to capture the subtle differences. To fuse the low-level features, a novel combination strategy, called consensus fusion, is proposed to capture all subtle information from several low-layers for finer retrieval. Extensive evaluation on two fine-grained datasets CUB-200-2011 and Stanford Dogs demonstrate that the proposed method achieves significant performance compared with the state-of-art baselines. Libing Geng, Hanjiang Lai, Yan Pan 0002, Jian Yin 0001 |
ICMR | 4 |
| 2019 | Improved Search in Hamming Space Using Deep Multi-Index HashingabstractSimilarity-preserving hashing is a widely used method for nearest neighbor search in large-scale image retrieval tasks. Considerable research has been conducted on deep-network-based hashing approaches to improve the performance. However, the binary codes generated from deep networks may be not uniformly distributed over the Hamming space, which will greatly increase the retrieval time. To this end, we propose a deep-network-based multi-index hashing (MIH) for retrieval efficiency. We first introduce the MIH mechanism into the proposed deep architecture, which divides the binary codes into multiple substrings. Each substring corresponds to one hash table. Then, we add the two balanced constraints to obtain more uniformly distributed binary codes: 1) balanced substrings, where the Hamming distances of each substring are equal for any two binary codes and 2) balanced hash buckets, where the sizes of each bucket are equal. Extensive evaluations on several benchmark image retrieval data sets show that the learned balanced binary codes bring dramatic speedups and achieve comparable performance over the existing baselines. Hanjiang Lai, Yan Pan 0002, Si Liu 0001, Zhenbin Weng, Jian Yin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Clothing Landmark Detection Using Deep Networks With Prior of Key Point AssociationsabstractThis paper considers a problem of landmark point detection in clothes, which is important and valuable for clothing industry. A novel method for landmark localization has been proposed, which is based on a deep end-to-end architecture using prior of key point associations. With the estimated landmark points as input, a deep network has been proposed to predict clothing categories and attributes. A systematic design of the proposed detecting system is implemented by using deep learning techniques and a large-scale clothes dataset containing 145 000 upper-body clothing images with landmark annotations. Experimental results indicate that clothing categories and attributes can be well classified by using the detected landmark points, which are associated with regions of interest in clothes (e.g., the sleeves and the collars) and share robust learning representation property with respect to large variances of human poses, nonfrontal views, or occlusion. A comprehensive performance evaluation over two newly released datasets is carried out in this paper, showing that the proposed system with deep architecture for clothing landmark detection outperforms the state-of-the-art techniques. Changqin Huang, Jikai Chen, Yan Pan 0002, Hanjiang Lai, Jian Yin 0001, Qionghao Huang |
IEEE Trans. Cybern. | 3 |
| 2018 | EGRank: An exponentiated gradient algorithm for sparse learning-to-rank
Yan Pan 0002, Jintang Ding, Hanjiang Lai, Changqin Huang |
Inf. Sci. | 2 |
| 2018 | Deep Recurrent Regression for Facial Landmark DetectionabstractWe propose a novel end-to-end deep architecture for face landmark detection, based on a deep convolutional and deconvolutional network followed by carefully designed recurrent network structures. The pipeline of this architecture consists of three parts. Through the first part, we encode an input face image to resolution-preserved deconvolutional feature maps via a deep network with stacked convolutional and deconvolutional layers. Then, in the second part, we estimate the initial coordinates of the facial key points by an additional convolutional layer on top of these deconvolutional feature maps. In the last part, by using the deconvolutional feature maps and the initial facial key points as input, we refine the coordinates of the facial key points by a recurrent network that consists of multiple long short-term memory components. Extensive evaluations on several benchmark data sets show that the proposed deep architecture has superior performance against the state-of-the-art methods. Hanjiang Lai, Shengtao Xiao, Yan Pan 0002, Zhen Cui 0001, Jiashi Feng, Chunyan Xu, Jian Yin 0001, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Optimizing Evaluation Metrics for Multitask Learning via the Alternating Direction Method of MultipliersabstractMultitask learning (MTL) aims to improve the generalization performance of multiple tasks by exploiting the shared factors among them. Various metrics (e.g., -score, area under the ROC curve) are used to evaluate the performances of MTL methods. Most existing MTL methods try to minimize either the misclassified errors for classification or the mean squared errors for regression. In this paper, we propose a method to directly optimize the evaluation metrics for a large family of MTL problems. The formulation of MTL that directly optimizes evaluation metrics is the combination of two parts: 1) a regularizer defined on the weight matrix over all tasks, in order to capture the relatedness of these tasks and 2) a sum of multiple structured hinge losses, each corresponding to a surrogate of some evaluation metric on one task. This formulation is challenging in optimization because both of its parts are nonsmooth. To tackle this issue, we propose a novel optimization procedure based on the alternating direction scheme of multipliers, where we decompose the whole optimization problem into a subproblem corresponding to the regularizer and another subproblem corresponding to the structured hinge losses. For a large family of MTL problems, the first subproblem has closed-form solutions. To solve the second subproblem, we propose an efficient primal-dual algorithm via coordinate ascent. Extensive evaluation results demonstrate that, in a large family of MTL problems, the proposed MTL method of directly optimization evaluation metrics has superior performance gains against the corresponding baseline methods. Ge-Yang Ke, Yan Pan 0002, Jian Yin 0001, Changqin Huang |
IEEE Trans. Cybern. | 2 |
| 2018 | Object-Location-Aware Hashing for Multi-Label Image Retrieval via Automatic Mask LearningabstractLearning-based hashing is a leading approach of approximate nearest neighbor search for large-scale image retrieval. In this paper, we develop a deep supervised hashing method for multi-label image retrieval, in which we propose to learn a binary "mask" map that can identify the approximate locations of objects in an image, so that we use this binary "mask" map to obtain length-limited hash codes which mainly focus on an image's objects but ignore the background. The proposed deep architecture consists of four parts: 1) a convolutional sub-network to generate effective image features; 2) a binary "mask" sub-network to identify image objects' approximate locations; 3) a weighted average pooling operation based on the binary "mask" to obtain feature representations and hash codes that pay most attention to foreground objects but ignore the background; and 4) the combination of a triplet ranking loss designed to preserve relative similarities among images and a cross entropy loss defined on image labels. We conduct comprehensive evaluations on four multi-label image data sets. The results indicate that the proposed hashing method achieves superior performance gains over the state-of-the-art supervised or unsupervised hashing baselines. Changqin Huang, Shang-Ming Yang, Yan Pan 0002, Hanjiang Lai |
IEEE Trans. Image Process. | 3 |
| 2017 | Multi-view Spectral Clustering via Tensor-SVD DecompositionabstractMulti-view clustering has attracted considerable attention in recent years, some related approaches always use matrices to represent views, and model by capturing two dimensional structure among views. The critical deficiency of these work is ignoring the space structure information of all views, which results in the mediocre performance of clustering. In this paper, we propose a novel Tensor-SVD decomposition based Multi-view Spectral Clustering algorithm(TMSC) to iron out flaws. Our method firstly puts transition probability matrices of all views into a three-order tensor, which naturally reserves the whole structure information of data. Then it establishes a low multi-rank tensor model based on tensor-SVD decomposition by fully mining the complementary information among multiple views. Another difficulty in this paper is that the optimal objective of TMSC has a low multirank constraint on the transition probability tensor, and a probabilistic simplex constraint on each fiber of the tensor. To tackle this challenging problem, we design an optimization procedure based on the Augmented Lagrangian Multiplier scheme. Experimental results on real word datasets show that TMSC has superior clustering quality over several state-of-theart multi-view clustering approaches. Bangtian Liu, Ge-Yang Ke, Yan Pan 0002, Jian Yin 0001 |
ICTAI | 5 |
| 2016 | Modelling Sentence Pairs with Tree-structured Attentive EncoderabstractWe describe an attentive encoder that combines tree-structured recursive neural networks and sequential recurrent neural networks for modelling sentence pairs. Since existing attentive models exert attention on the sequential structure, we propose a way to incorporate attention into the tree topology. Specially, given a pair of sentences, our attentive encoder uses the representation of one sentence, which generated via an RNN, to guide the structural encoding of the other sentence on the dependency parse tree. We evaluate the proposed attentive encoder on three tasks: semantic similarity, paraphrase identification and true-false question selection. Experimental results show that our encoder outperforms all baselines and achieves state-of-the-art results on two tasks. Cong Liu 0001, Yan Pan 0002 |
COLING | 3 |
| 2016 | Margin-based two-stage supervised hashing for image retrieval
Yan Pan 0002, Hanjiang Lai, Cong Liu 0001, Jian Yin 0001 |
Neurocomputing | 2 |
| 2015 | Simultaneous feature learning and hash coding with deep neural networksabstractSimilarity-preserving hashing is a widely-used method for nearest neighbour search in large-scale image retrieval tasks. For most existing hashing methods, an image is first encoded as a vector of hand-engineering visual features, followed by another separate projection or quantization step that generates binary codes. However, such visual feature vectors may not be optimally compatible with the coding process, thus producing sub-optimal hashing codes. In this paper, we propose a deep architecture for supervised hashing, in which images are mapped into binary codes via carefully designed deep neural networks. The pipeline of the proposed deep architecture consists of three building blocks: 1) a sub-network with a stack of convolution layers to produce the effective intermediate image features; 2) a divide-and-encode module to divide the intermediate image features into multiple branches, each encoded into one hash bit; and 3) a triplet ranking loss designed to characterize that one image is more similar to the second image than to the third one. Extensive evaluations on several benchmark image datasets show that the proposed simultaneous feature learning and hash coding pipeline brings substantial improvements over other state-of-the-art supervised or unsupervised hashing methods. Hanjiang Lai, Yan Pan 0002, Shuicheng Yan |
CVPR | 2 |
| 2015 | A Divide-and-Conquer Method for Scalable Robust Multitask LearningabstractMultitask learning (MTL) aims at improving the generalization performance of multiple tasks by exploiting the shared factors among them. An important line of research in the MTL is the robust MTL (RMTL) methods, which use trace-norm regularization to capture task relatedness via a low-rank structure. The existing algorithms for the RMTL optimization problems rely on the accelerated proximal gradient (APG) scheme that needs repeated full singular value decomposition (SVD) operations. However, the time complexity of a full SVD is O(min(md(2),m(2)d)) for an RMTL problem with m tasks and d features, which becomes unaffordable in real-world MTL applications that often have a large number of tasks and high-dimensional features. In this paper, we propose a scalable solution for large-scale RMTL, with either the least squares loss or the squared hinge loss, by a divide-and-conquer method. The proposed method divides the original RMTL problem into several size-reduced subproblems, solves these cheaper subproblems in parallel by any base algorithm (e.g., APG) for RMTL, and then combines the results to obtain the final solution. Our theoretical analysis indicates that, with high probability, the recovery errors of the proposed divide-and-conquer algorithm are bounded by those of the base algorithm. Furthermore, in order to solve the subproblems with the least squares loss or the squared hinge loss, we propose two efficient base algorithms based on the linearized alternating direction method, respectively. Experimental results demonstrate that, with little loss of accuracy, our method is substantially faster than the state-of-the-art APG algorithms for RMTL. Yan Pan 0002, Rongkai Xia, Jian Yin 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Robust Multi-View Spectral Clustering via Low-Rank and Sparse DecompositionabstractMulti-view clustering, which seeks a partition of the data inmultiple views that often provide complementary information to eachother, has received considerable attention in recent years. In reallife clustering problems, the data in each view may haveconsiderable noise. However, existing clustering methods blindlycombine the information from multi-view data with possiblyconsiderable noise, which often degrades their performance. In thispaper, we propose a novel Markov chain method for RobustMulti-view Spectral Clustering (RMSC). Our method has a flavor oflow-rank and sparse decomposition, where we firstly construct atransition probability matrix from each single view, and then usethese matrices to recover a shared low-rank transition probabilitymatrix as a crucial input to the standard Markov chain methodfor clustering. The optimization problem of RMSC has a low-rankconstraint on the transition probability matrix, and simultaneouslya probabilistic simplex constraint on each of its rows. To solvethis challenging optimization problem, we propose an optimization procedurebased on the Augmented Lagrangian Multiplier scheme. Experimentalresults on various real world datasets show that theproposed method has superior performance over severalstate-of-the-art methods for multi-view clustering. Rongkai Xia, Yan Pan 0002, Jian Yin 0001 |
AAAI | 2 |
| 2014 | Supervised Hashing for Image Retrieval via Image Representation LearningabstractHashing is a popular approximate nearest neighbor search approach for large-scale image retrieval. Supervised hashing, which incorporates similarity/dissimilarity information on entity pairs to improve the quality of hashing function learning, has recently received increasing attention. However, in the existing supervised hashing methods for images, an input image is usually encoded by a vector of hand-crafted visual features. Such hand-crafted feature vectors do not necessarily preserve the accurate semantic similarities of images pairs, which may often degrade the performance of hashing function learning. In this paper, we propose a supervised hashing method for image retrieval, in which we automatically learn a good image representation tailored to hashing as well as a set of hash functions. The proposed method has two stages. In the first stage, given the pairwise similarity matrix $S$ over training images, we propose a scalable coordinate descent method to decompose $S$ into a product of $HH^T$ where $H$ is a matrix with each of its rows being the approximate hash code associated to a training image. In the second stage, we propose to simultaneously learn a good feature representation for the input images as well as a set of hash functions, via a deep convolutional network tailored to the learned hash codes in $H$ and optionally the discrete class labels of the images. Extensive empirical evaluations on three benchmark datasets with different kinds of images show that the proposed method has superior performance gains over several state-of-the-art supervised and unsupervised hashing methods. Rongkai Xia, Yan Pan 0002, Hanjiang Lai, Cong Liu 0001, Shuicheng Yan |
AAAI | 2 |
| 2014 | Efficient k-Support Matrix Pursuit
Hanjiang Lai, Yan Pan 0002, Canyi Lu, Yong Tang 0001, Shuicheng Yan |
ECCV (2) | 2 |
| 2014 | A Localized Efficient Forwarding Algorithm in Large-Scale Delay Tolerant NetworksabstractThis paper proposes an efficient opportunistic forwarding algorithm in Delay Tolerant Networks (DTNs) using only local information: inter-meeting times collected locally. It tries to minimize delay with a controlled energy consumption through placing a limitation on the number of total copies per message. The proposed forwarding algorithm makes forwarding decisions based only on local information, which means that no information is needed to be exchanged among the nodes, except for the data to be transferred. The removal of information propagation is particularly important in large-scale DTNs with limited communication opportunities like vehicular communication networks. On the contrary, most existing algorithms either forward messages randomly without facilitating any information, or require the exchange of certain information to make wise forwarding decision. Extensive real trace-driven simulations are conducted, and the proposed algorithm significantly outperforms all of the comvehicular communication networks. Onpared localized algorithms in every simulation. Yuxing He, Cong Liu 0001, Yan Pan 0002, Jun Zhang 0003, Jie Wu 0001, Yaxiong Zhao, Mingming Lu |
MASS | 3 |
| 2014 | Scalable opportunistic forwarding algorithms in delay tolerant networks using similarity hashingabstractDue to intermittent connectivity and uncertain node mobility, opportunistic message forwarding algorithms have been widely adopted in delay tolerant networks (DTNs). While existing work proposes practical forwarding algorithms in terms of increasing the delivery rate and decreasing data overhead, little attention has been drawn to the control overhead induced by the algorithms. The control overhead could, however, make the forwarding algorithms infeasible when the network size scales. In this paper, we are interested in increasing scalability by reducing control overhead, while retaining the state-of-the-art forwarding performances. The basic idea is to use locality-sensitive hashing to map each node as a hash-code, and use these hash-codes to compute the pair-wise similarity that guides the forwarding decisions. The proposed SOFA algorithms have reduced control overhead and competitive forwarding performance, which are verified by extensive real trace-driven simulations. Cong Liu 0001, Yan Pan 0002, Ai Chen, Kaigui Bian, Jie Wu 0001 |
SECON | 2 |
| 2014 | A DFA with Extended Character-Set for Fast Deep Packet InspectionabstractDeep packet inspection (DPI), based on regular expressions, is expressive, compact, and efficient in specifying attack signatures. We focus on their implementations based on general-purpose processors that are cost-effective and flexible to update. In this paper, we propose a novel solution, called deterministic finite automata with extended character-set (DFA/EC), which can significantly decrease the number of states through doubling the size of the character-set. Unlike existing state reduction algorithms, our solution requires only a single main memory access for each byte in the traffic payload, which is the minimum. We perform experiments with several Snort rule-sets. Results show that, compared to DFAs, DFA/ECs are very compact and are over four orders of magnitude smaller in the best cases; DFA/ECs also have smaller memory bandwidth and run faster. We believe that DFA/EC will lay a groundwork for a new type of state compression technique in fast packet inspection. Cong Liu 0001, Yan Pan 0002, Ai Chen, Jie Wu 0001 |
IEEE Trans. Computers | 2 |
| 2013 | Rank Aggregation via Low-Rank and Structured-Sparse DecompositionabstractRank aggregation, which combines multiple individual rank lists toobtain a better one, is a fundamental technique in various applications such as meta-search and recommendation systems. Most existing rank aggregation methods blindly combine multiple rank lists with possibly considerable noises, which often degrades their performances. In this paper, we propose a new model for robust rank aggregation (RRA) via matrix learning, which recovers a latent rank list from the possibly incomplete and noisy input rank lists. In our model, we construct a pairwise comparison matrix to encode the order information in each input rank list. Based on our observations, each comparison matrix can be naturally decomposed into a shared low-rank matrix, combined with a deviation error matrix which is the sum of a column-sparse matrix and a row-sparse one. The latent rank list can be easily extracted from the learned low-rank matrix. The optimization formulation of RRA has an element-wise multiplication operator to handle missing values, a symmetric constraint on the noise structure, and a factorization trick to restrict the maximum rank of the low-rank matrix. To solve this challenging optimization problem, we propose a novel procedure based on the Augmented Lagrangian Multiplier scheme. We conduct extensive experiments on meta-search and collaborative filtering benchmark datasets. The results show that the proposed RRA has superior performance gain over several state-of-the-art algorithms for rank aggregation. Yan Pan 0002, Hanjiang Lai, Cong Liu 0001, Yong Tang 0001, Shuicheng Yan |
AAAI | 1 |
| 2013 | A Divide-and-Conquer Method for Scalable Low-Rank Latent Matrix PursuitabstractData fusion, which effectively fuses multiple prediction lists from different kinds of features to obtain an accurate model, is a crucial component in various computer vision applications. Robust late fusion (RLF) is a recent proposed method that fuses multiple output score lists from different models via pursuing a shared low-rank latent matrix. Despite showing promising performance, the repeated full Singular Value Decomposition operations in RLF's optimization algorithm limits its scalability in real world vision datasets which usually have large number of test examples. To address this issue, we provide a scalable solution for large-scale low-rank latent matrix pursuit by a divide-and-conquer method. The proposed method divides the original low-rank latent matrix learning problem into two size-reduced sub problems, which may be solved via any base algorithm, and combines the results from the sub problems to obtain the final solution. Our theoretical analysis shows that with fixed probability, the proposed divide-and-conquer method has recovery guarantees comparable to those of its base algorithm. Moreover, we develop an efficient base algorithm for the corresponding sub problems by factorizing a large matrix into the product of two size-reduced matrices. We also provide high probability recovery guarantees of the base algorithm. The proposed method is evaluated on various fusion problems in object categorization and video event detection. Under comparable accuracy, the proposed method performs more than $180$ times faster than the state-of-the-art baselines on the CCV dataset with about 4,500 test examples for video event detection. Yan Pan 0002, Hanjiang Lai, Cong Liu 0001, Shuicheng Yan |
CVPR | 1 |
| 2013 | Efficient gradient descent algorithm for sparse models with application in learning-to-rank
Hanjiang Lai, Yan Pan 0002, Yong Tang 0001 |
Knowl. Based Syst. | 2 |
| 2013 | Sparse Learning-to-Rank via an Efficient Primal-Dual AlgorithmabstractLearning-to-rank for information retrieval has gained increasing interest in recent years. Inspired by the success of sparse models, we consider the problem of sparse learning-to-rank, where the learned ranking models are constrained to be with only a few nonzero coefficients. We begin by formulating the sparse learning-to-rank problem as a convex optimization problem with a sparse-inducing ℓ1constraint. Since the ℓ1constraint is nondifferentiable, the critical issue arising here is how to efficiently solve the optimization problem. To address this issue, we propose a learning algorithm from the primal dual perspective. Furthermore, we prove that, after at most O(1/ε) iterations, the proposed algorithm can guarantee the obtainment of an ε-accurate solution. This convergence rate is better than that of the popular subgradient descent algorithm. i.e., O(1/ε2). Empirical evaluation on several public benchmark data sets demonstrates the effectiveness of the proposed algorithm: 1) Compared to the methods that learn dense models, learning a ranking model with sparsity constraints significantly improves the ranking accuracies. 2) Compared to other methods for sparse learning-to-rank, the proposed algorithm tends to obtain sparser models and has superior performance gain on both ranking accuracies and training time. 3) Compared to several state-of-the-art algorithms, the ranking accuracies of the proposed algorithm are very competitive and stable. Hanjiang Lai, Yan Pan 0002, Cong Liu 0001, Liang Lin 0004, Jie Wu 0001 |
IEEE Trans. Computers | 2 |
| 2013 | FSMRank: Feature Selection Algorithm for Learning to RankabstractIn recent years, there has been growing interest in learning to rank. The introduction of feature selection into different learning problems has been proven effective. These facts motivate us to investigate the problem of feature selection for learning to rank. We propose a joint convex optimization formulation which minimizes ranking errors while simultaneously conducting feature selection. This optimization formulation provides a flexible framework in which we can easily incorporate various importance measures and similarity measures of the features. To solve this optimization problem, we use the Nesterov's approach to derive an accelerated gradient algorithm with a fast convergence rate O(1/T(2)). We further develop a generalization bound for the proposed optimization problem using the Rademacher complexities. Extensive experimental evaluations are conducted on the public LETOR benchmark datasets. The results demonstrate that the proposed method shows: 1) significant ranking performance gain compared to several feature selection baselines for ranking, and 2) very competitive performance compared to several state-of-the-art learning-to-rank algorithms. Hanjiang Lai, Yan Pan 0002, Yong Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | Integrating Graph Partitioning and Matching for Trajectory Analysis in Video SurveillanceabstractIn order to track moving objects in long range against occlusion, interruption, and background clutter, this paper proposes a unified approach for global trajectory analysis. Instead of the traditional frame-by-frame tracking, our method recovers target trajectories based on a short sequence of video frames, e.g., 15 frames. We initially calculate a foreground map at each frame obtained from a state-of-the-art background model. An attribute graph is then extracted from the foreground map, where the graph vertices are image primitives represented by the composite features. With this graph representation, we pose trajectory analysis as a joint task of spatial graph partitioning and temporal graph matching. The task can be formulated by maximizing a posteriori under the Bayesian framework, in which we integrate the spatio-temporal contexts and the appearance models. The probabilistic inference is achieved by a data-driven Markov chain Monte Carlo algorithm. Given a period of observed frames, the algorithm simulates an ergodic and aperiodic Markov chain, and it visits a sequence of solution states in the joint space of spatial graph partitioning and temporal graph matching. In the experiments, our method is tested on several challenging videos from the public datasets of visual surveillance, and it outperforms the state-of-the-art methods. Liang Lin 0004, Yongyi Lu, Yan Pan 0002, Xiaowu Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2011 | Greedy feature selection for rankingabstractThis paper is concerned with a study on the feature selection for ranking. Learning to rank is a useful tool for collaborative filtering and many other collaborative systems, which many algorithms have been proposed for dealing this issue. But feature selection methods receive little attention, despite of their importance in collaborative filtering problems: First, recommender systems always have massive data. Using all these data in learning to rank is unrealistic and impossible. Second, we discuss that not all the features are useful for a user's query. So choosing the most relevant data is necessary and useful. To amend this problem, we describe an algorithm called FBPCRank to choose the most relevant features for ranking. Our method combines two measures of good subsets of features, which not only can decrease the loss objective, but also reduce total similarity scores between any two features. We adopt forward and backward methods to choose the most relative features and use Pearson correlation coefficient to measure the similarity of two features. The experiments indicate that our method can outperform other state-of-the-art algorithms by selecting just small amounts of features. Hanjiang Lai, Yong Tang 0001, Hai-Xia Luo, Yan Pan 0002 |
CSCWD | 4 |
| 2011 | Color style transfer by constraint locally linear embeddingabstractThis paper presents a new semi-automatic method for color style transfer between images, which enhances the artistic expression of a image while preserving the content. Our method consists of three steps. (1) We first parse an input image into several semantic objects according to different material properties using interactive algorithms. (2) We search for a proper reference image from library using the semantic information. (3) The dominant colors of the input image and reference image are computed by clustering, and then we propose a constrainted locally linear embedding (CLLE) algorithm to perform color style transfer on the input image. In the experiments, we apply the proposed method to several photos to produce expressive results. Ruimao Zhang, Xiaodan Lan, Yan Pan 0002, Liang Lin 0004 |
ICIP | 4 |
| 2011 | Learning to rank with document ranks and scores
Yan Pan 0002, Hai-Xia Luo, Yong Tang 0001, Changqin Huang |
Knowl. Based Syst. | 1 |
| 2010 | Learning to rank with a Weight MatrixabstractLearning to rank, a task applying machine learning techniques to rank the expected information of the users, such as movie items users might be interested in. It is useful for collaborative filtering, which is regarded as a hot subfield of computer supported collaborative work(CSCW). In this paper, we propose an algorithm based on RankBoost to rank expected information of the users more accurately. The main advantage of the algorithm against RankBoost is to add a Weight Matrix regularizer to rank the relevance levels of the information smoothly and locally based on graph methods. The experimental results on the public LETOR datasets show that the proposed algorithm performs better than the baseline algorithm, indicating that the method is promising. Zewu Peng, Yong Tang 0001, Luxian Lin, Yan Pan 0002 |
CSCWD | 4 |
| 2008 | Question classification with semantic tree kernelabstractQuestion Classification plays an important role in most Question Answering systems. In this paper, we exploit semantic features in Support Vector Machines (SVMs) for Question Classification. We propose a semantic tree kernel to incorporate semantic similarity information. A diverse set of semantic features is evaluated. Experimental results show that SVMs with semantic features, especially semantic classes, can significantly outperform the state-of-the-art systems. Yan Pan 0002, Yong Tang 0001, Luxian Lin, Yemin Luo |
SIGIR | 1 |
| 2006 | Upper Bound Estimation of Turnaround Time Based on Fuzzy Temporal Workflow NetabstractWorkflow management has been a hot issue in both academic and industrial research. Deadline assignment is of great significance in workflow management. In order to avoid deadline violation, designers need an approach to evaluate the upper bound of turnaround time of a workflow process, especially in the early stage of the process design. In this paper, a method is proposed to evaluate the latest finishing time of a workflow instance. This method is based on our previously proposed model fuzzy temporal workflow nets (FTWF-Nets), which could depict uncertain or imprecise temporal information. Then an efficient algorithm extended from critical path method (CPM) is devised. Finally a case study is given Yong Tang 0001, Yan Pan 0002 |
CSCWD | 3 |
| 2006 | Time Performance Evaluation for Workflow Based on Extended FTWF-netsabstractPerformance analysis for workflow models has been recognized as one of the most significant tasks in workflow management. Time performance evaluation methods are needed for workflow designers to estimate the average turnaround time of processes in the early stage of workflow system development. Based on extending fuzzy temporal workflow nets by introducing choice possibilities for transitions in structural conflict, a new workflow model named extended fuzzy temporal workflow nets (EFTWF-nets) has been proposed. And the calculation of temporal elements in EFTWF-nets is given. Then a decomposition algorithm for EFTWF-nets is put forward and a time performance evaluation method based on it is investigated. Finally a case study is given to illustrate how to use this approach Yan Pan 0002, Yong Tang 0001, Ji'an Xu, Kaishun Wu |
CSCWD | 1 |