Yangxi Li

dblp:62/7950 · DBLP profile ↗
← Back
45ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PathMind: A Retrieve-Prioritize-Reason Framework for Knowledge Graph Reasoning with Large Language Models
abstract
Knowledge graph reasoning (KGR) is the task of inferring new knowledge by performing logical deductions on knowledge graphs. Recently, large language models (LLMs) have demonstrated remarkable performance in complex reasoning tasks. Despite promising success, current LLM-based KGR methods still face two critical limitations. First, existing methods often extract reasoning paths indiscriminately, without assessing their different importance, which may introduce irrelevant noise that misleads LLMs. Second, while some methods leverage LLMs to dynamically explore potential reasoning paths, they require high retrieval demands and frequent LLM calls. To address these limitations, we propose PathMind, a novel framework designed to enhance faithful and interpretable reasoning by selectively guiding LLMs with important reasoning paths. Specifically, PathMind follows a "Retrieve-Prioritize-Reason" paradigm. First, it retrieves a query subgraph from KG through the retrieval module. Next, it introduces a path prioritization mechanism that identifies important reasoning paths using a semantic-aware path priority function, which simultaneously considers the accumulative cost and the estimated future cost for reaching the target. Finally, PathMind generates accurate and logically consistent responses via a dual-phase training strategy, including task-specific instruction tuning and path-wise preference alignment. Extensive experiments on benchmark datasets demonstrate that PathMind consistently outperforms competitive baselines, particularly on complex reasoning tasks with fewer input tokens, by identifying essential reasoning paths.
Yu Liu 0118, Xixun Lin, Yanmin Shang, Yangxi Li, Shi Wang 0002, Yanan Cao 0001
AAAI4
2026 BCIRT: Backscattering-corrected implicit representation tomography
Chuanhao Zhang, Yangxi Li, Jianping Song, Yingwei Fan, Guochen Ning, Canhong Xiang, Fang Chen 0007, Hongen Liao
Medical Image Anal.2
2025 Dynamic Evaluation with Cognitive Reasoning for Multi-turn Safety of Large Language Models
abstract
The rapid advancement of Large Language Models (LLMs) poses significant challenges for safety evaluation. Current static datasets struggle to identify emerging vulnerabilities due to three limitations: (1) they risk being exposed in model training data, leading to evaluation bias; (2) their limited prompt diversity fails to capture real-world application scenarios; (3) they are limited to provide human-like multi-turn interactions. To address these limitations, we propose a dynamic evaluation framework, CogSafe, for comprehensive and automated multi-turn safety assessment of LLMs. We introduce CogSafe based on cognitive theories to simulate the real chatting process. To enhance assessment diversity, we introduce scenario simulation and strategy decision to guide the dynamic generation, enabling coverage of application situations. Furthermore, we incorporate the cognitive process to simulate multi-turn dialogues that reflect the cognitive dynamics of real-world interactions. Extensive experiments demonstrate the scalability and effectiveness of our framework, which has been applied to evaluate the safety of widely used LLMs.
Lanxue Zhang, Yanan Cao 0001, Yuqiang Xie, Fang Fang 0009, Yangxi Li
ACL (1)5
2025 Reliably Bounding False Positives: A Zero-Shot Machine-Generated Text Detection Framework via Multiscaled Conformal Prediction
abstract
The rapid advancement of large language models has raised significant concerns regarding their potential misuse by malicious actors.As a result, developing effective detectors to mitigate these risks has become a critical priority.However, most existing detection methods focus excessively on detection accuracy, often neglecting the societal risks posed by high false positive rates (FPRs).This paper addresses this issue by leveraging Conformal Prediction (CP), which effectively constrains the upper bound of FPRs.While directly applying CP constraints FPRs, it also leads to a significant reduction in detection performance.To overcome this trade-off, this paper proposes a Zero-Shot Machine-Generated Text Detection Framework via Multiscaled Conformal Prediction (MCP), which both enforces the FPR constraint and improves detection performance.This paper also introduces RealDet, a high-quality dataset that spans a wide range of domains, ensuring realistic calibration and enabling superior detection performance when combined with MCP.Empirical evaluations demonstrate that MCP effectively constrains FPRs, significantly enhances detection performance, and increases robustness against adversarial attacks across multiple detectors and datasets.
Yubing Ren, Yanan Cao 0001, Xixun Lin, Fang Fang 0009, Yangxi Li
ACL (1)6
2025 Deep Graph Neural Point Process for Learning Temporal Interactive Networks
Xiaohua Qi, Xixun Lin, Yanmin Shang, Yangxi Li
NLPCC (3)6
2025 FZeroTC: fully zero-shot text classification for simultaneously discovering and labeling unseen classes
Dongsheng Duan, Cunchi Lv, Yangxi Li
Knowl. Inf. Syst.6
2025 Language and Attenuation-Driven Network for Robot-Assisted Cholangiocarcinoma Diagnosis From Optical Coherence Tomography
abstract
Automatic and accurate classification of cholangiocarcinoma (CCA) using optical coherence tomography (OCT) images is critical for confirming infiltration margins. Considering that the morphological representations in pathology stains can be implicitly captured in OCT imaging, we introduce the optical attenuation coefficient (OAC) and generalized visual-language information to focus on the optical properties of diseased tissue and exploit its inherent textured features. Maintaining the data within the appropriate working range during OCT scanning is crucial for reliable diagnosis. To this end, we propose an autonomous scanning method integrated with novel deep learning architecture to construct an efficient computer-aided system. We develop a cross-modal complementarity model, the language and attenuation-driven network (LA-OCT Net), designed to enhance the interaction between OAC and OCT information and leverage generalized image-text alignment for refined feature representation. The model incorporates a disentangled attenuation selection-based adversarial correlation loss to magnify the discrepancy between cross-modal features while maintaining discriminative consistency. The proposed robot-assisted pipeline ensures precise repositioning of the diseased cross-sectional location, allowing consistent measurements to treatment and precise tumor margin detection. Extensive experiments on a comprehensive clinical dataset demonstrate the effectiveness and superiority of our method. Specifically, our approach not only improves accuracy by 6% compared to state-of-the-art techniques, while also providing new insights into the potential of optical biopsy.
Chuanhao Zhang, Yangxi Li, Jianping Song, Yuxuan Zhai, Yingwei Fan, Canhong Xiang, Fang Chen 0007, Hongen Liao
IEEE Trans. Medical Imaging2
2024 An anomaly aware network embedding framework for unsupervised anomalous link detection
Dongsheng Duan, Lingling Tong, Jie Lu 0009, Cunchi Lv, Yangxi Li
Data Min. Knowl. Discov.7
2024 One-Stage Anchor-Free Online Multiple Target Tracking With Deformable Local Attention and Task-Aware Prediction
abstract
The tracking-by-detection paradigm currently dominates multiple target tracking algorithms. It usually includes three tasks: target detection, appearance feature embedding, and data association. Carrying out these three tasks successively usually leads to lower tracking efficiency. In this paper, we propose a one-stage anchor-free multiple task learning framework which carries out target detection and appearance feature embedding in parallel to substantially increase the tracking speed. This framework simultaneously predicts a target detection and produces a feature embedding for each location, by sharing a pyramid of feature maps. We propose a deformable local attention module which utilizes the correlations between features at different locations within a target to obtain more discriminative features. We further propose a task-aware prediction module which utilizes deformable convolutions to select the most suitable locations for the different tasks. At the selected locations, classification of samples into foreground or background, appearance feature embedding, and target box regression are carried out. Two effective training strategies, regression range overlapping and sample reweighting, are proposed to reduce missed detections in dense scenes. Ambiguous samples whose identities are difficult to determine are effectively dealt with to obtain more accurate feature embedding of target appearance. An appearance-enhanced non-maximum suppression is proposed to reduce over-suppression of true targets in crowded scenes. Based on the one-stage anchor-free network with the deformable local attention module and the task-aware prediction module, we implement a new online multiple target tracker. Experimental results show that our tracker achieves a very fast speed while maintaining a high tracking accuracy.
Weiming Hu 0004, Shaoru Wang, Zongwei Zhou, Yangxi Li, Stephen J. Maybank
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 TSFN: an Effective Time Series Anomaly Detection Approach via Transformer-based Self-feedback Network
abstract
As the scale of data on the Internet continues to increase, the management and monitoring of time series data are facing significant challenges. Efficient and stable time-series data anomaly detection methods are necessary for fields such as traffic detection, power grid operation and maintenance, financial stock market, and industry. However, there are fewer abnormal data labels in time series data, and the labeling cost is high. Traditional expert knowledge-based supervised methods have been difficult to adapt to large-scale data metric management and timely abnormal alarms. At the same time, the way based on the new neural network has an extensive time overhead when faced with massive data, and it isn’t easy to apply it in a real-time industrial environment. Therefore, we propose the TSFN model in this paper, an unsupervised method of a transformer-based self-feedback network. Which can capture timing dependencies, learn normal data distribution and improve the self-feedback ability for sensitive areas, and can be used to detect anomalies in multidimensional time series more quickly. Our experimental research on five public datasets shows that our method has fast training speed, good stability, excellent anomaly detection ability, and good generalization ability compared with the baseline method.
Hongwei Wu, Rong Yang 0008, Huang Qing, Kedong Liu, Zhuojun Jiang, Yangxi Li
CSCWD6
2023 A Dual Domain Attention Mechanism for Face Forgery Detection
abstract
Recently, deep face forgery detection has been attracting considerable attentions, due to the potential security consequences induced by this type of forgeries. Unfortunately, the existing techniques have not specifically considered the intrinsic differences between the frequency and spatial domain information. To explicitly accommodate different feature representations from different domains, we propose a novel Dual Domain Attention Mechanism (DDAM) for deep face forgery detection. Inspired by digital image processing, we construct a “soft” filter to adaptively filter the frequency information, which is irrelevant to our forgery detection. Besides, we construct a FC-based Attention Module to maintain a receptive field of the entire feature map, to better leverage contextual information from different domains. Extensive experiments demonstrate the effectiveness of the proposed method on widely used datasets.
Yucong Suo, Xiaohan Zhao, Yuanfang Guo, Yangxi Li, Yunhong Wang 0001
IJCB4
2023 DRSDetector: Detecting Gambling Websites by Multi-level Feature Fusion
abstract
With the development of the Internet, online gambling has gradually replaced the traditional way of gambling and became a popular way of making money for illegal organizations. In many countries, online gambling is prohibited by law. But in some countries, these online gambling activities can still attract a variety of victims through their secret promotion channels. In this paper, we propose a gambling website detection method called DRSDetector, which combines domain features, resource features, and semantic features. And this method uses the idea of ensemble learning to fuse different modules. Specifically, we learn the character features of domains based on two stacked Transformer Encoder structures, use the LightGBM to learn resource feature of websites, and learn the semantic feature of websites based on the HAN model. The experimental results show that the performance of DRSDetector is better than the traditional website detection methods. In addition, we also investigated the promotion channels of gambling websites and took China as an example to reveal the 10 major entertainment companies behind these websites. These will help the government to combat online gambling activities more accurately and effectively.
Rong Yang 0008, Yangxi Li
ISCC4
2023 Ranking-Based Color Constancy With Limited Training Samples
abstract
Computational color constancy is an important component of Image Signal Processors (ISP) for white balancing in many imaging devices. Recently, deep convolutional neural networks (CNN) have been introduced for color constancy. They achieve prominent performance improvements comparing with those statistics or shallow learning-based methods. However, the need for a large number of training samples, a high computational cost and a huge model size make CNN-based methods unsuitable for deployment on low-resource ISPs for real-time applications. In order to overcome these limitations and to achieve comparable performance to CNN-based methods, an efficient method is defined for selecting the optimal simple statistics-based method (SM) for each image. To this end, we propose a novel ranking-based color constancy method (RCC) that formulates the selection of the optimal SM method as a label ranking problem. RCC designs a specific ranking loss function, and uses a low rank constraint to control the model complexity and a grouped sparse constraint for feature selection. Finally, we apply the RCC model to predict the order of the candidate SM methods for a test image, and then estimate its illumination using the predicted optimal SM method (or fusing the results estimated by the top k SM methods). Comprehensive experiment results show that the proposed RCC outperforms nearly all the shallow learning-based methods and achieves comparable performance to (sometimes even better performance than) deep CNN-based methods with only 1/2000 of the model size and training time. RCC also shows good robustness to limited training samples and good generalization crossing cameras. Furthermore, to remove the dependence on the ground truth illumination, we extend RCC to obtain a novel ranking-based method without ground truth illumination (RCC_NO) that learns the ranking model using simple partial binary preference annotations provided by untrained annotators rather than experts. RCC_NO also achieves better performance than the SM methods and most shallow learning-based methods with low costs of sample collection and illumination measurement.
Bing Li 0001, Haina Qin, Weihua Xiong, Yangxi Li, Songhe Feng, Weiming Hu 0004, Stephen J. Maybank
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video Classification
abstract
Compared with image few-shot learning, most of the existing few-shot video classification methods perform worse on feature matching, because they fail to sufficiently exploit the temporal information and relation. Specifically, frames are usually evenly sampled, which may miss important frames. On the other hand, the heuristic model simply encodes the equally treated frames in sequence, which results in the lack of both long-term and short-term temporal modeling and interaction. To alleviate these limitations, we take advantage of the compressed domain knowledge and propose a long-short term Cross-Transformer (LSTC) for few-shot video classification. For short terms, the motion vector (MV) contains temporal cues and reflects the importance of each frame. For long terms, a video can be natively divided into a sequence of GOPs (Group Of Picture). Using this compressed domain knowledge helps to obtain a more accurate spatial-temporal feature space. Consequently, we design the long-short term selection module, short-term module, and long-term module to comprise the LSTC. Long-short term selection is performed to select informative compressed domain data. Long/short-term modules are utilized to sufficiently exploit the temporal information so that the query and support can be well-matched by cross-attention. Experimental results show the superiority of our method on various datasets.
Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Yanan Miao, Yangxi Li
IJCAI6
2021 Web Objectionable Video Recognition Based on Deep Multi-Instance Learning With Representative Prototypes Selection
abstract
To protect underage people from accessing objectionable videos in the Internet, an effective objectionable video recognition algorithm is necessary for web filtering. Recently, the multi-instance learning has been introduced for objectionable video recognition and achieves impressive results. However, hand-crafted features as well as redundant and noisy frames in objectionable videos become an intractable problem that inevitably degrades the recognition performance. In this paper, we propose a novel representative prototype selection algorithm embedding deep multi-instance representation learning. In the proposed method, an improved convolutional neural network is designed for multimodal multi-instance feature learning and a self-expressive dictionary learning model based on sparse and low rank constraint is designed to select the representative prototypes from each subspace of instances. Then the bag-level feature is constructed via mapping the bag to the selected prototypes. Experiments on three objectionable video sets show the effectiveness of our method for objectionable video recognition.
Xinmiao Ding, Bing Li 0001, Yangxi Li, Weihua Xiong, Weiming Hu 0004
IEEE Trans. Circuits Syst. Video Technol.3
2021 Robust Texture-Aware Computer-Generated Image Forensic: Benchmark and Algorithm
abstract
With advances in rendering techniques and generative adversarial networks, computer-generated (CG) images tend to be indistinguishable from photographic (PG) images. Revisiting previous works towards CG image forensic, we observed that existing datasets are constructed years ago and limited in both quantity and diversity. Besides, current algorithms only consider the global visual features for forensic, ignoring finer differences between CG and PG images. To mitigate these problems, we first contribute a Large-Scale CG images Benchmark (LSCGB), and then propose a simple yet strong baseline model to address the forensic task. On the one hand, the introduced benchmark has three superior properties, 1) large-scale: the benchmark contains 71168 CG and 71168 PG images with the corresponding expert-annotated labels. It is orders of magnitude bigger than previous datasets. 2) high diversity: we collect CG images from 4 different scenes generated by various rendering techniques. The PG images are varied in terms of image content, camera models, and photographer styles. 3) small bias: we carefully filter the collected images to ensure that the distributions of color, brightness, tone and saturation between CG and PG images are close. Furthermore, inspired by an empirical study on texture difference between CG and PG images, an effective texture-aware network is proposed to improve forensic accuracy. Concretely, we first strengthen texture information of multilevel features extracted from a backbone. Then, the relations among feature channels are explored by learning its gram matrix. Each feature channel represents a specific texture pattern. The gram matrix is thus able to embed the finer texture differences. Experimental results demonstrate that this baseline surpasses the existing methods. The benchmark is publically available at https://github.com/wmbai/LSCGB.
Weiming Bai, Bing Li 0001, Yangxi Li, Congxuan Zhang, Weiming Hu 0004
IEEE Trans. Image Process.5
2021 EDP: An Efficient Decomposition and Pruning Scheme for Convolutional Neural Network Compression
abstract
Model compression methods have become popular in recent years, which aim to alleviate the heavy load of deep neural networks (DNNs) in real-world applications. However, most of the existing compression methods have two limitations: 1) they usually adopt a cumbersome process, including pretraining, training with a sparsity constraint, pruning/decomposition, and fine-tuning. Moreover, the last three stages are usually iterated multiple times. 2) The models are pretrained under explicit sparsity or low-rank assumptions, which are difficult to guarantee wide appropriateness. In this article, we propose an efficient decomposition and pruning (EDP) scheme via constructing a compressed-aware block that can automatically minimize the rank of the weight matrix and identify the redundant channels. Specifically, we embed the compressed-aware block by decomposing one network layer into two layers: a new weight matrix layer and a coefficient matrix layer. By imposing regularizers on the coefficient matrix, the new weight matrix learns to become a low-rank basis weight, and its corresponding channels become sparse. In this way, the proposed compressed-aware block simultaneously achieves low-rank decomposition and channel pruning by only one single data-driven training stage. Moreover, the network of architecture is further compressed and optimized by a novel Pruning & Merging (PM) module which prunes redundant channels and merges redundant decomposed layers. Experimental results (17 competitors) on different data sets and networks demonstrate that the proposed EDP achieves a high compression ratio with acceptable accuracy degradation and outperforms state-of-the-arts on compression rate, accuracy, inference time, and run-time memory.
Xiaofeng Ruan, Yufan Liu 0001, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Yangxi Li, Stephen J. Maybank
IEEE Trans. Neural Networks Learn. Syst.6
2021 RLINK: Deep reinforcement learning for user identity linkage
abstract
Abstract User identity linkage is a task of recognizing the identities of the same user across different social networks (SN). Previous works tackle this problem via estimating the pairwise similarity between identities from different SN, predicting the label of identity pairs or selecting the most relevant identity pair based on the similarity scores. However, most of these methods fail to utilize the results of previously matched identities, which could contribute to the subsequent linkages in following matching steps. To address this problem, we transform user identity linkage into a sequence decision problem and propose a reinforcement learning model to optimize the linkage strategy from the global perspective. Our method makes full use of both the social network structure and the history matched identities, meanwhile explores the long-term influence of processing matching on subsequent decisions. We conduct extensive experiments on real-world datasets, the results show that our method outperforms the state-of-the-art methods.
Yanan Cao 0001, Qian Li 0003, Yanmin Shang, Yangxi Li, Yanbing Liu 0007, Guandong Xu
World Wide Web5
2020 Learning from Easy to Complex: Adaptive Multi-Curricula Learning for Neural Dialogue Generation
abstract
Current state-of-the-art neural dialogue systems are mainly data-driven and are trained on human-generated responses. However, due to the subjectivity and open-ended nature of human conversations, the complexity of training dialogues varies greatly. The noise and uneven complexity of query-response pairs impede the learning efficiency and effects of the neural dialogue generation models. What is more, so far, there are no unified dialogue complexity measurements, and the dialogue complexity embodies multiple aspects of attributes—specificity, repetitiveness, relevance, etc. Inspired by human behaviors of learning to converse, where children learn from easy dialogues to complex ones and dynamically adjust their learning progress, in this paper, we first analyze five dialogue attributes to measure the dialogue complexity in multiple perspectives on three publicly available corpora. Then, we propose an adaptive multi-curricula learning framework to schedule a committee of the organized curricula. The framework is established upon the reinforcement learning paradigm, which automatically chooses different curricula at the evolving learning process according to the learning status of the neural dialogue generation model. Extensive experiments conducted on five state-of-the-art models demonstrate its learning efficiency and effectiveness with respect to 13 automatic evaluation metrics and human judgments.
Hengyi Cai, Hongshen Chen, Yonghao Song, Yangxi Li, Dongsheng Duan, Dawei Yin 0001
AAAI6
2020 Type-Aware Anchor Link Prediction across Heterogeneous Networks Based on Graph Attention Network
abstract
Anchor Link Prediction (ALP) across heterogeneous networks plays a pivotal role in inter-network applications. The difficulty of anchor link prediction in heterogeneous networks lies in how to consider the factors affecting nodes alignment comprehensively. In recent years, predicting anchor links based on network embedding has become the main trend. For heterogeneous networks, previous anchor link prediction methods first integrate various types of nodes associated with a user node to obtain a fusion embedding vector from global perspective, and then predict anchor links based on the similarity between fusion vectors corresponding with different user nodes. However, the fusion vector ignores effects of the local type information on user nodes alignment. To address the challenge, we propose a novel type-aware anchor link prediction across heterogeneous networks (TALP), which models the effect of type information and fusion information on user nodes alignment from local and global perspective simultaneously. TALP can solve the network embedding and type-aware alignment under a unified optimization framework based on a two-layer graph attention architecture. Through extensive experiments on real heterogeneous network datasets, we demonstrate that TALP significantly outperforms the state-of-the-art methods.
Yanmin Shang, Yanan Cao 0001, Yangxi Li, Jianlong Tan, Yanbing Liu 0007
AAAI4
2020 AANE: Anomaly Aware Network Embedding For Anomalous Link Detection
abstract
Existing network embedding models regard all the links in a network as normal and model them without distinction. In real networks, there may be anomalous links like noise or adversarial links. We explicitly consider the existence of anomalous links in a network and propose anomaly aware network embedding (AANE) model. The key of AANE is the design of a new loss, which consists of anomaly aware loss and adjusted fitting loss. We adopt an anomaly indicator to iteratively select significant anomalous links from the network during model training, and removal loss and deviation loss are designed to model the reconstruction errors of selected anomalous and normal links respectively. To instantiate AANE, AAGAE and AAGCN are implemented on graph auto-encoder (GAE) and graph convolution based auto-encoder (GCNAE) respectively. For the purpose of evaluation, a heuristic anomalous link generation algorithm is proposed and by using the algorithm we generate anomalous links into six real world network datasets. Experimental results show that AANE outperforms both basic and competitive network embedding models in terms of anomalous link detection performance in most cases.
Dongsheng Duan, Lingling Tong, Yangxi Li, Jie Lu 0009
ICDM3
2020 Semi-supervised Online Multi-Task Metric Learning for Visual Recognition and Retrieval
abstract
Distance metric learning (DML) is critial in many multimedia application tasks. However, it is hard to learn a satisfactory distance metric given only a few labeled samples for each task. In this paper, we proposed a novel semi-supervised online multi-task DML method termed SOMTML, which enables the models describing different tasks to help each other during the metric learning procedure and thus improving their respective performance. Besides, unlabeled data are leveraged to further help alleviate the data deficiency issue in different tasks by designing a novel regularization term, which also allows some prior information to be incorporated. More importantly, a quite efficient algorithm is developed to update the metrics of all tasks adaptively. The proposed SOMTML is experimentally validated in two popular visual analytic-based applications: handwriting digits recognition and face retrieval. We compared the proposed method with competitive single-task and multi-task metric learning approaches. Extensive experimental results demonstrate the effectiveness and efficiency of the proposed SOMTML.
Yangxi Li, Han Hu 0003, Jin Li 0014, Yong Luo 0002, Yonggang Wen 0001
ACM Multimedia1
2020 Anchor-Free One-Stage Online Multi-object Tracking
Zongwei Zhou, Yangxi Li, Junliang Xing, Liang Li 0003, Weiming Hu 0004
PRCV (2)2
2020 Anisotropic Convolution for Image Classification
abstract
Convolutional neural networks are built upon simple but useful convolution modules. The traditional convolution has a limitation on feature extraction and object localization due to its fixed scale and geometric structure. Besides, the loss of spatial information also restricts the networks' performance and depth. To overcome these limitations, this paper proposes a novel anisotropic convolution by adding a scale factor and a shape factor into the traditional convolution. The anisotropic convolution augments the receptive fields flexibly and dynamically depending on the valid sizes of objects. In addition, the anisotropic convolution is a generalized convolution. The traditional convolution, dilated convolution and deformable convolution can be viewed as its special cases. Furthermore, in order to improve the training efficiency and avoid falling into a local optimum, this paper introduces a simplified implementation of the anisotropic convolution. The anisotropic convolution can be applied to arbitrary convolutional networks and the enhanced networks are called ACNs (anisotropic convolutional networks). Experimental results show that ACNs achieve better performance than many state-of-the-art methods and the baseline networks in tasks of image classification and object localization, especially in classification task of tiny images.
Bing Li 0001, Chunfeng Yuan, Yangxi Li, Haohao Wu, Weiming Hu 0004, Fangshi Wang
IEEE Trans. Image Process.4
2019 Knowledge Distillation via Instance Relationship Graph
abstract
The key challenge of knowledge distillation is to extract general, moderate and sufficient knowledge from a teacher network to guide a student network. In this paper, a novel Instance Relationship Graph (IRG) is proposed for knowledge distillation. It models three kinds of knowledge, including instance features, instance relationships and feature space transformation, while the latter two kinds of knowledge are neglected by previous methods. Firstly, the IRG is constructed to model the distilled knowledge of one network layer, by considering instance features and instance relationships as vertexes and edges respectively. Secondly, an IRG transformation is proposed to models the feature space transformation across layers. It is more moderate than directly mimicking the features at intermediate layers. Finally, hint loss functions are designed to force a student's IRGs to mimic the structures of a teacher's IRGs. The proposed method effectively captures the knowledge along the whole network via IRGs, and thus shows stable convergence and strong robustness to different network architectures. In addition, the proposed method shows superior performance over existing methods on datasets of various scales.
Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Chunfeng Yuan, Weiming Hu 0004, Yangxi Li, Yunqiang Duan
CVPR6
2019 PAAE: A Unified Framework for Predicting Anchor Links with Adversarial Embedding
abstract
The goal of predicting anchor links is to align accounts from multiple networks by whether they are held by the same natural person. Network structure is the key information for predicting anchor links. Exploring the intrinsic attributes of the network structure is an important way to align anchor users across social networks. Existing methods use a representation learning approach to embed network vertices into low dimension vectors space. But these methods suffer from lack of additional constraints for enhancing the robustness of the embedding vectors when aligning anchor nodes across networks with large structural differences. To offer a robust method, we propose a novel adversarial representation learning approach to align users, called PAAE(predicting anchor links with adversarial embedding), which employs an adversarial regularization to capture the robust embedding vectors and maps anchor users with an alignment autoencoders. PAAE can solve both the network embedding problem and the user alignment problem simultaneously under a unified optimization framework. Through extensive experiments on real social network datasets, we demonstrate that PAAE significantly outperforms the state-of-the-art methods.
Yanmin Shang, Zhezhou Kang, Yanan Cao 0001, Yangxi Li, Yanbing Liu 0007
ICME6
2018 Interaction-Aware Spatio-Temporal Pyramid Attention Networks for Action Classification
Chunfeng Yuan, Bing Li 0001, Yangxi Li, Weiming Hu 0004
ECCV (16)5
2018 Context-Dependent Random Walk Graph Kernels and Tree Pattern Graph Matching Kernels With Applications to Action Recognition
abstract
Graphs are effective tools for modeling complex data. Setting out from two basic substructures, random walks and trees, we propose a new family of context-dependent random walk graph kernels and a new family of tree pattern graph matching kernels. In our context-dependent graph kernels, context information is incorporated into primary random walk groups. A multiple kernel learning algorithm with a proposed l1,2-norm regularization is applied to combine context-dependent graph kernels of different orders. This improves the similarity measurement between graphs. In our tree-pattern graph matching kernel, a quadratic optimization with a sparse constraint is proposed to select the correctly matched tree-pattern groups. This augments the discriminative power of the tree-pattern graph matching. We apply the proposed kernels to human action recognition, where each action is represented by two graphs which record the spatiotemporal relations between local feature vectors. Experimental comparisons with state-of-the-art algorithms on several benchmark datasets demonstrate the effectiveness of the proposed kernels for recognizing human actions. It is shown that our kernel based on tree-pattern groups, which have more complex structures and exploit more local topologies of graphs than random walks, yields more accurate results but requires more runtime than the context-dependent walk graph kernel.
Weiming Hu 0004, Baoxin Wu, Chunfeng Yuan, Yangxi Li, Stephen J. Maybank
IEEE Trans. Image Process.5
2016 Knowledge Graph Completion via Local Semantic Contexts
Xiangling Zhang, Cuilan Du, Peishan Li, Yangxi Li
DASFAA (1)4
2016 Random Partition Factorization Machines for Context-Aware Recommendations
Cuilan Du, Kankan Zhao, Cuiping Li 0001, Yangxi Li, Hong Chen 0001
WAIM (1)5
2016 Manifold regularized multi-view feature selection for social image annotation
Yangxi Li, Cuilan Du, Yang Liu 0003, Yonggang Wen 0001
Neurocomputing1
2015 Large-margin multi-view Gaussian process
Chang Xu 0002, Dacheng Tao, Yangxi Li, Chao Xu 0006
Multim. Syst.3
2013 FIM: A Real-Time Content Based Sample Image Matching System
abstract
Sample Image Matching is to decide if a queried image is belongs to the database or not. In this paper, we focus on real-time image matching, which is critical in many real world applications. Although traditional image retrieval methods can be directly utilized for image matching, they usually suffer the high computational cost problem and thus is not applicable here. To resolve this problem, we first introduce ORB, a recently proposed and well-established image feature, for image matching. Then we compare several variants of the descriptors, different size of the codebook, and two approaches to compute the matching scores, based on which we propose a strategy for final matching decision. According to the comparison results, we finally present a real-time image matching system, fast image matching (FIM), which can process about 33 images per second, with a satisfactory accuracy.
Yong Luo 0002, Yangxi Li, Jinhui Tu, Chao Xu 0006
ICIG2
2013 A comprehensive study on learning to rank for content-based image retrieval
Yangxi Li, Bo Geng, Chao Xu 0006, Hong Liu 0008
Signal Process.1
2012 Learning to rerank images with enhanced spatial verification
abstract
Reranking is one of the commonly used schemes to improve the initial ranking performance for content based image retrieval (CBIR). The state-of-the-art reranking methods for CBIR are mainly based on spatial verification and global feature. To mine the complementary properties of different reranking strategies, we combine features representing images from different perspectives with RankSVM to obtain a reranking model to refine the initial ranking list. Besides, compared with traditional spatial verification based methods which measure image similarity only with single inlier's statistical properties, we bind close inlier visual words together to mine more geometric information from images. Through organizing inliers into sequence and computing the relative positions among inliers, we define an efficient similarity measurement with the order consistency between inlier sequences. Experimental results on both Oxford and imageNet datasets demonstrate that our proposed reranking method is effective and promising.
Chang Xu 0002, Yangxi Li, Chao Xu 0006
ICIP2
2012 Query difficulty estimation for image retrieval
Yangxi Li, Bo Geng, Linjun Yang, Chao Xu 0006
Neurocomputing1
2012 Parallel Lasso for Large-Scale Video Concept Detection
abstract
Existing video concept detectors are generally built upon the kernel based machine learning techniques, e.g., support vector machines, regularized least squares, and logistic regression, just to name a few. However, in order to build robust detectors, the learning process suffers from the scalability issues including the high-dimensional multi-modality visual features and the large-scale keyframe examples. In this paper, we propose parallel lasso (Plasso) by introducing the parallel distributed computation to significantly improve the scalability of lasso (thel1regularized least squares). We apply the parallel incomplete Cholesky factorization to approximate the covariance statistics in the preprocess step, and the parallel primal-dual interior-point method with the Sherman-Morrison-Woodbury formula to optimize the model parameters. For a dataset withnsamples in ad-dimensional space, compared with lasso, Plasso significantly reduces complexities from the originalO(d3) for computational time andO(d2) for storage space toO(h2d/m) andO(hd/m) , respectively, if the system hasmprocessors and the reduced dimensionhis much smaller than the original dimensiond. Furthermore, we develop the kernel extension of the proposed linear algorithm with the sample reweighting schema, and we can achieve similar time and space complexity improvements [time complexity fromO(n3) toO(h2n/m) and the space complexity fromO(n2) toO(hn/m), for a dataset withntraining examples]. Experimental results on TRECVID video concept detection challenges suggest that the proposed method can obtain significant time and space savings for training effective detectors with limited communication overhead.
Bo Geng, Yangxi Li, Dacheng Tao, Meng Wang 0001, Zhengjun Zha, Chao Xu 0006
IEEE Trans. Multim.2
2012 Difficulty Guided Image Retrieval Using Linear Multiple Feature Embedding
abstract
Existing image retrieval systems suffer from a performance variance for different queries. Severe performance variance may greatly degrade the effectiveness of the subsequent query-dependent ranking optimization algorithms, especially those that utilize the information mined from the initial search results. In this paper, we tackle this problem by proposing a query difficulty guided image retrieval system, which can predict the queries' ranking performance in terms of their difficulties and adaptively apply ranking optimization approaches. We estimate the query difficulty by comprehensively exploring the information residing in the query image, the retrieval results, and the target database. To handle the high-dimensional and multi-model image features in the large-scale image retrieval setting, we propose a linear multiple feature embedding algorithm which learns a linear transformation from a small set of data by integrating a joint subspace in which the neighborhood information is preserved. The transformation can be effectively and efficiently used to infer the subspace features of the newly observed data in the online setting. We prove the significance of query difficulty to image retrieval by applying it to guide the conduction of three retrieval refinement applications, i.e., reranking, federated search, and query suggestion. Thorough empirical studies on three datasets suggest the effectiveness and scalability of the proposed image query difficulty estimation algorithm, as well as the promising of the image difficulty guided retrieval system.
Yangxi Li, Bo Geng, Dacheng Tao, Zhengjun Zha, Linjun Yang, Chao Xu 0006
IEEE Trans. Multim.1
2011 Query expansion by spatial co-occurrence for image retrieval
abstract
The well-known bag-of-features (BoF) model is widely utilized for large scale image retrieval. However, BoF model lacks the spatial information of visual words, which is informative for local features to build up meaningful visual patches. To compensate for the spatial information loss, in this paper, we propose a novel query expansion method called Spatial Co-occurrence Query Expansion (SCQE), by utilizing the spatial co-occurrence information of visual words mined from the database images to boost the retrieval performance. In offline phase, for each visual word in the vocabulary, we treat the visual words that are frequently co-occurred with it in the database images as neighbors, base on which a spatial co-occurrence graph is built. In online phase, a query image can be expanded with some spatial co-occurred but unseen visual words according to the spatial co-occurrence graph, and the retrieval performance can be improved by expanding these visual words appropriately. Experimental results demonstrate that, SCQE achieves promising improvements over the typical BoF baseline on two datasets comprising 5K and 505K images respectively.
Yingfei Li, Bo Geng, Zhengjun Zha, Yangxi Li, Dacheng Tao, Chao Xu 0006
ACM Multimedia4
2011 Difficulty guided image retrieval using linear multiview embedding
abstract
Existing image retrieval systems suffer from a radical performance variance for different queries. The bad initial search results for "difficult" queries may greatly degrade the performance of their subsequent refinements, especially the refinement that utilizes the information mined from the search results, e.g., pseudo relevance feedback based reranking. In this paper, we tackle this problem by proposing a query difficulty guided image retrieval system, which selectively performs reranking according to the estimated query difficulty. To improve the performance of both reranking and difficulty estimation, we apply multiview embedding (ME) to images represented by multiple different features for integrating a joint subspace by preserving the neighborhood information in each feature space. However, existing ME approaches suffer from both "out of sample" and huge computational cost problems, and cannot be applied to online reranking or offline large-scale data processing for practical image retrieval systems. Therefore, we propose a linear multiview embedding algorithm which learns a linear transformation from a small set of data and can effectively infer the subspace features of new data. Empirical evaluations on both Oxford and 500K ImageNet datasets suggest the effectiveness of the proposed difficulty guided retrieval system with LME.
Yangxi Li, Bo Geng, Zhengjun Zha, Dacheng Tao, Linjun Yang, Chao Xu 0006
ACM Multimedia1
2011 Query Difficulty Guided Image Retrieval System
Yangxi Li, Yong Luo 0002, Dacheng Tao, Chao Xu 0006
MMM (2)1
2009 Color Correction and Compression for Multi-view Video Using H.264 Features
Boxin Shi, Yangxi Li, Chao Xu 0006
ACCV (3)2
2009 Integrating Color Constancy into Multi-view Video Coding
abstract
Color constancy is the ability to remove the dependency of illuminant and show the intrinsic color of objects. A novel multi-view video coding scheme integrated with color constancy algorithms is introduced in this paper. Color constancy algorithms based on gray-edge hypothesis are selected. The new scheme takes color constancy as a preprocessing to make the sequence independent of the illuminant conditions. Based on the examination of the scene change in the sequence, the key frames are picked out and the parameters of color constancy are determined. The illuminant information is sent into the encoder along with the illuminant independent sequence, and then it is used to solve the color cast problem of multi-view coding at the decoder. Both visual quality and coding performance prove that this novel scheme can lower the bitrates and gain better color quality.
Yangxi Li, Boxin Shi, Chao Xu 0006
ICIG1
2009 Intrinsic Image Decomposition Using Color Invariant Edge
abstract
The intrinsic image composed of reflectance and shading images plays important roles in various computer vision applications. This paper focuses on solving the problem of intrinsic image decomposition. Based on the assumption that the image derivatives can be classified into either reflectance-related or shading-related, the reflectance and shading image can be restored from the classified derivatives. We improve the classification result using only color information by introducing the color invariant edge. Considering some color invariant properties in the image, the color invariant edge can provide more useful information in guiding the classification and producing more robust decomposition result as it is shown in the experiment.
Boxin Shi, Yangxi Li, Chao Xu 0006
ICIG2
2009 Block-based color correction algorithm for multi-view video coding
abstract
The color variations among different viewpoints in multiview video sequences may deteriorate the visual quality and coding efficiency. Various color correction methods have been proposed, however, the color appearance and histogram of corrected target frames are not similar enough to the reference frames in details. Focusing on restoring more similar color, a block-based color correction algorithm is proposed. The blocks in reference frames are matched into target frames through spatial prediction, and the colorization scheme is then adopted to expand color as a coarse correction. Finally the mixture with global color transfer result yields the fine correction. The experiment results show this novel method can provide better visual effect in detail and also provide the corrected frames with histograms more similar to reference histograms.
Boxin Shi, Yangxi Li, Chao Xu 0006
ICME2