Lijiang Chen

dblp:76/3584 · DBLP profile ↗
← Back
49ranked-venue papers
9as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 1
YearPublicationVenuePosition
2026 UKANCNet: Multi-scale feature fusion with UKAN enhancement for micro wind turbine blade defect segmentation
Jizheng Yi, Xiangyu Shen, Lijiang Chen, Ze Jin
Expert Syst. Appl.5
2026 ViTaL: A multimodality dataset and benchmark for multi-pathological ovarian tumor recognition
abstract
Ovarian tumor, as a common gynecological disease, can rapidly deteriorate into serious health crises when undetected early, thus posing significant threats to the health of women. Deep neural networks have the potential to identify ovarian tumors, thereby reducing mortality rates, but limited public datasets hinder its progress. To address this gap, we introduce a vital ovarian tumor pathological recognition dataset called ViTaL that contains V isual, T abular and L inguistic modality data of 496 patients across six pathological categories. The ViTaL dataset comprises three subsets corresponding to different patient data modalities: visual data from 2216 two-dimensional ultrasound images, tabular data from medical examinations of 496 patients, and linguistic data from ultrasound reports of 496 patients. It is insufficient to merely distinguish between benign and malignant ovarian tumors in clinical practice. To enable multi-pathology classification of ovarian tumor, we propose a ViTaL-Net based on the Triplet Hierarchical Offset Attention Mechanism (THOAM) to minimize the loss incurred during feature fusion of multi-modal data. This mechanism could effectively enhance the relevance and complementarity between information from different modalities. ViTaL-Net serves as a benchmark for the task of multi-pathology, multi-modality classification of ovarian tumors. In our comprehensive experiments, the proposed method exhibits satisfactory performance, achieving accuracies exceeding 90 % on the two most common pathological types of ovarian tumors and an overall performance of 85 %. Our dataset and code are available at https://github.com/GGbond-study/vitalnet .
Lijiang Chen, Guangxia Cui, Wenpei Bai, Shuchang Lyu, Qi Zhao 0037
Expert Syst. Appl.2
2026 Unsupervised cross-domain semantic segmentation on multi-modality ovarian tumor ultrasound data
Shuchang Lyu, Qi Zhao 0037, Wenpei Bai, Linghan Cai, Guangxia Cui, Lijiang Chen, Huiyu Zhou 0001
Pattern Recognit.8
2025 Look in different views: Multi-scheme regression guided cell instance segmentation
Yunmeng Huang, Wenquan Feng, Shuchang Lyu, Qi Zhao 0037, Lijiang Chen
Knowl. Based Syst.7
2025 A Voronoi Density-Based Locally Unique Network for Fine-Grained Multi-Label Classification
abstract
Multi-label image classification aims to classify all categories in images simultaneously. When current multi-label classification methods meet fine-grained objects in a single image, the extreme inter-class similarity and over-prediction problems are two major challenges that hinder model performance. To solve the above two problems, we propose Voronoi density based Locally Unique Network (VoLUNet). First, due to high correlation between predictions of different classes, following the Kolmogorov-Arnold Network (KAN), we design the Weak Inter-class Correlation Classifier (WIC-Classifier) to replace linear weights setting in MLP architecture, promoting the potential of fine-grained discrimination. Second, we propose a Local Non-Maximum Suppression (Local-NMS) loss to multi-label classification model, predicting only one unique class with high prediction value for each local region. Third, different classes may have different pixel proportions and Local-NMS loss will be imbalanced for diverse fine-grained classes, we design the Voronoi Density based Superpixel Module (VDSM) to balance the quantities of local feature vectors with different classes. Finally, comprehensive experiments are conducted on four datasets, TreeSatAI, GeoLifeCLEF, FothemNet ShipRSImageNet, and our VoLUNet can significantly improve the classification performance compared to current state-of-the-art models. Codes of this paper are public available at https://github.com/cv516Buaa/BinghaoLiu/tree/main/VoLUNet.
Binghao Liu, Qi Zhao 0037, Lijiang Chen
IEEE Trans. Circuits Syst. Video Technol.5
2025 Graph Reconstruction Attention Fusion Network for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) has become increasingly popular due to the exponential surge of user comments on social media. The MSA aims to efficiently integrate various modalities through a superior fusion framework. However, previous studies have primarily focused on the integration of sequence data while neglecting its structural information. In addition, effectively modeling the continuous expression of human sentiment polarity remains a significant challenge. Therefore, we propose the graph reconstruction attention fusion network, which availably promotes the multimodal fusion process by combining sequence learning with graph learning. First, we design a graph reconstruction learning module to obtain multimodal graph embeddings. Second, a text-guided cross-modal enhancement architecture is adopted to acquire multimodal representations, where a sentiment attenuation factor is introduced to promote emotional continuity modeling. Finally, we propose a feature-wised attention structure adapted for the classifier, it dynamically adjusts weights of multimodal features that are beneficial for downstream tasks. Extensive experiments on three challenging datasets, CMU-MOSI, CMU-MOSEI, and CH-SIMS demonstrate that our model significantly outperforms existing state-of-the-art methods.
Ronglong Hu, Jizheng Yi, Lijiang Chen, Ze Jin
IEEE Trans. Ind. Informatics3
2024 OV-VG: A benchmark for open-vocabulary visual grounding
abstract
Open-vocabulary learning has emerged as a cutting-edge research area, particularly in light of the widespread adoption of vision-based foundational models. Its primary objective is to comprehend novel concepts that are not encompassed within a predefined vocabulary. One key facet of this endeavor is Visual Grounding (VG), which entails locating a specific region within an image based on a corresponding language description. While current foundational models excel at various visual language tasks, there is a noticeable absence of models specifically tailored for open-vocabulary visual grounding (OV-VG). This research endeavor introduces novel and challenging OV tasks, namely Open-Vocabulary Visual Grounding (OV-VG) and Open-Vocabulary Phrase Localization (OV-PL). The overarching aim is to establish connections between language descriptions and the localization of novel objects. To facilitate this, we have curated a comprehensive annotated benchmark, encompassing 7272 OV-VG images (comprising 10,000 instances) and 1000 OV-PL images. In our pursuit of addressing these challenges, we delved into various baseline methodologies rooted in existing open-vocabulary object detection (OV-D), VG, and phrase localization (PL) frameworks. Surprisingly, we discovered that state-of-the-art (SOTA) methods often falter in diverse scenarios. Consequently, we developed a novel framework that integrates two critical components: Text-Image Query Selection (TIQS) and Language-Guided Feature Attention (LGFA). These modules are designed to bolster the recognition of novel categories and enhance the alignment between visual and linguistic information. Extensive experiments demonstrate the efficacy of our proposed framework, which consistently attains SOTA performance across the OV-VG task. Additionally, ablation studies provide further evidence of the effectiveness of our innovative models. Codes and datasets will be made publicly available at https://github.com/cv516Buaa/OV-VG.
Wenquan Feng, Xiangtai Li, Shuchang Lyu, Binghao Liu, Lijiang Chen, Qi Zhao 0037
Neurocomputing7
2024 Know your orientation: A viewpoint-aware framework for polyp segmentation
abstract
Automatic polyp segmentation in endoscopic images is critical for the early diagnosis of colorectal cancer. Despite the availability of powerful segmentation models, two challenges still impede the accuracy of polyp segmentation algorithms. Firstly, during a colonoscopy, physicians frequently adjust the orientation of the colonoscope tip to capture underlying lesions, resulting in viewpoint changes in the colonoscopy images. These variations increase the diversity of polyp visual appearance, posing a challenge for learning robust polyp features. Secondly, polyps often exhibit properties similar to the surrounding tissues, leading to indistinct polyp boundaries. To address these problems, we propose a viewpoint-aware framework named VANet for precise polyp segmentation. In VANet, polyps are emphasized as a discriminative feature and thus can be localized by class activation maps in a viewpoint classification process. With these polyp locations, we design a viewpoint-aware Transformer (VAFormer) to alleviate the erosion of attention by the surrounding tissues, thereby inducing better polyp representations. Additionally, to enhance the polyp boundary perception of the network, we develop a boundary-aware Transformer (BAFormer) to encourage self-attention towards uncertain regions. As a consequence, the combination of the two modules is capable of calibrating predictions and significantly improving polyp segmentation performance. Extensive experiments on seven public datasets across six metrics demonstrate the state-of-the-art results of our method, and VANet can handle colonoscopy images in real-world scenarios effectively. The source code is available at https://github.com/1024803482/Viewpoint-Aware-Network.
Linghan Cai, Lijiang Chen, Yifeng Wang 0001, Yongbing Zhang 0002
Medical Image Anal.2
2024 Multichannel Cross-Modal Fusion Network for Multimodal Sentiment Analysis Considering Language Information Enhancement
abstract
With the popularity of short videos, analyzing human emotions is crucial for understanding individual attitudes and guiding social public opinions. Consequently, multimodal sentiment analysis (MSA) has garnered significant attention in the field of human–computer interaction. The main challenge of MSA is to explore a high-quality multimodal fusion framework, as multiple modalities contribute inconsistently to sentiment prediction. However, most of the existing methods assume equal importance among different modalities, resulting in inadequate expression of the main modality. In addition, auxiliary modalities often contain redundant information, which hinders the multimodal fusion process. Therefore, we propose the multichannel cross-modal fusion network (MCFNet) to promote the multimodal fusion procedure by constructing a multichannel various modality fusion framework comprising three channels: obtaining multimodal representation through the first channel; eliminating information redundancy from auxiliary modalities via the second channel; and enhancing significance attributed to the main modality adopting the third channel. Subsequently, we design a multichannel information fusion gate to integrate feature representations from these three channels for downstream sentiment classification tasks. Numerous experiments on three benchmark datasets, CMU-multimodal opinion sentiment intensity (MOSI), CMU-multimodal opinion sentiment and emotion intensity (MOSEI), and Twitter2019, show that the MCFNet has made a significant progress compared to recent state-of-the-art methods.
Ronglong Hu, Jizheng Yi, Aibin Chen, Lijiang Chen
IEEE Trans. Ind. Informatics4
2023 Learn by Oneself: Exploiting Weight-Sharing Potential in Knowledge Distillation Guided Ensemble Network
abstract
Recent CNNs (convolutional neural networks) have become more and more compact. The elegant structure design highly improves the performance of CNNs. With the development of knowledge distillation technique, the performance of CNNs gets further improved. However, existing knowledge distillation guided methods either rely on offline pretrained high-quality large teacher models or online heavy training burden. To solve the above problems, we propose a feature-sharing and weight-sharing based ensemble network (training framework) guided by knowledge distillation (EKD-FWSNet) to make baseline models stronger in terms of representation ability with less training computation and memory cost involved. Specifically, motivated by getting rid of the dependence of offline pretrained teacher model, we design an end-to-end online training scheme to optimize EKD-FWSNet. Motivated by decreasing the online training burden, we only introduce one auxiliary classmate branch to construct multiple forward branches, which will then be integrated as ensemble teacher to guide baseline model. Compared to previous online ensemble training frameworks, EKD-FWSNet can provide diverse output predictions without relying on increasing auxiliary classmate branches. Motivated by maximizing the optimization power of EKD-FWSNet, we exploit the representation potential of weight-sharing blocks and design efficient knowledge distillation mechanism in EKD-FWSNet. Extensive comparison experiments and visualization analysis on benchmark datasets (CIFAR-10/100, tiny-ImageNet, CUB-200 and ImageNet) show that self-learned EKD-FWSNet can boost the performance of baseline models by large margin, which has obvious superiority compared to previous related methods. Extensive analysis also proves the interpretability of EKD-FWSNet. Our code is available at https://github.com/cv516Buaa/EKD-FWSNet.
Qi Zhao 0037, Shuchang Lyu, Lijiang Chen, Binghao Liu, Ting-Bing Xu, Wenquan Feng
IEEE Trans. Circuits Syst. Video Technol.3
2023 MGML: Multigranularity Multilevel Feature Ensemble Network for Remote Sensing Scene Classification
Qi Zhao 0037, Shuchang Lyu, Yujing Ma, Lijiang Chen
IEEE Trans. Neural Networks Learn. Syst.5
2022 A Similarity Distillation Guided Feature Refinement Network for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation is a challenging task of predicting object categories in pixel-wise with only few annotated samples. Existing methods mainly have two problems, which are representation inconsistency of query and support images and semantic-level feature insufficient. To tackle the two problems, we propose a similarity distillation guided feature refinement network (SD-FRNet). Specifically, we first use support label to generate support similarity feature map and coarse prediction of query image. Then, we use this coarse prediction to generate query similarity feature. To compensate feature representation inconsistency, we conduct knowledge distillation mechanism to align similarity features of query and support images. To enrich semantic-level feature, we further design a feature refinement module, which achieves high-quality segmentation. Extensive experiments show the effectiveness of SD-FRNet. On benchmark datasets, PASCAL-5iand COCO-20i, our proposed SD-FRNet outperform the previous SOTA (state-of-the-art) results.
Shuchang Lyu, Binghao Liu, Lijiang Chen, Qi Zhao 0037
ICIP3
2022 Using Guided Self-Attention with Local Information for Polyp Segmentation
Linghan Cai, Meijing Wu, Lijiang Chen, Wenpei Bai, Shuchang Lyu, Qi Zhao 0037
MICCAI (4)3
2022 ML-FDA: Meta-Learning via Feature Distribution Alignment for Few-Shot Learning
abstract
Computer vision tasks suffer from the high cost of collecting large amounts of labeled data. Few-shot Learning (FSL) is a dominant approach to solve this problem because it provides an insight to learn the knowledge of novel categories with few training samples. In FSL task, Meta-learning and metric learning have achieved impressive results. However, the performance of this task is still limited by large intra-class variance and small inter-class distance caused by limited number of few samples. To solve this problem, In this paper, we propose a new method, which integrates meta-learning and metric learning techniques. Specifically, we first propose a feature representation module (FR) to construct representative support class prototypes and query features. Then, we design bias loss to minimize the bias between support and query samples. Furthermore, we design an intra-class loss to minimize the distance between query class prototype and each query sample. We denote this model as ML-FDA and validate it on standard few-shot classification benchmark datasets (MiniImageNet, CIFAR-FS, FC100). The results show that our method improves the performance over other same paradigm methods and achieves the best performance on most benchmarks. The ablation study and visulization analysis also demonstrate the effectiveness of our method.
Binghao Liu, Shuchang Lyu, Lijiang Chen, Qi Zhao 0037, Wenquan Feng
VCIP4
2022 Limited text speech synthesis with electroglottograph based on Bi-LSTM and modified Tacotron-2
abstract
Abstract This paper proposes a framework of applying only the EGG signal for speech synthesis in the limited categories of contents scenario. EGG is a sort of physiological signal which can reflect the trends of the vocal cord movement. Note that EGG’s different acquisition method contrasted with speech signals, we exploit its application in speech synthesis under the following two scenarios. (1) To synthesize speeches under high noise circumstances, where clean speech signals are unavailable. (2) To enable dumb people who retain vocal cord vibration to speak again. Our study consists of two stages, EGG to text and text to speech. The first is a text content recognition model based on Bi-LSTM, which converts each EGG signal sample into the corresponding text with a limited class of contents. This model achieves 91.12% accuracy on the validation set in a 20-class content recognition experiment. Then the second step synthesizes speeches with the corresponding text and the EGG signal. Based on modified Tacotron-2, our model gains the Mel cepstral distortion (MCD) of 5.877 and the mean opinion score (MOS) of 3.87, which is comparable with the state-of-the-art performance and achieves an improvement by 0.42 and a relatively smaller model size than the origin Tacotron-2. Considering to introduce the characteristics of speakers contained in EGG to the final synthesized speech, we put forward a fine-grained fundamental frequency modification method, which adjusts the fundamental frequency according to EGG signals and achieves a lower MCD of 5.781 and a higher MOS of 3.94 than that without modification.
Lijiang Chen, Xia Mao, Qi Zhao 0037
Appl. Intell.1
2022 Embedded Self-Distillation in Compact Multibranch Ensemble Network for Remote Sensing Scene Classification
abstract
Remote sensing (RS) image scene classification task faces many challenges due to the interference from different characteristics of different geographical elements. To solve this problem, we propose a multi-branch ensemble network to enhance the feature representation ability by fusing features in final output logits and intermediate feature maps. However, simply adding branches will increase the complexity of models and decline the inference efficiency. On this issue, we embed self-distillation (SD) method to transfer knowledge from ensemble network to main-branch in it. Through optimizing with SD, main-branch will have close performance as ensemble network. During inference, we can cut other branches to simplify the whole model. In this paper, we first design compact multi-branch ensemble network, which can be trained in an end-to-end manner. Then, we insert SD method on output logits and feature maps. Compared to previous methods, our proposed architecture (ESD-MBENet) performs strongly on classification accuracy with compact design. Extensive experiments are applied on three benchmark RS datasets AID, NWPU-RESISC45 and UC-Merced with three classic baseline models, VGG16, ResNet50 and DenseNet121. Results prove that our proposed ESD-MBENet can achieve better accuracy than previous state-of-the-art (SOTA) complex models. Moreover, abundant visualization analysis make our method more convincing and interpretable.
Qi Zhao 0037, Yujing Ma, Shuchang Lyu, Lijiang Chen
IEEE Trans. Geosci. Remote. Sens.4
2021 Make Baseline Model Stronger: Embedded Knowledge Distillation in Weight-Sharing Based Ensemble Network
Shuchang Lyu, Qi Zhao 0037, Yujing Ma, Lijiang Chen
BMVC4
2021 Few-shot prototype alignment regularization network for document image layout segementation
Yujie Li 0001, Xing Xu 0001, Yi Lai, Fumin Shen, Lijiang Chen, Pengxiang Gao
Pattern Recognit.6
2020 Co-Clustering to Reveal Salient Facial Features for Expression Recognition
abstract
Facial expressions are a strong visual intimation of gestural behaviors. The intelligent ability to learn these non-verbal cues of the humans is the key characteristic to develop efficient human computer interaction systems. Extracting an effective representation from facial expression images is a crucial step that impacts the recognition accuracy. In this paper, we propose a novel feature selection strategy using singular value decomposition (SVD) based co-clustering to search for the most salient regions in terms of facial features that possess a high discriminating ability among all expressions. To the best of our knowledge, this is the first known attempt to explicitly perform co-clustering in the facial expression recognition domain. In our method, Gabor filters are used to extract local features from an image and then discriminant features are selected based on the class membership in co-clusters. Experiments demonstrate that co-clustering localizes the salient regions of the face image. Not only does the procedure reduce the dimensionality but also improves the recognition accuracy. Experiments on CK plus, JAFFE and MMI databases validate the existence and effectiveness of these learned facial features.
Sheheryar Khan, Lijiang Chen, Hong Yan 0001
IEEE Trans. Affect. Comput.2
2019 Sparsity Regularization Discriminant Projection for Feature Extraction
Sen Yuan, Xia Mao, Lijiang Chen
Neural Process. Lett.3
2019 Residual attention-based LSTM for video captioning
Zhilong Zhou, Lijiang Chen, Lianli Gao
World Wide Web3
2018 Learning deep features to recognise speech emotion using merged deep CNN
abstract
This study aims at learning deep features from different data to recognise speech emotion. The authors designed a merged convolutional neural network (CNN), which had two branches, one being one‐dimensional (1D) CNN branch and another 2D CNN branch, to learn the high‐level features from raw audio clips and log‐mel spectrograms. The building of the merged deep CNN consists of two steps. First, one 1D CNN and one 2D CNN architectures were designed and evaluated; then, after the deletion of the second dense layers, the two CNN architectures were merged together. To speed up the training of the merged CNN, transfer learning was introduced in the training. The 1D CNN and 2D CNN were trained first. Then, the learned features of the 1D CNN and 2D CNN were repurposed and transferred to the merged CNN. Finally, the merged deep CNN initialised with transferred features was fine‐tuned. Two hyperparameters of the designed architectures were chosen through Bayesian optimisation in the training. The experiments conducted on two benchmark datasets show that the merged deep CNN can improve emotion classification performance significantly.
Jianfeng Zhao 0005, Xia Mao, Lijiang Chen
IET Signal Process.3
2018 Elastic preserving projections based on L1-norm maximization
Sen Yuan, Xia Mao, Lijiang Chen
Multim. Tools Appl.3
2018 A closed-form solution to the graph total variation problem for continuous emotion profiling in noisy environment
Shaoling Jing, Xia Mao, Lijiang Chen, Maria Colomba Comes, Arianna Mencattini, Grazia Raguso, Fabien Ringeval, Björn W. Schuller, Corrado Di Natale, Eugenio Martinelli
Speech Commun.3
2017 Robust Math Formula Recognition in Degraded Chinese Document Images
abstract
In this paper, we study the problem of math formula recognition (MFR) in degraded Chinese document images. Compared to traditional optical character recognition (OCR), the MFR problem brings new challenges in terms of character segmentation and structural analysis, especially in degraded images. To tackle these issues, we propose an over-segmentation strategy to split and recognize adhesive formula elements based on convolutional neural network (CNN). In addition, we propose a hierarchical framework for formula structure analysis that constructs the formula in a top-down manner to iteratively split the regions into recognizable units. Due to the lack of degraded Chinese document images with math formulas in the community, we also harvest a diverse ground-truth dataset containing 100 images submitted from our system users. Extended experiments demonstrate the effectiveness and robustness of our proposed method in comparison with state-of-the-art methods.
Dongxiang Zhang, Long Guo, Lijiang Chen, Dengfeng Ke
ICDAR5
2017 An Iterative Refinement Framework for Image Document Binarization with Bhattacharyya Similarity Measure
abstract
Background noise and illumination condition are two primary factors degrading the performance of document image binarization. In this paper, we propose an iterative refinement framework to support robust binarization. Initially, an input image is transformed into a Bhattacharyya similarity matrix with Gaussian kernel, which is subsequently converted into a binary image using maximum entropy classifier. Then, we adopt the run-length histogram to estimate the character stroke width, an important indicator to determine the length of filter window. After noise elimination, the output image is used for the next round of refinement and the process terminates when the estimated stroke width is stable. Extensive experiments were conducted on the standard DIBCO datasets as well as a new benchmark harvested from our user query log. Results show that our proposed method outperforms state-of-the-art methods and is more robust to handle low-quality images.
Dongxiang Zhang, Dengfeng Ke, Long Guo, Shengkun Shi, Lijiang Chen
ICDAR9
2017 Human interaction recognition fusing multiple features of depth sequences
abstract
Human interaction recognition has played a major role in building intelligent video surveillance systems. Recently, depth data captured by the emerging RGB‐D sensors began to show its importability in human interaction recognition. This study proposes a novel framework for human interaction recognition using depth information including an algorithm to reconstruct depth sequence with as few key frames as possible. The proposed framework includes two essential modules. First, key frames extraction by sparse constraint, then the fusion multi‐feature, is constructed by using two types of available features and Max‐pooling, respectively. Finally, multiple features are directly sent to the SVM for the recognition of the human activity. This study explores the static and dynamic feature fusion method to improve the recognition performance with contextual relevance of continuous frames. A weight is used to fuse shape and optical flow features, which not only enhance the description capability of human behavioural characteristics in the spatiotemporal domain, but also effectively reduces the adverse impact of certain distortion point of interest for target recognition. Experimental results show that the proposed approach yields considerable performance improvement over the state‐of‐the‐art approaches with respect to accuracy on a public action dataset.
Xia Mao, Lijiang Chen
IET Comput. Vis.3
2017 Multilinear Spatial Discriminant Analysis for Dimensionality Reduction
abstract
In the last few years, great efforts have been made to extend the linear projection technique (LPT) for multidimensional data (i.e., tensor), generally referred to as the multilinear projection technique (MPT). The vectorized nature of LPT requires high-dimensional data to be converted into vector, and hence may lose spatial neighborhood information of raw data. MPT well addresses this problem by encoding multidimensional data as general tensors of a second or even higher order. In this paper, we propose a novel multilinear projection technique, called multilinear spatial discriminant analysis (MSDA), to identify the underlying manifold of high-order tensor data. MSDA considers both the nonlocal structure and the local structure of data in the transform domain, seeking to learn the projection matrices from all directions of tensor data that simultaneously maximize the nonlocal structure and minimize the local structure. Different from multilinear principal component analysis (MPCA) that aims to preserve the global structure and tensor locality preserving projection (TLPP) that is in favor of preserving the local structure, MSDA seeks a tradeoff between the nonlocal (global) and local structures so as to drive its discriminant information from the range of the non-local structure and the range of the local structure. This spatial discriminant characteristic makes MSDA have more powerful manifold preserving ability than TLPP and MPCA. Theoretical analysis shows that traditional MPTs, such as multilinear linear discriminant analysis, TLPP, MPCA, and tensor maximum margin criterion, could be derived from the MSDA model by setting different graphs and constraints. Extensive experiments on face databases (ORL, CMU PIE, and the extended Yale-B) and the Weizmann action database demonstrate the effectiveness of the proposed MSDA method.
Sen Yuan, Xia Mao, Lijiang Chen
IEEE Trans. Image Process.3
2016 Illumination compensation for facial feature point localization in a single 2D face image
Jizheng Yi, Xia Mao, Lijiang Chen, Alberto Rovetta
Neurocomputing3
2016 Text-Independent Phoneme Segmentation Combining EGG and Speech Data
abstract
A new approach for text-independent phoneme segmentation at sampling point level is proposed in this paper. The algorithm consists of two phases: First, the voiced sections in speech data are detected using the information of vocal folds vibration contained in electroglottograph (EGG). A Hilbert envelope feature is adopted to achieve sampling point level detection accuracy. Second, the voiced sections and other sections are treated separately. Each voiced section is divided into several candidate phonemes using the Viterbi algorithm. Then adjacent candidate phonemes are merged based on a Hotellings T-square test method. For other sections, the unvoiced consonants are detected from silence based on a singularity exponent feature. Comparison experiments show that the proposed method has better performance than the existing ones for a variety of tolerances, and is more robust to noise.
Lijiang Chen, Xia Mao, Hong Yan 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Trajectory-based view-invariant hand gesture recognition by fusing shape and orientation
abstract
Traditional studies in vision‐based hand gesture recognition remain rooted in view‐dependent representations, and hence users are forced to be fronto‐parallel to the camera. To solve this problem, view‐invariant gesture recognition aims to make the recognition result independent of viewpoint changes. However, in current works the view‐invariance is achieved at the price of mixing different gesture patterns that have similar trajectory curve shape but different semantic meanings. For example, the gesture ‘push’ can be mistaken as ‘drag’ from another viewpoint. To address this shortcoming, in this study, the authors use a shape descriptor to extract the view‐invariant features of a three‐dimensional (3D) trajectory. As the shape features are invariant to omnidirectional viewpoint changes, the orientation features are then added into weight different rotation angles so that similar trajectory shapes are better separated. The proposed method was conducted on two different databases, including a popular Australian Sign Language database and a challenging Kinect Hand Trajectory database. Experimental results show that the proposed algorithm achieves a higher average recognition rate than the state‐of‐the‐art approaches, and can better distinguish confusing gestures while meeting the view‐invariant condition.
Xia Mao, Lijiang Chen, Yu-Li Xue
IET Comput. Vis.3
2015 Kernel optimization using nonparametric Fisher criterion in the subspace
Xia Mao, Lijiang Chen, Yu-Li Xue, Alberto Rovetta
Pattern Recognit. Lett.3
2014 View-Invariant Gesture Recognition Using Nonparametric Shape Descriptor
abstract
In this paper we propose a new method for view-invariant gesture recognition, based on what we call nonparametric shape descriptor. We represent gestures as 3D motion trajectories and then we prove that the shape of a trajectory is equivalent to the Euclidean distances between all its points. The set of point-to-point distances description is mapped to a high-dimensional kernel space by kernel principal component analysis (KPCA), and then nonparametric discriminant analysis (NDA) is used to extract the view-invariant shape features as the input for pattern classification. The algorithm is performed on a public dataset, and shows better view-invariant performance than other state-of-the-art methods.
Xia Mao, Lijiang Chen, Yu-Li Xue, Angelo Compare
ICPR3
2014 Facial expression recognition considering individual differences in facial structure and texture
abstract
Facial expression recognition (FER) plays an important role in human–computer interaction. The recent years have witnessed an increasing trend of various approaches for the FER, but these approaches usually do not consider the effect of individual differences to the recognition result. When the face images change from neutral to a certain expression, the changing information constituted of the structural characteristics and the texture information can provide rich important clues not seen in either face image. Therefore it is believed to be of great importance for machine vision. This study proposes a novel FER algorithm by exploiting the structural characteristics and the texture information hiding in the image space. Firstly, the feature points are marked by an active appearance model. Secondly, three facial features, which are feature point distance ratio coefficient, connection angle ratio coefficient and skin deformation energy parameter, are proposed to eliminate the differences among the individuals. Finally, a radial basis function neural network is utilised as the classifier for the FER. Extensive experimental results on the Cohn–Kanade database and the Beihang University (BHU) facial expression database show the significant advantages of the proposed method over the existing ones.
Jizheng Yi, Xia Mao, Lijiang Chen, Yu-Li Xue, Angelo Compare
IET Comput. Vis.3
2013 VoteTrust: Leveraging friend invitation graph to defend against social network Sybils
abstract
Online social networks (OSNs) currently face a significant challenge by the existence and continuous creation of fake user accounts (Sybils), which can undermine the quality of social network service by introducing spam and manipulating online rating. Recently, there has been much excitement in the research community over exploiting social network structure to detect Sybils. However, they rely on the assumption that Sybils form a tight-knit community, which may not hold in real OSNs. In this paper, we present VoteTrust, a Sybil detection system that further leverages user interactions of initiating and accepting links. VoteTrust uses the techniques of trust-based vote assignment and global vote aggregation to evaluate the probability that the user is a Sybil. Using detailed evaluation on real social network (Renren), we show VoteTrust's ability to prevent Sybils gathering victims (e.g., spam audience) by sending a large amount of unsolicited friend requests and befriending many normal users, and demonstrate it can significantly outperform traditional ranking systems (such as TrustRank or BadRank) in Sybil detection.
Jilong Xue, Zhi Yang 0001, Xiaoyong Yang, Xiao Wang 0018, Lijiang Chen, Yafei Dai
INFOCOM5
2013 iPLUG: Personalized List Recommendation in Twitter
Lijiang Chen, Yibing Zhao, Shimin Chen, Hui Fang 0001, Chengkai Li 0001, Min Wang 0001
WISE (2)1
2013 Wiki3C: exploiting wikipedia for context-aware concept categorization
abstract
Wikipedia is an important human generated knowledge base containing over 21 million articles organized by millions of categories. In this paper, we exploit Wikipedia for a new task of text mining: Context-aware Concept Categorization. In the task, we focus on categorizing concepts according to their context. We exploit article link feature and category structure in Wikipedia, followed by introducing Wiki3C, an unsupervised and domain independent concept categorization approach based on context. In the approach, we investigate two strategies to select and filter Wikipedia articles for the category representation. Besides, a probabilistic model is employed to compute the semantic relatedness between two concepts in Wikipedia. Experimental evaluation using manually labeled ground truth shows that our proposed Wiki3C can achieve a noticeable improvement over the baselines without considering contextual information.
Peng Jiang 0002, Huiman Hou, Lijiang Chen, Shimin Chen, Conglei Yao, Chengkai Li 0001, Min Wang 0001
WSDM3
2013 PROSE: Proactive, Selective CDN Participation for P2P Streaming
Zhihui Lu 0002, Lijiang Chen, Jie Wu 0003, Da Deng 0002, Sijia Huang, Yi Huang 0020
J. Comput. Sci. Technol.2
2013 Speech Emotional Features Extraction Based on Electroglottograph
abstract
This study proposes two classes of speech emotional features extracted from electroglottography (EGG) and speech signal. The power-law distribution coefficients (PLDC) of voiced segments duration, pitch rise duration, and pitch down duration are obtained to reflect the information of vocal folds excitation. The real discrete cosine transform coefficients of the normalized spectrum of EGG and speech signal are calculated to reflect the information of vocal tract modulation. Two experiments are carried out. One is of proposed features and traditional features based on sequential forward floating search and sequential backward floating search. The other is the comparative emotion recognition based on support vector machine. The results show that proposed features are better than those commonly used in the case of speaker-independent and content-independent speech emotion recognition.
Lijiang Chen, Xia Mao, Pengfei Wei 0001, Angelo Compare
Neural Comput.1
2012 CPDID: A Novel CDN-P2P Dynamic Interactive Delivery Scheme for Live Streaming
abstract
Although many streaming application providers have relied on CDN services, there are several barriers to making CDN a more common service: expensive construction cost and fixed service mode. The rise of cloud computing requires "CDN as a Service" with a more open service mode, which requires content services available on-demand and be utilized in an open and loosely-coupled fashion. In most of the current streaming systems, CDN provides serve requesters with a passive and static state, there is no dynamic interaction with P2P systems, so the total CDN-P2P-Hybrid efficiency is not high. In this paper, we present CPDID: a CDN-P2P Dynamic Interactive Delivery scheme for Live Streaming. We explore CPDID architecture based on REST interface and JSON message format. And then we propose an identifying and selecting 'upload amplification nodes' algorithm to more efficiently utilize CDN resource. Our experimental results show that CPDID achieves at least 10-25% performance improvement compared with the existing native CDN-P2P-Hybrid schemes. At last, we analyze the prospective research direction and propose our future work.
Zhihui Lu 0002, Jie Wu 0003, Yi Huang 0020, Lijiang Chen, Da Deng 0002
ICPADS4
2012 Mandarin emotion recognition combining acoustic and emotional point information
Lijiang Chen, Xia Mao, Pengfei Wei 0001, Yu-Li Xue, Mitsuru Ishizuka
Appl. Intell.1
2011 Constrained Skyline Query Processing against Distributed Data Sites
abstract
The skyline of a multidimensional point set is a subset of interesting points that are not dominated by others. In this paper, we investigate constrained skyline queries in a large-scale unstructured distributed environment, where relevant data are distributed among geographically scattered sites. We first propose a partition algorithm that divides all data sites into incomparable groups such that the skyline computations in all groups can be parallelized without changing the final result. We then develop a novel algorithm framework called PaDSkyline for parallel skyline query processing among partitioned site groups. We also employ intragroup optimization and multifiltering technique to improve the skyline query processes within each group. In particular, multiple (local) skyline points are sent together with the query as filtering points, which help identify unqualified local skyline points early on a data site. In this way, the amount of data to be transmitted via network connections is reduced, and thus, the overall query response time is shortened further. Cost models and heuristics are proposed to guide the selection of a given number of filtering points from a superset. A cost-efficient model is developed to determine how many filtering points to use for a particular data site. The results of an extensive experimental study demonstrate that our proposals are effective and efficient.
Lijiang Chen, Bin Cui 0001, Hua Lu 0001
IEEE Trans. Knowl. Data Eng.1
2011 Approximate entity extraction in temporal databases
Wei Lu 0015, Gabriel Pui Cheong Fung, Xiaoyong Du 0001, Xiaofang Zhou 0001, Lijiang Chen
World Wide Web5
2010 Distributed Cache Indexing for Efficient Subspace Skyline Computation in P2P Networks
Lijiang Chen, Bin Cui 0001, Linhao Xu, Heng Tao Shen
DASFAA (1)1
2010 CPH-VoD: A Novel CDN-P2P-Hybrid Architecture Based VoD Scheme
Zhihui Lu 0002, Jie Wu 0003, Lijiang Chen, Sijia Huang, Yi Huang 0020
WISE3
2009 Efficient information retrieval in mobile peer-to-peer networks
abstract
Mobile devices have become indispensable in daily life, and hence how to take advantage of these portable and powerful facilities to share resources and information begins to emerge as an interesting problem. In this paper, we investigate the problem of information retrieval in a mobile peer-to-peer network. The prevailing approach to information retrieval is to apply flooding methods because of its quick response and easy maintenance. Obviously, this kind of approach wastes a huge amount of communication bandwidth which greatly affects the availability of the network, and the battery power which significantly shortens the serving time of mobile devices in the network. To tackle this problem, we propose a novel approach by mimicking different human behaviors of social networks, which takes advantages of Intelligence Accuracy (IA) mechanism that evaluates the distance from a node to certain resources in the network. Extensive experimental results show the efficiency and effectiveness of our approach as well as its scalability in a volatile environment.
Lijiang Chen, Bin Cui 0001, Heng Tao Shen, Wei Lu 0015, Xiaofang Zhou 0001
CIKM1
2009 Efficient Skyline Computation in Structured Peer-to-Peer Systems
abstract
An increasing number of large-scale applications exploit peer-to-peer network architecture to provide highly scalable and flexible services. Among these applications, data management in peer-to-peer systems is one of the interesting domains. In this paper, we investigate the multidimensional skyline computation problem on a structured peer-to-peer network. In order to achieve low communication cost and quick response time, we utilize the iMinMax(\theta ) method to transform high-dimensional data to one-dimensional value and distribute the data in a structured peer-to-peer network called BATON. Thereafter, we propose a progressive algorithm with adaptive filter technique for efficient skyline computation in this environment. We further discuss some optimization techniques for the algorithm, and summarize the key principles of our algorithm into a query routing protocol with detailed analysis. Finally, we conduct an extensive experimental evaluation to demonstrate the efficiency of our approach.
Bin Cui 0001, Lijiang Chen, Linhao Xu, Hua Lu 0001, Guojie Song, Quanqing Xu
IEEE Trans. Knowl. Data Eng.2
2008 iSky: Efficient and Progressive Skyline Computing in a Structured P2P Network
abstract
An interesting problem in peer-based data management is efficient support for skyline queries within a multiattribute space. A skyline query retrieves from a set of multidimensional data points a subset of interesting points, compared to which no other points are better. Skyline queries play an important role in multi-criteria decision making and user preference applications. In this paper, we address the skyline computing problem in a structured P2P network. We exploit the iMinMax(thetas) transformation to map high-dimensional data points to 1-dimensional values. All transformed data points are then distributed on a structured P2P network called BATON, where all peers are virtually organized as a balanced binary search tree. Subsequently, a progressive algorithm is proposed to compute skyline in the distributed P2P network. Further, we propose an adaptive skyline filtering technique to reduce both processing cost and communication cost during distributed skyline computing. Our performance study, with both synthetic and real datasets, shows that the proposed approach can dramatically reduce transferred data volume and gain quick response time.
Lijiang Chen, Bin Cui 0001, Hua Lu 0001, Linhao Xu, Quanqing Xu
ICDCS1
2008 Parallel Distributed Processing of Constrained Skyline Queries by Filtering
abstract
Skyline queries are capable of retrieving interesting points from a large data set according to multiple criteria. Most work on skyline queries so far has assumed a centralized storage, whereas in practice relevant data are often distributed among geographically scattered sites. In this work, we tackle constrained skyline queries in large-scale distributed environments without the assumption of any overlay structures, and propose a novel algorithm named PaDSkyline (Parallel distributed Skyline query processing). PaDSkyline significantly shortens the response time by performing parallel processing over site groups produced by a partition algorithm. Within each group, it locally optimizes the query processing over distributed sites. It also drastically enhances the network transmission efficiency by performing early reduction of skyline candidates with deliberately selected multiple filtering points. Results of extensive experiments demonstrate the efficiency and robustness of our proposals.
Bin Cui 0001, Hua Lu 0001, Quanqing Xu, Lijiang Chen, Yafei Dai, Yongluan Zhou
ICDE4