Shiming Xiang

dblp:81/6575 · DBLP profile ↗
← Back
14ranked-venue papers in the field
4as first author
5since 2021 · last 2024
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 9 (3 first)Information Retrieval & Web Search · 3Database Systems & Data Management · 1 (1 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2024 An Efficient Graph Autoencoder with Lightweight Desmoothing Decoder and Long-Range Modeling
abstract
Graph self-supervised learning provides a powerful guarantee for learning high-quality representations in an unsupervised manner. Despite its early birth, the performance of generative graph self-supervised learning has long lagged behind that of up-and-coming contrastive learning, especially on node classification tasks. In this paper, we investigate potential issues in existing graph autoencoders and attribute their poor performance to three main aspects: complex decoder design, lack of desmoothing process in feature remap, and overemphasis on local topological proximity. To tackle these issues, we propose an effective and efficient graph autoencoder framework for unsupervised representation learning, which contains two key components: lightweight smoothness-aware feature reconstructor and global structural dependency catcher. After performing a desmoothing operation on encoded representations via a learnable high-pass filter, the feature decoder reconstructs the original features through a simple linear projection. The lightweight design liberates the decoder from self-supervised pretext tasks and puts the encoder more accountable for achieving optimization objectives, which promotes effective training of the encoder. Global structural dependency catcher utilizes graph diffusion to build a structural regularization to capture long-range topological dependency on a graph. The empirical studies demonstrate the effectiveness of our approach, which can surpass dominant contrastive learning methods.
Jinyong Wen, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan
ICDM4
2023 Continual Semantic Segmentation via Scalable Contrastive Clustering and Background Diversity
abstract
Despite the efficacy towards static data distribution, traditional semantic segmentation methods encounter Catastrophic forgetting when tackling continually changing data streams. Another fundamental challenge is Background shift, which results from the semantic drift of the background class during continual learning steps. To extend the applicability of semantic segmentation methods, we introduce a novel, scalable segmentation architecture called ScaleSeg, designed to adapt the incremental scenarios. The architecture of ScaleSeg consists of a series of prototypes updated by online contrastive clustering. Additionally, we propose a background diversity strategy to enhance the model’s plasticity and stability, thus overcoming background shift. Comprehensive experiments and ablation studies on challenging benchmarks demonstrate that ScaleSeg surpasses previous state-of-the-art methods, particularly when dealing with extensive task sequences.
Qi Yang 0015, Xing Nie, Linsu Shi, Jiazhong Yu, Shiming Xiang
ICDM6
2022 Spatiotemporal Contextual Consistency Network for Precipitation Nowcasting
abstract
Precipitation nowcasting is forecasting rainfall in the short-term conditioned by the known meteorological parameters. Recently, deep neural networks (DNNs) have shown outstanding performance in this task. But, there are several challenges imposed by the multiple meteorological elements, including the multimodal modeling, the considerable variation in scales of precipitation region, as well as the long-tailed distribution of rainfall data. To solve these problems, this paper proposes Spatiotemporal Contextual Consistency Network (SCCN) for learning from the multi meteorological elements. Architecturally, a parameter-shared multimodal fusion CNN encoder, which dynamically exchanges features between different modalities, is used to encode the multimodal meteorological data. To improve the spatial modeling of the multiple meteorological features, we compose the multi-scale filters and deconstruction convolution to modify the gate operators in ConvLSTM to propose a spatial contextual consistency ConvLSTM (SCC-ConvLSTM). Furthermore, considering the temporal consistency in rainfall, a temporal consistency module (TCM) is designed to gear to long-tailed distribution. Under this module, different long-tailed meteorological elements are calculated to encode features and residuals fused with the previous precipitation distribution in sequence. The experimental results of precipitation nowcasting demonstrate the effectiveness of our method on the ERA5 dataset and WeatherBench dataset.
Xinyu Xiao, Qizhao Jin, Gaofeng Meng, Shiming Xiang, Chunhong Pan
ICDM4
2022 Adversarial Gradient Driven Exploration for Deep Click-Through Rate Prediction
abstract
Exploration-Exploitation (E& E) algorithms are commonly adopted to deal with the feedback-loop issue in large-scale online recommender systems. Most of existing studies believe that high uncertainty can be a good indicator of potential reward, and thus primarily focus on the estimation of model uncertainty. We argue that such an approach overlooks the subsequent effect of exploration on model training. From the perspective of online learning, the adoption of an exploration strategy would also affect the collecting of training data, which further influences model learning. To understand the interaction between exploration and training, we design a Pseudo-Exploration module that simulates the model updating process after a certain item is explored and the corresponding feedback is received. We further show that such a process is equivalent to adding an adversarial perturbation to the model input, and thereby name our proposed approach as an the Adversarial Gradient Driven Exploration (AGE). For production deployment, we propose a dynamic gating unit to pre-determine the utility of an exploration. This enables us to utilize the limited amount of resources for exploration, and avoid wasting pageview resources on ineffective exploration. The effectiveness of AGE was firstly examined through an extensive number of ablation studies on an academic dataset. Meanwhile, AGE has also been deployed to one of the world-leading display advertising platforms, and we observe significant improvements on various top-line evaluation metrics.
Kailun Wu, Weijie Bian, Zhangming Chan, Lejian Ren, Shiming Xiang, Shuguang Han, Hongbo Deng, Bo Zheng 0007
KDD5
2022 CAN: Feature Co-Action Network for Click-Through Rate Prediction
abstract
Feature interaction has been recognized as an important problem in machine learning, which is also very essential for click-through rate (CTR) prediction tasks. In recent years, Deep Neural Networks (DNNs) can automatically learn implicit nonlinear interactions from original sparse features, and therefore have been widely used in industrial CTR prediction tasks. However, the implicit feature interactions learned in DNNs cannot fully retain the complete representation capacity of the original and empirical feature interactions (e.g., cartesian product) without loss. For example, a simple attempt to learn the combination of feature A and feature B < A, B > as the explicit cartesian product representation of new features can outperform previous implicit feature interaction models including factorization machine (FM)-based models and their variations. This indicates there is still a big gap between explicit and implicit feature interaction models. However, to learn all the explicit feature interaction (cartesian product) representations requires a very large sample size along with N times of original parameter space (where N is quite large in most industrial applications). In this paper, we propose a Co-Action Network (CAN) to approximate the explicit pairwise feature interactions without introducing too many additional parameters. More specifically, giving feature A and its associated feature B, their feature interaction is modeled by learning two sets of parameters: 1) the embedding of feature A, and 2) a Multi-Layer Perceptron (MLP) to represent feature B. The approximated feature interaction can be obtained by passing the embedding of feature A through the MLP network of feature B. We refer to such pairwise feature interaction as feature co-action, and such a Co-Action Network unit can provide a very powerful capacity to fitting complex feature interactions. In addition, FM can be viewed as a special case of the CAN unit when the MLP is a single layer with only one output. Experimental results on public and industrial datasets show that CAN outperforms state-of-the-art CTR models and the cartesian product method. Moreover, CAN has been deployed in the display advertisement system in Alibaba, obtaining 12% improvement on CTR and 8% on Revenue Per Mille (RPM), which is a great improvement to the business. The code for experiments in this paper is open-sourced\footnotehttps://github.com/CAN-Paper/Co-Action-Network.
Weijie Bian, Kailun Wu, Lejian Ren, Qi Pi, Can Xiao, Xiang-Rong Sheng, Yong-Nan Zhu, Zhangming Chan, Na Mou, Xinchen Luo, Shiming Xiang, Guorui Zhou, Xiaoqiang Zhu, Hongbo Deng
WSDM12
2017 Deep Networks for Degraded Document Image Binarization through Pyramid Reconstruction
abstract
Binarization of document images is an important processing step for document images analysis and recognition. However, this problem is quite challenging in some cases because of the quality degradation of document images, such as varying illumination, complicated backgrounds, image noises due to ink spots, water stains or document creases. In this paper, we propose a framework based on deep convolutional neural-network (DCNN) for adaptive binarization of degraded document images. The basic idea of our method is to decompose a degraded document image into a spatial pyramid structure by using DCNN, with each layer at different scale. Then the foreground image is sequentially reconstructed from these layers in a coarse-to-fine manner by using deconvolutional network. Such kind of decomposition is quite beneficial, since multi-resolution supervision information can be directly introduced into network learning. We also define several loss functions about label consistency and foregrounds smoothing to further regularize the training of the network. Experimental results demonstrate the effectiveness of the proposed method.
Gaofeng Meng, Kun Yuan 0003, Ying Wu 0001, Shiming Xiang, Chunhong Pan
ICDAR4
2015 Cross-Modal Similarity Learning: A Low Rank Bilinear Formulation
abstract
The cross-media retrieval problem has received much attention in recent years due to the rapid increasing of multimedia data on the Internet. A new approach to the problem has been raised which intends to match features of different modalities directly. In this research, there are two critical issues: how to get rid of the heterogeneity between different modalities and how to match the cross-modal features of different dimensions. Recently metric learning methods show a good capability in learning a distance metric to explore the relationship between data points. However, the traditional metric learning algorithms only focus on single-modal features, which suffer difficulties in addressing the cross-modal features of different dimensions. In this paper, we propose a cross-modal similarity learning algorithm for the cross-modal feature matching. The proposed method takes a bilinear formulation, and with the nuclear-norm penalization, it achieves low-rank representation. Accordingly, the accelerated proximal gradient algorithm is successfully imported to find the optimal solution with a fast convergence rate O(1/t2). Experiments on three well known image-text cross-media retrieval databases show that the proposed method achieves the best performance compared to the state-of-the-art algorithms.
Cuicui Kang, Shengcai Liao, Yonghao He, Jian Wang 0068, Wenjia Niu, Shiming Xiang, Chunhong Pan
CIKM6
2015 Large Scale Image Annotation via Deep Representation Learning and Tag Embedding Learning
abstract
In this paper, we focus on the issue of large scale image annotation, whereas most existing methods are devised for small datasets. A novel model based on deep representation learning and tag embedding learning is proposed. Specifically, the proposed model learns an unified latent space for image visual features and tag embeddings simultaneously. Furthermore, a metric matrix is introduced to estimate the relevance scores between images and tags. Finally, an objective function modeling triplet relationships (irrelevant tag, image, relevant tag) is proposed with maximum margin pursuit. The proposed model is easy to tackle new images and tags via online learning and has a relatively low test computation complexity. Experimental results on NUS-WIDE dataset demonstrate the effectiveness of the proposed model.
Yonghao He, Jian Wang 0068, Cuicui Kang, Shiming Xiang, Chunhong Pan
ICMR4
2015 Image-Text Cross-Modal Retrieval via Modality-Specific Feature Learning
abstract
Cross-modal retrieval extends the ability of search engines to deal with the massive cross-modal data. The goal of image-text cross-modal retrieval is to search images (texts) by using text (image) queries by computing the similarities of images and texts directly. Many existing methods rely on low-level visual features and textual features for cross-modal retrieval, ignoring the characteristics existing in the raw data of different modalities. In this paper, a novel model based on modality-specific feature learning is proposed. Considering the characteristics of different modalities, the model uses two types of convolutional neural networks to map the raw data to the latent space representations for images and texts, respectively. Particularly, the convolution based network used for texts involves word embedding learning, which has been proved effective to extract meaningful textual features for text classification. In the latent space, the mapped features of images and texts form relevant and irrelevant image-text pairs, which are used by the one-vs-more learning scheme. This learning scheme can achieve ranking functionality by allowing for one relevant and more irrelevant pairs. The standard back-propagation technique is employed to update the parameters of two convolutional networks. Extensive cross-modal retrieval experiments are carried out on three challenging datasets that consist of image-document pairs or image-query click-through data from a search engine, and the results firmly demonstrate that the proposed model is much more effective.
Jian Wang 0068, Yonghao He, Cuicui Kang, Shiming Xiang, Chunhong Pan
ICMR4
2009 Subspace Regularization: A New Semi-supervised Learning Method
Yan-Ming Zhang 0001, Xinwen Hou, Shiming Xiang, Cheng-Lin Liu 0001
ECML/PKDD (2)3
2009 Embedding new data points for manifold learning via coordinate propagation
Shiming Xiang, Feiping Nie 0001, Yangqiu Song, Changshui Zhang, Chunxia Zhang 0001
Knowl. Inf. Syst.1
2009 Nonlinear Dimensionality Reduction with Local Spline Embedding
abstract
This paper presents a new algorithm for Nonlinear Dimensionality Reduction (NLDR). Our algorithm is developed under the conceptual framework of compatible mapping. Each such mapping is a compound of a tangent space projection and a group of splines. Tangent space projection is estimated at each data point on the manifold, through which the data point itself and its neighbors are represented in tangent space with local coordinates. Splines are then constructed to guarantee that each of the local coordinates can be mapped to its own single global coordinate with respect to the underlying manifold. Thus, the compatibility between local alignments is ensured. In such a work setting, we develop an optimization framework based on reconstruction error analysis, which can yield a global optimum. The proposed algorithm is also extended to embed out of samples via spline interpolation. Experiments on toy data sets and real-world data sets illustrate the validity of our method.
Shiming Xiang, Feiping Nie 0001, Changshui Zhang, Chunxia Zhang 0001
IEEE Trans. Knowl. Data Eng.1
2007 Embedding New Data Points for Manifold Learning Via Coordinate Propagation
Shiming Xiang, Feiping Nie 0001, Yangqiu Song, Changshui Zhang, Chunxia Zhang 0001
PAKDD1
2006 Spline Embedding for Nonlinear Dimensionality Reduction
Shiming Xiang, Feiping Nie 0001, Changshui Zhang, Chunxia Zhang 0001
ECML1