VLDB 2026 Research / reviewers in the wild / expert
Ruxin Wang 0002
dblp:149/7989-2
· DBLP profile ↗
36ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-2730-9409ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 6 since 2021Security and privacy · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Concept Drift and Long-Tailed Distribution in Fine-Grained Visual Categorization: Benchmark and MethodabstractData is the foundation for the development of computer vision, and the establishment of datasets plays an important role in advancing the techniques of fine-grained visual categorization (FGVC). In the existing FGVC datasets used in computer vision, it is generally assumed that each collected instance has fixed characteristics and the distribution of different categories is relatively balanced. In contrast, the real world scenario reveals the fact that the characteristics of instances tend to vary with time and exhibit a long-tailed distribution. Hence, the collected datasets may mislead the optimization of the fine-grained classifiers, resulting in unpleasant performance in real applications. Starting from the real-world conditions and to promote the practical progress of fine-grained visual categorization, we present a Concept Drift and Long-Tailed Distribution (CDLT) dataset. Specifically, the dataset is collected by gathering 11195 images of 250 instances in different species for 47 consecutive months in their natural contexts. The collection process involves dozens of crowd workers for photographing and domain experts for labeling. Meanwhile, we propose a feature recombination framework to address the learning challenges associated with CDLT. Experimental results validate the efficacy of our method while also highlighting the limitations of popular large vision-language models (e.g., CLIP) in the context of long-tailed distributions. This emphasizes the significance of CDLT as a benchmark for investigating these challenges. Shuo Ye, Shiming Chen 0002, Ruxin Wang 0002, Tianxu Wu, Salman Khan 0001, Fahad Shahbaz Khan, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | A comprehensive review of network pruning based on pruning granularity and pruning time perspectives
Kehan Zhu, Fuyi Hu, Yuanbing Ding, Wei Zhou 0011, Ruxin Wang 0002 |
Neurocomputing | 5 |
| 2024 | CNFA: Conditional Normalizing Flow for Query-Limited AttackabstractTraditional black-box attack methods rely on sufficient feedback from the victim model through a large number of queries until the attack is successful. This may not be acceptable in real applications, since the deployed system may be equipped with certain defense mechanisms and only return the final result (i.e., hard label) to the client. In contrast, one possible approach is formulating a hard label attack, which can be successfully executed within limited queries. To implement this idea, in this paper, we bypass the reliance on victim models and benefit from the intrinsic characteristics of adversarial examples (AEs) and the transferability of examples across different data-driven models. This motivates us to generatively reformulate the attack problem and propose a conditional normalized flow-based attack (CNFA), which builds up a statistical mapping from the benign example to its adversarial counterpart by tackling the conditional likelihood under the hard-label black-box setting. A well-trained CNFA model can directly and efficiently generate a batch of AEs for specific condition inputs. Extensive experiments validate the effectiveness of the proposed idea in a hard-label black-box setting and the superiority of CNFA over SOTA techniques. Renyang Liu 0001, Wei Zhou 0011, Haoran Li 0023, Ruxin Wang 0002 |
ICASSP | 5 |
| 2024 | Mitigating Label Noise on Graphs via Topological Sample SelectionabstractDespite the success of the carefully-annotated benchmarks, the effectiveness of existing graph neural networks (GNNs) can be considerably impaired in practice when the real-world graph data is noisily labeled. Previous explorations in sample selection have been demonstrated as an effective way for robust learning with noisy labels, however, the conventional studies focus on i.i.d data, and when moving to non-iid graph data and GNNs, two notable challenges remain: (1) nodes located near topological class boundaries are very informative for classification but cannot be successfully distinguished by the heuristic sample selection. (2) there is no available measure that considers the graph topological information to promote sample selection in a graph. To address this dilemma, we propose a $\textit{Topological Sample Selection}$ (TSS) method that boosts the informative sample selection process in a graph by utilising topological information. We theoretically prove that our procedure minimizes an upper bound of the expected risk under target clean distribution, and experimentally show the superiority of our method compared with state-of-the-art baselines. Jiangchao Yao, Xiaobo Xia, Jun Yu 0001, Ruxin Wang 0002, Bo Han 0003, Tongliang Liu |
ICML | 5 |
| 2024 | DTA: distribution transform-based attack for query-limited scenarioabstractAbstract In generating adversarial examples, the conventional black-box attack methods rely on sufficient feedback from the to-be-attacked models by repeatedly querying until the attack is successful, which usually results in thousands of trials during an attack. This may be unacceptable in real applications since Machine Learning as a Service Platform (MLaaS) usually only returns the final result (i.e., hard-label) to the client and a system equipped with certain defense mechanisms could easily detect malicious queries. By contrast, a feasible way is a hard-label attack that simulates an attacked action being permitted to conduct a limited number of queries. To implement this idea, in this paper, we bypass the dependency on the to-be-attacked model and benefit from the characteristics of the distributions of adversarial examples to reformulate the attack problem in a distribution transform manner and propose a distribution transform-based attack (DTA). DTA builds a statistical mapping from the benign example to its adversarial counterparts by tackling the conditional likelihood under the hard-label black-box settings. In this way, it is no longer necessary to query the target model frequently. A well-trained DTA model can directly and efficiently generate a batch of adversarial examples for a certain input, which can be used to attack un-seen models based on the assumed transferability. Furthermore, we surprisingly find that the well-trained DTA model is not sensitive to the semantic spaces of the training dataset, meaning that the model yields acceptable attack performance on other datasets. Extensive experiments validate the effectiveness of the proposed idea and the superiority of DTA over the state-of-the-art. Renyang Liu 0001, Wei Zhou 0011, Xin Jin 0005, Yuanyu Wang, Ruxin Wang 0002 |
Cybersecur. | 6 |
| 2024 | Hiding image with inception transformerabstractAbstract Image steganography aims to hide secret data in the cover media for covert communication. Though many deep‐learning‐based image steganography methods have been presented, these approaches suffer from the inefficiency of building long‐distance connections between the cover and secret images, leading to noticeable modification traces and poor steganalysis resistance. To improve the visual imperceptibility of generated stego images, it is essential to establish a global correlation between the cover and secret images. In this way, the secret image can be dispersed throughout the cover image globally. To bridge this gap, a novel image steganography framework called HiiT is proposed, which takes advantage of CNN and Transformer to learn both the local and global pixel correlation in image hiding. Specifically, a new Transformer structure called Inception Transformer is proposed, which incorporates the Inception Net in the attention‐based Transformer architecture. The Inception Net can learn different scaled image features using multiple convolution kernels, while the attention mechanism can learn the global pixel correlation. By this, the proposed Inception Transformer learns the long‐distance pixel dependency between the cover and secret images. Furthermore, we propose a ‘Skip Connection’ mechanism in the proposed Inception Transformer, which merges the low‐level visual features and high‐level semantic features and achieves better model performance. In detail, The HiiT generates higher‐quality stego images with 45.46 PSNR and 0.9915 SSIM. Besides, accurately restored secret images achieve 47.27 PSNR and 0.9952 SSIM. Extensive experimental results show the proposed HiiT significantly improves the image‐hiding performance compared with state‐of‐the‐art methods. Yunyun Dong, Ruxin Wang 0002, Bingbing Song, Tingchu Wei, Wei Zhou 0011 |
IET Image Process. | 3 |
| 2024 | An efficient training-from-scratch framework with BN-based structural compressor
Fuyi Hu, Wei Zhou 0011, Ruxin Wang 0002 |
Pattern Recognit. | 6 |
| 2023 | Improving the Adversarial Robustness of Object Detection with Contrastive Learning
Weiwei Zeng, Wei Zhou 0011, Yunyun Dong, Ruxin Wang 0002 |
PRCV (9) | 5 |
| 2023 | Adversarial attacks on multi-focus image fusion models
Xin Jin 0005, Xin Jin 0021, Ruxin Wang 0002, Shin-Jye Lee, Shaowen Yao 0001, Wei Zhou 0011 |
Comput. Secur. | 3 |
| 2023 | Detecting Adversarial Examples on Deep Neural Networks With Mutual Information Neural EstimationabstractDespite achieving exceptional performance, deep neural networks (DNNs) suffer from the harassment caused by adversarial examples, which are produced by corrupting clean examples with tiny perturbations. Many powerful defense methods have been presented such as training data augmentation and input reconstruction which, however, usually rely on the prior knowledge of the targeted models or attacks. A clean example and its adversarial version are very similar but have different high-level representations in a victim model. If we can obtain a space in which the representations of similar examples are also similar, then adversarial examples can be picked out by comparing the representations of input examples in this space and the high-level space of the victim model. Inspired by this, we propose a novel approach for detecting adversarial images, which can protect any pre-trained DNN classifiers and resist an endless stream of new attacks. Specifically, we first adopt a dual autoencoder to project images to a latent space. The dual autoencoder uses the self-supervised learning to ensure that small modifications to samples do not significantly alter their latent representations. Next, the mutual information neural estimation is utilized to enhance the discrimination of the latent representations. We then leverage the prior distribution matching to regularize the latent representations. To easily compare the representations of examples in the two spaces, and not rely on the prior knowledge of the targeted model, a simple fully connected neural network is used to embed the learned representations into an eigenspace, which is consistent with the output eigenspace of the targeted model. Through the distribution similarity of an input example in the two eigenspaces, we can judge whether the input example is adversarial or not. Extensive experiments on MNIST, CIFAR-10, and ImageNet show that the proposed method has superior defense performance and transferability than state-of-the-arts. Ruxin Wang 0002, Shui Yu 0001, Yunyun Dong, Shaowen Yao 0001, Wei Zhou 0011 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | Type-I Generative Adversarial AttackabstractDeep neural networks are vulnerable to adversarial attacks either by examples with indistinguishable perturbations which produce incorrect predictions, or by examples with noticeable transformations that are still predicted as the original label. The latter case is known as the Type I attack which, however, has achieved limited attention in literature. We advocate that the vulnerability comes from the ambiguous distributions among different classes in the resultant feature space of the model, which is saying that the examples with different appearances may present similar features. Inspired by this, we propose a novel Type I attack method called generative adversarial attack (GAA). Specifically, GAA aims at exploiting the distribution mapping from the source domain of multiple classes to the target domain of a single class by using generative adversarial networks. A novel loss and a U-net architecture with latent modification are elaborated to ensure the stable transformation between the two domains. In this way, the generated adversarial examples have similar appearances with examples of the target domain, yet obtaining the original prediction by the model being attacked. Extensive experiments on multiple benchmarks demonstrate that the proposed method generates adversarial images that are more visually similar to the target images than the competitors, and the state-of-the-art performance is achieved. Shenghong He, Ruxin Wang 0002, Tongliang Liu, Chao Yi, Xin Jin 0005, Renyang Liu 0001, Wei Zhou 0011 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | Securing Deep Learning as a Service Against Adaptive High Frequency Attacks With MMCATabstractMost cloud providers offer Deep Learning as a Service (DLaaS) for different business, science and engineering domains. However, it is known that deep neural networks (DNNs) are vulnerable to adversarial examples, which can cause well-trained DNN models to misbehave by injecting human-imperceptible perturbations to the query input data. Securing deep learning as a service becomes a critical challenge in mitigating such adversarial input perturbations, and enhancing the robustness of DNNs. In this article, we report two important facts: First, most adversarial perturbations are high frequency signals or are added to high frequency signals. Second, due to Frequency Principle that neural networks overly pay attention to fit the low frequency signals during training, the models could be easily misled by the high frequency signals of adversarial examples. These facts consequently contribute to the vulnerability of DNNs service in the Cloud. We conjecture that the more robust the neural networks are in learning from high frequency signals, the more resilient these neural networks are against adversarial perturbed examples. We propose a novel method for generating high-frequency-enhanced adversarial examples, which is achieved by a high-pass filter in the frequency domain via Fourier Transform. This method enhances the learning ability for high frequency signals and ameliorates to over-fit useless low frequency signals. In order to improve the robustness of DNNs service under such signal frequency attacks, we propose a multi-modal collaborative adversarial training framework, named as MMCAT, which uses the multi-modal information of the input images for cross-modal collaborative training, delivering excellent extension for effectively learning of multi-modal image information. Extensive experiments show that under strong adaptive frequency attacks, the DNNs service trained with the proposed MMCAT method achieve superior performance and robustness over the state-of-the-art adversarial training approaches. Bingbing Song, Ruxin Wang 0002, Yunyun Dong, Ling Liu 0001, Wei Zhou 0011 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | Towards Query-limited Adversarial Attacks on Graph Neural NetworksabstractGraph Neural Network (GNN) is a graph representation learning approach for graph-structured data, which has witnessed a remarkable progress in the past few years. As a counterpart, the robustness of such a model has also received considerable attention. Previous studies show that the performance of a well-trained GNN can be faded by black-box adversarial examples significantly. In practice, the attacker can only query the target model with very limited counts, yet the existing methods require hundreds of thousand queries to extend attacks, leading the attacker to be exposed easily. To perform a step forward in addressing this issue, in this paper, we propose a novel attack methods, namely Graph Query-limited Attack (GQA), in which we generate adversarial examples on the surrogate model to fool the target model. Specifically, in GQA, we use contrastive learning to fit the feature extraction layers of the surrogate model in a query-free manner, which can reduce the need of queries. Furthermore, in order to utilize query results sufficiently, we obtain a series of queries with rich information by changing the input iteratively, and storing them in a buffer for recycling usage. Experiments show that GQA can decrease the accuracy of the target model by 4.8%, with only 1% edges modified and 100 queries performed. Haoran Li 0023, Liwen Wu, Wei Zhou 0011, Ruxin Wang 0002 |
ICTAI | 6 |
| 2022 | Abstract Painting Synthesis via Decremental optimizationabstractAbstract Existing stroke‐based painting synthesis methods usually fail to achieve good results with limited strokes because these methods use semantically irrelevant metrics to calculate the similarity between the painting and photo domains. Hence, it is hard to see meaningful semantical information from the painting. This paper proposes a painting synthesis method that uses a CLIP (Contrastive‐Language‐Image‐Pretraining) model to build a semantically‐aware metric so that the cross‐domain semantic similarity is explicitly involved. To ensure the convergence of the objective function, we design a new strategy called decremental optimization. Specifically, we define painting as a set of strokes and use a neural renderer to obtain a rasterized painting by optimizing the stroke control parameters through a CLIP‐based loss. The optimization process is initialized with an excessive number of brush strokes, and the number of strokes is then gradually reduced to generate paintings of varying levels of abstraction. Experiments show that our method can obtain vivid paintings, and the results are better than the comparison stroke‐based painting synthesis methods when the number of strokes is limited. Zhengpeng Zhao, Dan Xu 0001, Qiuxia Yang, Ruxin Wang 0002 |
Comput. Graph. Forum | 7 |
| 2022 | CASR-Net: A color-aware super-resolution network for panchromatic image
Ling Liu 0010, Xin Jin 0005, Jianan Feng, Ruxin Wang 0002, Hangying Liao, Shin-Jye Lee, Shaowen Yao 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2022 | Adversarial UV-Transformation Texture Estimation for 3D Face AgingabstractFace aging aims to estimate aged facial textures given a certain face image. A number of 2D face-aging methods have been developed, but there have been few studies on 3D face aging, which would be valuable in several real-world applications. The lack of 3D face-aging data has had a significant impact on the development of 3D face aging, but we hypothesized that the large amounts of 2D face-aging data on the internet could be leveraged for 3D aged facial textures. In this paper, we propose a novel 3D aging framework, which we call UV-transformation texture estimation based on generative adversarial networks (UVTE-GAN), to achieve 3D face aging. Specifically, the proposed framework has three parts: 1) a 3D vertex and texture estimator, which accurately estimates the face’s spatial vertices and textures; 2) a texture-aging GAN, which is responsible for aging the estimated texture map via adversarial learning; and 3) a 2D & 3D rendering rebuilder, which recovers 2D & 3D faces using the estimated facial vertex map and aged facial texture map. In addition, we also design a plugin layer that allows us to train the whole model in an end-to-end manner. Experimental results demonstrate the effectiveness of the proposed method in synthesizing visually pleasing 3D aged face pictures, and state-of-the-art performance is achieved on several public datasets. Yiqiang Wu, Ruxin Wang 0002, Mingming Gong, Jun Cheng 0002, Zhengtao Yu 0001, Dapeng Tao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Remote Sensing Scene Classification Based on Attention-Enabled Progressively SearchingabstractRemote sensing image scene classification plays a significant role in remote sensing image analysis. Aiming at the problems of large transformation and scale variation of background and key objects in remote sensing images, we propose a neural architecture search (NAS) method based on attention search space. The network adaptively searches convolution, pooling, and attention operations in the appropriate layers. To ensure the stability of the searching process, a multistage network progressive fusion search method is proposed, which discards useless operations in stages, reduces the burden of search algorithm, and improves the search efficiency. Finally, paying attention to the association information between objects and scenes, a bottom-up multiscale fusion network connection strategy is proposed to fully reuse the semantics of multiscale feature maps in each stage. The experimental results show that the proposed method performs better than the manual method and the current neural network architecture search method. Junge Shen, Bin Cao 0006, Ruxin Wang 0002, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | LR-SVM+: Learning Using Privileged Information with Noisy LabelsabstractThe paradigm of Learning Using Privileged Information (LUPI) always assumes that labels are annotated precisely. However, in practice, this assumption may be violated, as the labels may be heavily noisy, which inevitably degenerates the performance of learning algorithms in the LUPI paradigm. To handle the side effect of noisy labels, we propose a novel Label Noise Robust SVM+ (LR-SVM+) algorithm. Specifically, as the privileged information contains rich information of the latent labels, we first utilize it to infer underlying clean labels. Then we use the inference to modify the noisy labels. Comprehensive experiments demonstrate the necessity of studying label noise robust SVM+ and the effectiveness of the proposed method. Zhengning Wu, Xiaobo Xia, Ruxin Wang 0002, Jun Yu 0001, Yinian Mao, Tongliang Liu |
IEEE Trans. Multim. | 3 |
| 2021 | Scientific Rumors Detection in Short Online TextsabstractThe huge amount of information on social media contains much false information that has not been confirmed or has been confirmed but not known by all users, which may cause improper public attention or mislead the lives of Internet users. Among the widely spread false information, there are many rumors that require professional knowledge to be distinguished. In this paper, through in-depth analysis of the data collected from Sina Weibo official false news refuting account and questionnaires sent out by us, we propose the indicator of the relative influence of scientific rumors, which demonstrates that the rumors which requires professional scientific knowledge for identification are small but have strong propagation ability. This makes it particularly important to use technical manners to automatically detect them from a great amount of information. Targeting at this, we build up an Internet rumor dataset which contains three categories: the first is called scientific rumors which needs to be clarified by professional knowledge; the second is named as social rumors that need to be validated by authoritative media, organizations, or officials; and the third is ordinary texts. A long short term memory (LSTM)-based model is developed to detect scientific rumors and social rumors among the collected data. A benchmark on the proposed dataset is provided by comparing multiple competitors, which shows that our model is competitive to the state-of-the-arts. As far as we know, this is the first work that detects scientific rumors uses deep learning tools. Ruxin Wang 0002, Xiaohui Cui |
SMC | 2 |
| 2021 | EnsembleFool: A method to generate adversarial examples based on model fusion strategy
Wenyu Peng, Renyang Liu 0001, Ruxin Wang 0002, Taining Cheng, Zifeng Wu, Wei Zhou 0011 |
Comput. Secur. | 3 |
| 2020 | Pairwise Similarity Regularization for Adversarial Domain AdaptationabstractDomain adaptation aims at learning a predictive model that can generalize to a new target domain different from the source (training) domain. To mitigate the domain gap, adversarial training has been developed to learn domain invariant representations. State-of-the-art methods further make use of pseudo labels generated by the source domain classifier to match conditional feature distributions between the source and target domains. However, if the target domain is more complex than the source domain, the pseudo labels are unreliable to characterize the class-conditional structure of the target domain data, undermining prediction performance. To resolve this issue, we propose a Pairwise Similarity Regularization (PSR) approach that exploits cluster structures of the target domain data and minimizes the divergence between the pairwise similarity of clustering partition and that of pseudo predictions. Therefore, PSR guarantees that two target instances in the same cluster have the same class prediction and thus eliminate the negative effect of unreliable pseudo labels. Extensive experimental results show that our PSR method significantly boosts the current adversarial domain adaptation methods by a large margin on four visual benchmarks. In particular, PSR achieves a remarkable improvement of more than 5% over the state-of-the-art on several hard-to-transfer tasks. Haotian Wang 0001, Wenjing Yang 0002, Ji Wang 0001, Ruxin Wang 0002, Long Lan, Mingyang Geng |
ACM Multimedia | 4 |
| 2020 | Receptive Field Size Versus Model Depth for Single Image Super-ResolutionabstractThe performance of single image super-resolution (SISR) has been largely improved by innovative designs of deep architectures. An important claim raised by these designs is that the deep models have large receptive field size and strong nonlinearity. However, we are concerned about the question that which factor, receptive field size or model depth, is more critical for SISR. Towards revealing the answers, in this paper, we propose a strategy based on dilated convolution to investigate how the two factors affect the performance of SISR. Our findings from exhaustive investigations suggest that SISR is more sensitive to the changes of receptive field size than to the model depth variations, and that the model depth must be congruent with the receptive field size to produce improved performance. These findings inspire us to design a shallower architecture which can save computational and memory cost while preserving comparable effectiveness with respect to a much deeper architecture. Ruxin Wang 0002, Mingming Gong, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2020 | A Cuboid CNN Model With an Attention Mechanism for Skeleton-Based Action RecognitionabstractThe introduction of depth sensors such as Microsoft Kinect have driven research in human action recognition. Human skeletal data collected from depth sensors convey a significant amount of information for action recognition. While there has been considerable progress in action recognition, most existing skeleton-based approaches neglect the fact that not all human body parts move during many actions, and they fail to consider the ordinal positions of body joints. Here, and motivated by the fact that an action's category is determined by local joint movements, we propose a cuboid model for skeleton-based action recognition. Specifically, a cuboid arranging strategy is developed to organize the pairwise displacements between all body joints to obtain a cuboid action representation. Such a representation is well structured and allows deep CNN models to focus analyses on actions. Moreover, an attention mechanism is exploited in the deep model, such that the most relevant features are extracted. Extensive experiments on our new Yunnan University-Chinese Academy of Sciences-Multimodal Human Action Dataset (CAS-YNU MHAD), the NTU RGB+D dataset, the UTD-MHAD dataset, and the UTKinect-Action3D dataset demonstrate the effectiveness of our method compared to the current state-of-the-art. Kaijun Zhu, Ruxin Wang 0002, Jun Cheng 0002, Dapeng Tao |
IEEE Trans. Multim. | 2 |
| 2019 | Embedded Block Residual Network: A Recursive Restoration Model for Single-Image Super-ResolutionabstractSingle-image super-resolution restores the lost structures and textures from low-resolved images, which has achieved extensive attention from the research community. The top performers in this field include deep or wide convolutional neural networks, or recurrent neural networks. However, the methods enforce a single model to process all kinds of textures and structures. A typical operation is that a certain layer restores the textures based on the ones recovered by the preceding layers, ignoring the characteristics of image textures. In this paper, we believe that the lower-frequency and higher-frequency information in images have different levels of complexity and should be restored by models of different representational capacity. Inspired by this, we propose a novel embedded block residual network (EBRN) which is an incremental recovering progress for texture super-resolution. Specifically, different modules in the model restores information of different frequencies. For lower-frequency information, we use shallower modules of the network to recover; for higher-frequency information, we use deeper modules to restore. Extensive experiments indicate that the proposed EBRN model achieves superior performance and visual improvements against the state-of-the-arts. Yajun Qiu, Ruxin Wang 0002, Dapeng Tao, Jun Cheng 0002 |
ICCV | 2 |
| 2019 | A tensor framework for geosensor data forecasting of significant societal events
Lihua Zhou, Guowang Du, Ruxin Wang 0002, Dapeng Tao, Lizhen Wang 0001, Jun Cheng 0002 |
Pattern Recognit. | 3 |
| 2019 | Dual-Transfer Face Sketch-Photo SynthesisabstractRecognizing the identity of a sketched face from a face photograph dataset is a critical yet challenging task in many applications, not least law enforcement and criminal investigations. An intelligent sketched face identification system would rely on automatic face sketch synthesis from photographs, thereby avoiding the cost of artists manually drawing sketches. However, conventional face sketch-photo synthesis methods tend to generate sketches that are consistent with the artists'drawing styles. Identity-specific information is often overlooked, leading to unsatisfactory identity verification and recognition performance. In this paper, we discuss the reasons why conventional methods fail to recover identity-specific information. Then, we propose a novel dual-transfer face sketch-photo synthesis framework composed of an inter-domain transfer process and an intra-domain transfer process. In the inter-domain transfer, a regressor of the test photograph with respect to the training photographs is learned and transferred to the sketch domain, ensuring the recovery of common facial structures during synthesis. In the intra-domain transfer, a mapping characterizing the relationship between photographs and sketches is learned and transferred across different identities, such that the loss of identity-specific information is suppressed during synthesis. The fusion of information recovered by the two processes is straightforward by virtue of an ad hoc information splitting strategy. We employ both linear and nonlinear formulations to instantiate the proposed framework. Experiments on The Chinese University of Hong Kong face sketch database demonstrate that compared to the current state-of-the-art the proposed framework produces more identifiable facial structures and yields higher face recognition performance in both the photo and sketch domains. Mingjin Zhang, Ruxin Wang 0002, Xinbo Gao 0001, Jie Li 0001, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2018 | Face Sketch Synthesis From Coarse to FineabstractSynthesizing fine face sketches from photos is a valuable yet challenging problem in digital entertainment. Face sketches synthesized by conventional methods usually exhibit coarse structures of faces, whereas fine details are lost especially on some critical facial components. In this paper, by imitating the coarse-to-fine drawing process of artists, we propose a novel face sketch synthesis framework consisting of a coarse stage and a fine stage. In the coarse stage, a mapping relationship between face photos and sketches is learned via the convolutional neural network. It ensures that the synthesized sketches keep coarse structures of faces. Given the test photo and the coarse synthesized sketch, a probabilistic graphic model is designed to synthesize the delicate face sketch which has fine and critical details. Experimental results on public face sketch databases illustrate that our proposed framework outperforms the state-of-the-art methods in both quantitive and visual comparisons. Mingjin Zhang, Nannan Wang 0001, Yunsong Li 0001, Ruxin Wang 0002, Xinbo Gao 0001 |
AAAI | 4 |
| 2018 | Joint medical image fusion, denoising and enhancement via discriminative low-rank sparse dictionaries learning
Huafeng Li 0001, Xiaoge He, Dapeng Tao, Yuan Yan Tang, Ruxin Wang 0002 |
Pattern Recognit. | 5 |
| 2018 | Training Very Deep CNNs for General Non-Blind DeconvolutionabstractNon-blind image deconvolution is an ill-posed problem. The presence of noise and band-limited blur kernels makes the solution of this problem non-unique. Existing deconvolution techniques produce a residual between the sharp image and the estimation that is highly correlated with the sharp image, the kernel, and the noise. In most cases, different restoration models must be constructed for different blur kernels and different levels of noise, resulting in low computational efficiency or highly redundant model parameters. Here we aim to develop a single model that handles different types of kernels and different levels of noise: general non-blind deconvolution. Specifically, we propose a very deep convolutional neural network that predicts the residual between a pre-deconvolved image and the sharp image rather than the sharp image. The residual learning strategy makes it easier to train a single model for different kernels and different levels of noise, encouraging high effectiveness and efficiency. Quantitative evaluations demonstrate the practical applicability of the proposed model for different blur kernels. The model also shows state-of-the-art performance on synthesized blurry images. Ruxin Wang 0002, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2018 | Multiclass Learning With Partially Corrupted LabelsabstractTraditional classification systems rely heavily on sufficient training data with accurate labels. However, the quality of the collected data depends on the labelers, among which inexperienced labelers may exist and produce unexpected labels that may degrade the performance of a learning system. In this paper, we investigate the multiclass classification problem where a certain amount of training examples are randomly labeled. Specifically, we show that this issue can be formulated as a label noise problem. To perform multiclass classification, we employ the widely used importance reweighting strategy to enable the learning on noisy data to more closely reflect the results on noise-free data. We illustrate the applicability of this strategy to any surrogate loss functions and to different classification settings. The proportion of randomly labeled examples is proved to be upper bounded and can be estimated under a mild condition. The convergence analysis ensures the consistency of the learned classifier to the optimal classifier with respect to clean data. Two instantiations of the proposed strategy are also introduced. Experiments on synthetic and real data verify that our approach yields improvements over the traditional classifiers as well as the robust classifiers. Moreover, we empirically demonstrate that the proposed strategy is effective even on asymmetrically noisy data. Ruxin Wang 0002, Tongliang Liu, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Coupled Deep Autoencoder for Single Image Super-ResolutionabstractSparse coding has been widely applied to learning-based single image super-resolution (SR) and has obtained promising performance by jointly learning effective representations for low-resolution (LR) and high-resolution (HR) image patch pairs. However, the resulting HR images often suffer from ringing, jaggy, and blurring artifacts due to the strong yet ad hoc assumptions that the LR image patch representation is equal to, is linear with, lies on a manifold similar to, or has the same support set as the corresponding HR image patch representation. Motivated by the success of deep learning, we develop a data-driven model coupled deep autoencoder (CDA) for single image SR. CDA is based on a new deep architecture and has high representational capability. CDA simultaneously learns the intrinsic representations of LR and HR image patches and a big-data-driven function that precisely maps these LR representations to their corresponding HR representations. Extensive experimentation demonstrates the superior effectiveness and efficiency of CDA for single image SR compared to other state-of-the-art methods on Set5 and Set14 datasets. Jun Yu 0002, Ruxin Wang 0002, Cuihua Li, Dacheng Tao |
IEEE Trans. Cybern. | 3 |
| 2016 | Superpixel-guided nonlocal means for image denoising and super-resolution
Ruxin Wang 0002, Jun Cheng 0002 |
Signal Process. | 3 |
| 2016 | Multi-View Object Retrieval via Multi-Scale Topic ModelsabstractThe increasing number of 3D objects in various applications has increased the requirement for effective and efficient 3D object retrieval methods, which attracted extensive research efforts in recent years. Existing works mainly focus on how to extract features and conduct object matching. With the increasing applications, 3D objects come from different areas. In such circumstances, how to conduct object retrieval becomes more important. To address this issue, we propose a multi-view object retrieval method using multi-scale topic models in this paper. In our method, multiple views are first extracted from each object, and then the dense visual features are extracted to represent each view. To represent the 3D object, multi-scale topic models are employed to extract the hidden relationship among these features with respect to varied topic numbers in the topic model. In this way, each object can be represented by a set of bag of topics. To compare the objects, we first conduct topic clustering for the basic topics from two data sets, and then generate the common topic dictionary for new representation. Then, the two objects can be aligned to the same common feature space for comparison. To evaluate the performance of the proposed method, experiments are conducted on two data sets. The 3D object retrieval experimental results and comparison with existing methods demonstrate the effectiveness of the proposed method. Richang Hong, Zhenzhen Hu 0004, Ruxin Wang 0002, Meng Wang 0001, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2016 | Non-Local Auto-Encoder With Collaborative Stabilization for Image RestorationabstractDeep neural networks have been applied to image restoration to achieve the top-level performance. From a neuroscience perspective, the layerwise abstraction of knowledge in a deep neural network can, to some extent, reveal the mechanisms of how visual cues are processed in human brain. A pivotal property of human brain is that similar visual cues can stimulate the same neuron to induce similar neurological signals. However, conventional neural networks do not consider this property, and the resulting models are, as a result, unstable regarding their internal propagation. In this paper, we develop the (stacked) non-local auto-encoder, which exploits self-similar information in natural images for stability. We propose that similar inputs should induce similar network propagation. This is achieved by constraining the difference between the hidden representations of non-local similar image blocks during training. By applying the proposed model to image restoration, we then develop a collaborative stabilization step to further rectify forward propagation. To obtain a reliable deep model, we employ several strategies to simplify training and improve testing. Extensive image restoration experiments, including image denoising and super-resolution, demonstrate the effectiveness of the proposed method. Ruxin Wang 0002, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2015 | Single Image Superresolution via Directional Group Sparsity and Directional FeaturesabstractSingle image superresolution (SR) aims to construct a high-resolution version from a single low-resolution (LR) image. The SR reconstruction is challenging because of the missing details in the given LR image. Thus, it is critical to explore and exploit effective prior knowledge for boosting the reconstruction performance. In this paper, we propose a novel SR method by exploiting both the directional group sparsity of the image gradients and the directional features in similarity weight estimation. The proposed SR approach is based on two observations: 1) most of the sharp edges are oriented in a limited number of directions and 2) an image pixel can be estimated by the weighted averaging of its neighbors. In consideration of these observations, we apply the curvelet transform to extract directional features which are then used for region selection and weight estimation. A combined total variation regularizer is presented which assumes that the gradients in natural images have a straightforward group sparsity structure. In addition, a directional nonlocal means regularization term takes pixel values and directional information into account to suppress unwanted artifacts. By assembling the designed regularization terms, we solve the SR problem of an energy function with minimal reconstruction error by applying a framework of templates for first-order conic solvers. The thorough quantitative and qualitative results in terms of peak signal-to-noise ratio, structural similarity, information fidelity criterion, and preference matrix demonstrate that the proposed approach achieves higher quality SR reconstruction than the state-of-the-art algorithms. Ruxin Wang 0002, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2014 | Diverse Expected Gradient Active Learning for Relative AttributesabstractThe use of relative attributes for semantic understanding of images and videos is a promising way to improve communication between humans and machines. However, it is extremely labor- and time-consuming to define multiple attributes for each instance in large amount of data. One option is to incorporate active learning, so that the informative samples can be actively discovered and then labeled. However, most existing active-learning methods select samples one at a time (serial mode), and may therefore lose efficiency when learning multiple attributes. In this paper, we propose a batch-mode active-learning method, called diverse expected gradient active learning. This method integrates an informativeness analysis and a diversity analysis to form a diverse batch of queries. Specifically, the informativeness analysis employs the expected pairwise gradient length as a measure of informativeness, while the diversity analysis forces a constraint on the proposed diverse gradient angle. Since simultaneous optimization of these two parts is intractable, we utilize a two-step procedure to obtain the diverse batch of queries. A heuristic method is also introduced to suppress imbalanced multiclass distributions. Empirical evaluations of three different databases demonstrate the effectiveness and efficiency of the proposed approach. Xinge You, Ruxin Wang 0002, Dacheng Tao |
IEEE Trans. Image Process. | 2 |