Jun Wang 0006

dblp:w/JunWang6 · DBLP profile ↗
← Back
99ranked-venue papers
12as first author
37since 2021 · last 2026
0000-0003-2787-8932ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 70 · 8 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 5 first-author · 17 since 2021Databases, data management, data science and information retrieval · 20 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
Yuge Huang, Yuxi Mi, Guodong Mu, Shouhong Ding, Jun Wang 0006, Rizen Guo, Shuigeng Zhou
AAAI7
2026 Attention Residual Fusion Network With Contrast for Source-Free Domain Adaptation
abstract
Source-free domain adaptation (SFDA) involves training a model on source domain and then applying it to a related target domain without access to the source data and labels during adaptation. The complexity of scene information and lack of the source domain make SFDA a difficult task. Recent studies have shown promising results, but many approaches to domain adaptation concentrate on domain shift and neglect the effects of negative transfer, which may impede enhancements of model performance during adaptation. In this paper, addressing this issue, we propose a novel framework of Attention Residual Fusion Network (ARFNet) based on contrast learning for SFDA to alleviate negative transfer and domain shift during the progress of adaptation, in which attention residual fusion, global-local attention contrast, and dynamic centroid evaluation are exploited. Concretely, the attention mechanism is first exploited to capture the discriminative region of the target object. Then, in each block, attention features are decomposed into spatial-wise and channel-wise attentions. The spatial-wise attentions are aggregated with original semantic features to achieve the cross-layer attention residual fusion progressively while the channel-wise attentions are exploited for self-distillation. During adaptation progress, we contrast global and local representations to improve the perceptual capabilities of different categories, which enables the model to discriminate variations between inner-class and intra-class. Finally, a dynamic centroid evaluation strategy is exploited to evaluate the trustworthy centroids and labels for self-supervised self-distillation, which aims to accurately approximate the center of the source domain and pseudo-labels to mitigate domain shift. To validate the efficacy of our methods, we execute comprehensive experiments on five benchmarks of varying scales, i.e., Office-31, Office-Home, VisDA-C, DomainNet-126, Cub-Paintings. Experimental outcomes indicate that our method surpasses other techniques, attaining superior performance across SFDA benchmarks.Code is available at https://github.com/RoryShao/ARFNet.git.
Renrong Shao, Wei Zhang 0056, Jun Wang 0006
IEEE Trans. Circuits Syst. Video Technol.3
2025 Coherency Improved Explainable Recommendation via Large Language Model
abstract
Explainable recommender systems are designed to elucidate the explanation behind each recommendation, enabling users to comprehend the underlying logic. Previous works perform rating prediction and explanation generation in a multi-task manner. However, these works suffer from incoherence between predicted ratings and explanations. To address the issue, we propose a novel framework that employs a large language model (LLM) to generate a rating, transforms it into a rating vector, and finally generates an explanation based on the rating vector and user-item information. Moreover, we propose utilizing publicly available LLMs and pre-trained sentiment analysis models to automatically evaluate the coherence without human annotations. Extensive experimental results on three datasets of explainable recommendation show that the proposed framework is effective, outperforming state-of-the-art baselines with improvements of 7.3% in explainability and 4.4% in text quality.
Ruixin Ding, Weihai Lu, Jun Wang 0006, Mo Yu, Wei Zhang 0056
AAAI4
2025 STAIR: Manipulating Collaborative and Multimodal Information for E-Commerce Recommendation
abstract
While the mining of modalities is the focus of most multimodal recommendation methods, we believe that how to fully utilize both collaborative and multimodal information is pivotal in e-commerce scenarios where, as clarified in this work, the user behaviors are rarely determined entirely by multimodal features. In order to combine the two distinct types of information, some additional challenges are encountered: 1) Modality erasure: Vanilla graph convolution, which proves rather useful in collaborative filtering, however erases multimodal information; 2) Modality forgetting: Multimodal information tends to be gradually forgotten as the recommendation loss essentially facilitates the learning of collaborative information. To this end, we propose a novel approach named STAIR, which employs a novel stepwise graph convolution to enable a co-existence of collaborative and multimodal information in e-commerce recommendation. Besides, it starts with the raw multimodal features as an initialization, and the forgetting problem can be significantly alleviated through constrained embedding updates. As a result, STAIR achieves state-of-the-art recommendation performance on three public e-commerce datasets with minimal computational and memory costs.
Cong Xu 0005, Yunhang He, Jun Wang 0006, Wei Zhang 0056
AAAI3
2025 From Enhancement to Understanding: Build a Generalized Bridge for Low-Light Vision via Semantically Consistent Unsupervised Fine-Tuning
Shao Zeng, Tianjun Gu, Zhizhong Zhang 0001, Shouhong Ding, Jun Wang 0006, Xin Tan 0002, Yuan Xie 0006, Lizhuang Ma
ICCV8
2025 Negotiated Reasoning: On Provably Addressing Relative Over-Generalization
Junjie Sheng, Wenhao Li 0001, Bo Jin 0003, Hongyuan Zha, Jun Wang 0006, Xiangfeng Wang 0001
AAMAS5
2025 Dream2Drive: A Large Language Model Powered Agent for Real World Multi-task Vehicle Motion Planning
abstract
Vehicle motion planning is a critical component of Autonomous Driving (AD). While conventional motion planning methods have shown satisfactory performance on mainstream public datasets, they struggle to generalize to private and diverse real-world scenarios. They also struggle to adapt to multi-task scenarios due to the lack of common sense and understanding of real-world scenarios. Additionally, they lack the ability to accumulate past driving experience. To address these issues, we present Dream2Drive, a highly adaptable multi-task motion planning agent based on Large Language Models (LLMs), which exploits the inherent knowledge of the LLMs and long and short-term memory experience. Dream2Drive can plan different tasks without needing adaptation, leveraging five key capabilities: Observation, Short-term Memory, Long-term Experience, Planning, and Rethinking. To validate the adaptability of Dream2Drive to real-world conditions, we collect a novel dataset DreamReal including 225 scenarios from the real world. Extensive experiments are conducted to demonstrate the effectiveness and compatibility of Dream2Drive, providing valuable insights and references for multi-task motion planning of vehicles in private and diverse real-world scenarios.
Hongru Wang 0014, Jun Wang 0006
IJCNN2
2025 Collaborative Filtering Meets Spectrum Shift: Connecting User-Item Interaction with Graph-Structured Side Information
abstract
Graph Neural Networks (GNNs) have demonstrated their superiority in collaborative filtering, where the user-item (U-I) interaction bipartite graph serves as the fundamental data format. However, when graph-structured side information (e.g., multimodal similarity graphs or social networks) is integrated into the U-I bipartite graph, existing graph collaborative filtering methods fall short of achieving satisfactory performance. We quantitatively analyze this problem from a spectral perspective. Recall that a bipartite graph possesses a full spectrum within the range of [-1, 1], with the highest frequency exactly achievable at -1 and the lowest frequency at 1; however, we observe as more side information is incorporated, the highest frequency of the augmented adjacency matrix progressively shifts rightward. This spectrum shift phenomenon has caused previous approaches built for the full spectrum [-1, 1] to assign mismatched importance to different frequencies. To this end, we propose Spectrum Shift Correction (dubbed SSC), incorporating shifting and scaling factors to enable spectral GNNs to adapt to the shifted spectrum. Unlike previous paradigms of leveraging side information, which necessitate tailored designs for diverse data types, SSC directly connects traditional graph collaborative filtering with any graph-structured side information. Experiments on social and multimodal recommendation demonstrate the effectiveness of SSC, achieving relative improvements of up to 23% without incurring any additional computational overhead. Our code is available at https://github.com/yhhe2004/SSC-KDD.
Yunhang He, Cong Xu 0005, Jun Wang 0006, Wei Zhang 0056
KDD (2)3
2025 Switchable Token-Specific Codebook Quantization For Face Image Compression
abstract
With the ever-increasing volume of visual data, the efficient and lossless transmission, along with its subsequent interpretation and understanding, has become a critical bottleneck in modern information systems. The emerged codebook-based solution utilize a globally shared codebook to quantize and dequantize each token, controlling the bpp by adjusting the number of tokens or the codebook size. However, for facial images—which are rich in attributes—such global codebook strategies overlook both the category-specific correlations within images and the semantic differences among tokens, resulting in suboptimal performance, especially at low bpp. Motivated by these observations, we propose a Switchable Token-Specific Codebook Quantization for face image compression, which learns distinct codebook groups for different image categories and assigns an independent codebook to each token. By recording the codebook group to which each token belongs with a small number of bits, our method can reduce the loss incurred when decreasing the size of each codebook group. This enables a larger total number of codebooks under a lower overall bpp, thereby enhancing the expressive capability and improving reconstruction performance. Owing to its generalizable design, our method can be integrated into any existing codebook-based representation learning approach and has demonstrated its effectiveness on face recognition datasets, achieving an average accuracy of 93.51\% for reconstructed images at 0.05 bpp.
Guodong Mu, Jun Wang 0006, Yuan Xie 0001, Zhizhong Zhang 0001, Shouhong Ding
NeurIPS7
2025 Consistent Assistant Domains Transformer for Source-Free Domain Adaptation
abstract
Source-free domain adaptation (SFDA) aims to address the challenge of adapting to a target domain without accessing the source domain directly. However, due to the inaccessibility of source domain data, deterministic invariable features cannot be obtained. Current mainstream methods primarily focus on evaluating invariant features in the target domain that closely resemble those in the source domain, subsequently aligning the target domain with the source domain. However, these methods are susceptible to hard samples and influenced by domain bias. In this paper, we propose a Consistent Assistant Domains Transformer for SFDA, abbreviated as CADTrans, which solves the issue by constructing invariable feature representations of domain consistency. Concretely, we develop an assistant domain module for CADTrans to obtain diversified representations from the intermediate aggregated global attentions, which addresses the limitation of existing methods in adequately representing diversity. Based on assistant and target domains, invariable feature representations are obtained by multiple consistent strategies, which can be used to distinguish easy and hard samples. Finally, to align the hard samples to the corresponding easy samples, we construct a conditional multi-kernel max mean discrepancy (CMK-MMD) strategy to distinguish between samples of the same category and those of different categories. Extensive experiments are conducted on various benchmarks such as Office-31, Office-Home, VISDA-C, and DomainNet-126, proving the significant performance improvements achieved by our proposed approaches. Code is available at https://github.com/RoryShao/CADTrans.git.
Renrong Shao, Wei Zhang 0056, Kangyang Luo, Qin Li 0002, Jun Wang 0006
IEEE Trans. Image Process.5
2025 Explainable Session-Based Recommendation via Path Reasoning
abstract
This paper explores explaining session-based recommendation (SR) by path reasoning. Current SR models emphasize accuracy but lack explainability, while traditional path reasoning prioritizes knowledge graph exploration, ignoring sequential patterns present in the session history. Therefore, we propose a generalized hierarchical reinforcement learning framework for SR, which improves the explainability of existing SR models via Path Reasoning, namely PR4SR. Considering the different importance of items to the session, we design the session-level agent to select the items in the session as the starting nodes for path reasoning and the path-level agent to perform path reasoning. In particular, we design a multi-target reward mechanism to adapt to the skip behaviors of sequential patterns in SR and introduce path midpoint reward to enhance the exploration efficiency and accuracy in knowledge graphs. To improve the knowledge graph’s completeness and diversify the paths of explanation, we incorporate extracted feature information from images into the knowledge graph. We instantiate PR4SR in five state-of-the-art SR models (i.e., GRU4REC, NARM, GCSAN, SR-GNN, SASRec) and compare it with other explainable SR frameworks to demonstrate the effectiveness of PR4SR for recommendation and explanation tasks through extensive experiments with these approaches on four datasets.
Yang Cao 0021, Shuo Shang, Jun Wang 0006, Wei Zhang 0056
IEEE Trans. Knowl. Data Eng.3
2025 Understanding Adversarial Robustness From Feature Maps of Convolutional Layers
abstract
The adversarial robustness of a neural network mainly relies on two factors: model capacity and antiperturbation ability. In this article, we study the antiperturbation ability of the network from the feature maps of convolutional layers. Our theoretical analysis discovers that larger convolutional feature maps before average pooling can contribute to better resistance to perturbations, but the conclusion is not true for max pooling. It brings new inspiration to the design of robust neural networks and urges us to apply these findings to improve existing architectures. The proposed modifications are very simple and only require upsampling the inputs or slightly modifying the stride configurations of downsampling operators. We verify our approaches on several benchmark neural network architectures, including AlexNet, VGG, RestNet18, and PreActResNet18. Nontrivial improvements in terms of both natural accuracy and adversarial robustness can be achieved under various attack and defense mechanisms. The code is available at https://github.com/MTandHJ/rcm.
Cong Xu 0005, Wei Zhang 0056, Jun Wang 0006, Min Yang 0009
IEEE Trans. Neural Networks Learn. Syst.3
2024 D2LLM: Decomposed and Distilled Large Language Models for Semantic Search
abstract
The key challenge in semantic search is to create models that are both accurate and efficient in pinpointing relevant sentences for queries.While BERT-style bi-encoders excel in efficiency with pre-computed embeddings, they often miss subtle nuances in search tasks.Conversely, GPT-style LLMs with crossencoder designs capture these nuances but are computationally intensive, hindering realtime applications.In this paper, we present D2LLMs-Decomposed and Distilled LLMs for semantic search-that combines the best of both worlds.We decompose a cross-encoder into an efficient bi-encoder integrated with Pooling by Multihead Attention and an Interaction Emulation Module, achieving nuanced understanding and pre-computability.Knowledge from the LLM is distilled into this model using contrastive, rank, and feature imitation techniques.Our experiments show that D2LLM surpasses five leading baselines in terms of all metrics across three tasks, particularly improving NLI task performance by at least 6.45%.
Hang Yu 0002, Jun Wang 0006, Wei Zhang 0056
ACL (1)4
2024 SDPose: Tokenized Pose Estimation via Circulation-Guide Self-Distillation
abstract
Recently, transformer-based methods have achieved state-of-the-art prediction quality on human pose estimation(HPE). Nonetheless, most of these top-performing transformer-based models are too computation-consuming and storage-demanding to deploy on edge computing platforms. Those transformer-based models that require fewer resources are prone to under-fitting due to their smaller scale and thus perform notably worse than their larger counterparts. Given this conundrum, we introduce SD-Pose, a new self-distillation method for improving the performance of small transformer-based models. To mitigate the problem of under-fitting, we design a transformer module named Multi-Cycled Transformer(MCT) based on multiple-cycled forwards to more fully exploit the potential of small model parameters. Further, in order to prevent the additional inference compute-consuming brought by MCT, we introduce a self-distillation scheme, extracting the knowledge from the MCT module to a naive forward model. Specifically, on the MSCOCO validation dataset, SDPose-T obtains 69.7% mAP with 4.4M parameters and 1.8 GFLOPs. Furthermore, SDPose-S-V2 obtains 73.5% mAP on the MSCOCO validation dataset with 6.2M parameters and 4.7 GFLOPs, achieving a new state-of-the-art among predominant tiny neural network methods.
Sichen Chen, Siming Huang, Ran Yi 0002, Peixian Chen, Jun Wang 0006, Shouhong Ding, Lizhuang Ma
CVPR8
2024 Privacy-Preserving Face Recognition Using Trainable Feature Subtraction
abstract
The widespread adoption of face recognition has led to increasing privacy concerns, as unauthorized access to face images can expose sensitive personal information. This paper explores face image protection against viewing and recovery attacks. Inspired by image compression, we propose creating a visually uninformative face image through feature subtraction between an original face and its model-produced regeneration. Recognizable identity features within the image are encouraged by co-training a recognition model on its high-dimensional feature represen-tation. To enhance privacy, the high-dimensional represen-tation is crafted through random channel shuffling, resulting in randomized recognizable images devoid of attacker-leverageable texture details. We distill our methodologies into a novel privacy-preserving face recognition method, MinusFace. Experiments demonstrate its high recognition accuracy and effective privacy protection. Its code is avail-able at https://github.com/Tencent/TFace.
Yuxi Mi, Zhizhou Zhong, Yuge Huang, Jiazhen Ji, Jianqing Xu, Jun Wang 0006, Shaoming Wang, Shouhong Ding, Shuigeng Zhou
CVPR6
2024 DMT: Comprehensive Distillation with Multiple Self-Supervised Teachers
abstract
Numerous self-supervised learning paradigms, such as contrastive learning and masked image modeling, have been proposed to acquire powerful and general representations from unlabeled data. However, these models are commonly pretrained within their specific framework alone, failing to consider the complementary nature of visual representations. To tackle this issue, we introduce Comprehensive Distillation with Multiple Self-supervised Teachers (DMT) for pretrained model compression, which leverages the strengths of multiple off-the-shelf self-supervised models. Our experimental results on prominent benchmark datasets exhibit that the proposed method significantly surpasses state-of-the-art competitors while retaining favorable efficiency metrics. On classification tasks, our DMT framework utilizing three different self-supervised ViT-Base teachers enhances the performance of both small/tiny models and the base model itself. For dense tasks, DMT elevates the AP/mIoU of standard SSL models on MS-COCO and ADE20K datasets by 4.0%.
Yuang Liu, Jing Wang 0224, Qiang Zhou 0001, Fan Wang 0019, Jun Wang 0006, Wei Zhang 0056
ICASSP5
2024 Graph-enhanced Optimizers for Structure-aware Recommendation Embedding Evolution
abstract
Embedding plays a key role in modern recommender systems because they are virtual representations of real-world entities and the foundation for subsequent decision-making models. In this paper, we propose a novel embedding update mechanism, Structure-aware Embedding Evolution (SEvo for short), to encourage related nodes to evolve similarly at each step. Unlike GNN (Graph Neural Network) that typically serves as an intermediate module, SEvo is able to directly inject graph structural information into embedding with minimal computational overhead during training. The convergence properties of SEvo along with its potential variants are theoretically analyzed to justify the validity of the designs. Moreover, SEvo can be seamlessly integrated into existing optimizers for state-of-the-art performance. Particularly SEvo-enhanced AdamW with moment estimate correction demonstrates consistent improvements across a spectrum of models and datasets, suggesting a novel technical route to effectively utilize graph structural information beyond explicit GNN modules.
Cong Xu 0005, Jun Wang 0006, Jianyong Wang 0001, Wei Zhang 0056
NeurIPS2
2024 Dynamic Token-Pass Transformers for Semantic Segmentation
abstract
Vision transformers (ViT) usually extract features via forwarding all the tokens in the self-attention layers from top to toe. In this paper, we introduce dynamic token-pass vision transformers (DoViT) for semantic segmentation, which can adaptively reduce the inference cost for images with different complexity. DoViT gradually stops partial easy tokens from self-attention calculation and keeps the hard tokens forwarding until meeting the stopping criteria. We employ lightweight auxiliary heads to make the token-pass decision and divide the tokens into keeping/stopping parts. With a token separate calculation, the self-attention layers are speeded up with sparse tokens and still work friendly with hardware. A token reconstruction module is built to collect and reset the grouped tokens to their original position in the sequence, which is necessary to predict correct semantic masks. We conduct extensive experiments on two common semantic segmentation tasks, and demonstrate that our method greatly reduces about 40% ∼ 60% FLOPs and the drop of mIoU is within 0.8% for various segmentation transformers. The throughput and inference speed of ViT-L/B are increased to more than 2× on Cityscapes. Code is available at https://github.com/FLHonker/DoViT-code.
Yuang Liu, Qiang Zhou 0001, Jing Wang 0224, Zhibin Wang 0004, Fan Wang 0019, Jun Wang 0006, Wei Zhang 0056
WACV6
2024 Modeling Dynamic Item Tendency Bias in Sequential Recommendation With Causal Intervention
abstract
Sequential recommendation is a critical but challenging task in capturing users’ potential preferences due to inherent biases in the data. Existing debiasing recommendation methods aim to eliminate biases from historical interaction data collected by recommender systems and have shown promising results. However, there is another significant bias that hinders the improvement of sequential recommendation models: dynamic item tendency bias. This bias arises because a period might have some unique tendencies consisting of items interacted with by users with the same intent, leading to a dynamic tendency distribution that biases the model training towards these tendencies. To address this issue, we propose a causal approach to model dynamic item tendency bias in sequential recommendation. We first extract tendencies on carefully designed item-item graphs through community detection. We then use causal intervention to conduct deconfounded training to capture true user preferences and introduce the beneficial item tendency bias to the inference process through optimal transport techniques. Experimental results on four real-world datasets demonstrate that our proposed method consistently outperforms state-of-the-art debiasing recommendation methods, confirming that our model is effective in reducing dynamic item tendency bias and dealing with tendency drifts.
Shuo Shang, Jun Wang 0006, Wei Zhang 0056
IEEE Trans. Knowl. Data Eng.4
2024 StableGCN: Decoupling and Reconciling Information Propagation for Collaborative Filtering
abstract
Graph Convolutional Networks (GCNs) have been widely applied to collaborative filtering, where each layer typically contains neighborhood aggregation and feature transformation. Recent studies have found that feature transformation contributes little to the final recommendation performance. They however eliminated it directly without further exploration, leading to a degradation of model expressive power. In this paper, we show that this problem arises from inconsistent information propagation process, in which the dominance of feature transformation prevents features from being properly smoothed by neighborhood aggregation. To this end, we present StableGCN to decouple and reconcile this contradictory process in an orderly rather than intertwined manner. The coarse-grained node features are first refined by an elaborate extractor, and then smoothed by a specific kind of GCN concerning feature denoising. Consequently, feature transformation and neighborhood aggregation can support each other without sacrificing expressive power. Extensive experiments on six public datasets demonstrate the effectiveness and state-of-the-art performance of StableGCN.
Cong Xu 0005, Jun Wang 0006, Wei Zhang 0056
IEEE Trans. Knowl. Data Eng.2
2023 Self-Decoupling and Ensemble Distillation for Efficient Segmentation
abstract
Knowledge distillation (KD) is a promising teacher-student learning paradigm that transfers information from a cumbersome teacher to a student network. To avoid the training cost of a large teacher network, the recent studies propose to distill knowledge from the student itself, called Self-KD. However, due to the limitations of the performance and capacity of the student, the soft-labels or features distilled by the student barely provide reliable guidance. Moreover, most of the Self-KD algorithms are specific to classification tasks based on soft-labels, and not suitable for semantic segmentation. To alleviate these contradictions, we revisit the label and feature distillation problem in segmentation, and propose Self-Decoupling and Ensemble Distillation for Efficient Segmentation (SDES). Specifically, we design a decoupled prediction ensemble distillation (DPED) algorithm that generates reliable soft-labels with multiple expert decoders, and a decoupled feature ensemble distillation (DFED) mechanism to utilize more important channel-wise feature maps for encoder learning. The extensive experiments on three public segmentation datasets demonstrate the superiority of our approach and the efficacy of each component in the framework through the ablation study.
Yuang Liu, Wei Zhang 0056, Jun Wang 0006
AAAI3
2023 DistilPose: Tokenized Pose Regression with Heatmap Distillation
abstract
In the field of human pose estimation, regression-based methods have been dominated in terms of speed, while heatmap-based methods are far ahead in terms of performance. How to take advantage of both schemes remains a challenging problem. In this paper, we propose a novel human pose estimation framework termed DistilPose, which bridges the gaps between heatmap-based and regression-based methods. Specifically, DistilPose maximizes the transfer of knowledge from the teacher model (heatmap-based) to the student model (regression-based) through Token-distilling Encoder (TDE) and Simulated Heatmaps. TDE aligns the feature spaces of heatmap-based and regression-based models by introducing tokenization, while Simulated Heatmaps transfer explicit guidance (distribution and confidence) from teacher heatmaps into student models. Extensive experiments show that the proposed DistilPose can significantly improve the performance of the regression-based models while maintaining efficiency. Specifically, on the MSCOCO validation dataset, DistilPose-S obtains 71.6% mAP with 5.36M parameters, 2.38 GFLOPs, and 40.2 FPS, which saves 12.95×, 7.16× computational cost and is 4.9× faster than its teacher model with only 0.9 points performance drop. Furthermore, DistilPose-L obtains 74.4% mAP on MSCOCO validation dataset, achieving a new state-of-the-art among predominant regression-based models. Code will be available at https://github.com/yshMars/DistilPose.
Suhang Ye, Jie Hu 0018, Liujuan Cao, Shengchuan Zhang, Jun Wang 0006, Shouhong Ding, Rongrong Ji
CVPR7
2023 Data-free Knowledge Distillation for Fine-grained Visual Categorization
abstract
Data-free knowledge distillation (DFKD) is a promising approach for addressing issues related to model compression, security privacy, and transmission restrictions. Although the existing methods exploiting DFKD have achieved inspiring achievements in coarse-grained classification, in practical applications involving fine-grained classification tasks that require more detailed distinctions between similar categories, sub-optimal results are obtained. To address this issue, we propose an approach called DFKD-FGVC that extends DFKD to fine-grained visual categorization (FGVC) tasks. Our approach utilizes an adversarial distillation framework with attention generator, mixed high-order attention distillation, and semantic feature contrast learning. Specifically, we introduce a spatial-wise attention mechanism to the generator to synthesize fine-grained images with more details of discriminative parts. We also utilize the mixed high-order attention mechanism to capture complex interactions among parts and the subtle differences among discriminative features of the fine-grained categories, paying attention to both local features and semantic context relationships. Moreover, we leverage the teacher and student models of the distillation framework to contrast high-level semantic feature maps in the hyperspace, comparing variances of different categories. We evaluate our approach on three widely-used FGVC benchmarks (Aircraft, Cars196, and CUB200) and demonstrate its superior performance. Code is available at https://github.com/RoryShao/DFKD-FGVC.git
Renrong Shao, Wei Zhang 0056, Jianhua Yin 0001, Jun Wang 0006
ICCV4
2023 Learning Cooperative Oversubscription for Cloud by Chance-Constrained Multi-Agent Reinforcement Learning
abstract
Oversubscription is a common practice for improving cloud resource utilization. It allows the cloud service provider to sell more resources than the physical limit, assuming not all users would fully utilize the resources simultaneously. However, how to design an oversubscription policy that improves utilization while satisfying some safety constraints remains an open problem. Existing methods and industrial practices are over-conservative, ignoring the coordination of diverse resource usage patterns and probabilistic constraints. To address these two limitations, this paper formulates the oversubscription for cloud as a chance-constrained optimization problem and proposes an effective Chance-Constrained Multi-Agent Reinforcement Learning (C2MARL) method to solve this problem. Specifically, C2MARL reduces the number of constraints by considering their upper bounds and leverages a multi-agent reinforcement learning paradigm to learn a safe and optimal coordination policy. We evaluate our C2MARL on an internal cloud platform and public cloud datasets. Experiments show that our C2MARL outperforms existing methods in improving utilization () under different levels of safety constraints.
Junjie Sheng, Lu Wang 0029, Fangkai Yang, Bo Qiao 0001, Hang Dong 0004, Xiangfeng Wang 0001, Bo Jin 0003, Jun Wang 0006, Si Qin, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001
WWW8
2023 Quantize Sequential Recommenders Without Private Data
abstract
Deep neural networks have achieved great success in sequential recommendation systems. While maintaining high competence in user modeling and next-item recommendation, these models have long been plagued by the numerous parameters and computation, which inhibit them to be deployed on resource-constrained mobile devices. Model quantization, as one of the main paradigms for compression techniques, converts float parameters to low-bit values to reduce parameter redundancy and accelerate inference. To avoid drastic performance degradation, it usually requests a fine-tuning phase with an original dataset. However, the training set of user-item interactions is not always available due to transmission limits or privacy concerns. In this paper, we propose a novel framework to quantize sequential recommenders without access to any real private data. A generator is employed in the framework to synthesize fake sequence samples to feed the quantized sequential recommendation model and minimize the gap with a full-precision sequential recommendation model. The generator and the quantized model are optimized with a min-max game — alternating discrepancy estimation and knowledge transfer. Moreover, we devise a two-level discrepancy modeling strategy to transfer information between the quantized model and the full-precision model. The extensive experiments of various recommendation networks on three public datasets demonstrate the effectiveness of the proposed framework.
Lingfeng Shi, Yuang Liu, Jun Wang 0006, Wei Zhang 0056
WWW3
2023 Conditional pseudo-supervised contrast for data-Free knowledge distillation
Renrong Shao, Wei Zhang 0056, Jun Wang 0006
Pattern Recognit.3
2022 Multi-Knowledge Aggregation and Transfer for Semantic Segmentation
abstract
As a popular deep neural networks (DNN) compression technique, knowledge distillation (KD) has attracted increasing attentions recently. Existing KD methods usually utilize one kind of knowledge in an intermediate layer of DNN for classification tasks to transfer useful information from cumbersome teacher networks to compact student networks. However, this paradigm is not very suitable for semantic segmentation, a comprehensive vision task based on both pixel-level and contextual information, since it cannot provide rich information for distillation. In this paper, we propose a novel multi-knowledge aggregation and transfer (MKAT) framework to comprehensively distill knowledge within an intermediate layer for semantic segmentation. Specifically, the proposed framework consists of three parts: Independent Transformers and Encoders module (ITE), Auxiliary Prediction Branch (APB), and Mutual Label Calibration (MLC) mechanism, which can take advantage of abundant knowledge from intermediate features. To demonstrate the effectiveness of our proposed approach, we conduct extensive experiments on three segmentation datasets: Pascal VOC, Cityscapes, and CamVid, showing that MKAT outperforms the other KD methods.
Yuang Liu, Wei Zhang 0056, Jun Wang 0006
AAAI3
2022 Structural Landmarking and Interaction Modelling: A "SLIM" Network for Graph Classification
abstract
Graph neural networks are a promising architecture for learning and inference with graph-structured data. Yet, how to generate informative, fixed dimensional features for graphs with varying size and topology can still be challenging. Typically, this is achieved through graph-pooling, which summarizes a graph by compressing all its nodes into a single vector. Is such a “collapsing-style” graph-pooling the only choice for graph classification? From complex system’s point of view, properties of a complex system arise largely from the interaction among its components. Therefore, we speculate that preserving the interacting relation between parts, instead of pooling them together, could benefit system level prediction. To verify this, we propose SLIM, a graph neural network model for Structural Landmarking and Interaction Modelling. The main idea is to compute a set of end-to-end optimizable sub-structure landmarks, so that any input graph can be projected onto these (spatially) local structural representatives for a faithful, global characterization. By doing so, explicit interaction between component parts of a graph can be leveraged directly in generating discriminative graph representation. Encouraging results are observed on benchmark datasets for graph classification, demonstrating the value of interaction modelling in the design of graph neural networks.
Yaokang Zhu, Kai Zhang 0001, Jun Wang 0006, Haibin Ling, Jie Zhang 0012, Hongyuan Zha
AAAI3
2022 Multi-view Pre-trained Model for Code Vulnerability Identification
Xuxiang Jiang, Yinhao Xiao, Jun Wang 0006, Wei Zhang 0056
WASA (3)3
2022 Learning structured communication for multi-agent reinforcement learning
Junjie Sheng, Xiangfeng Wang 0001, Bo Jin 0003, Junchi Yan, Wenhao Li 0001, Tsung-Hui Chang, Jun Wang 0006, Hongyuan Zha
Auton. Agents Multi Agent Syst.7
2022 Node Embedding and Classification with Adaptive Structural Fingerprint
Yaokang Zhu, Jun Wang 0006, Jie Zhang 0012, Kai Zhang 0001
Neurocomputing2
2022 Learning to schedule multi-NUMA virtual machines via reinforcement learning
Junjie Sheng, Yiqiu Hu, Bo Jin 0003, Jun Wang 0006, Xiangfeng Wang 0001
Pattern Recognit.6
2021 Source-Free Domain Adaptation for Semantic Segmentation
abstract
Unsupervised Domain Adaptation (UDA) can tackle the challenge that convolutional neural network (CNN)-based approaches for semantic segmentation heavily rely on the pixel-level annotated data, which is labor-intensive. However, existing UDA approaches in this regard inevitably require the full access to source datasets to reduce the gap between the source and target domains during model adaptation, which are impractical in the real scenarios where the source datasets are private, and thus cannot be released along with the well-trained source models. To cope with this issue, we propose a source-free domain adaptation framework for semantic segmentation, namely SFDA, in which only a well-trained source model and an unlabeled target domain dataset are available for adaptation. SFDA not only enables to recover and preserve the source domain knowledge from the source model via knowledge transfer during model adaptation, but also distills valuable information from the target domain for self-supervised learning. The pixel-and patch-level optimization objectives tailored for semantic segmentation are seamlessly integrated in the framework. The extensive experimental results on numerous benchmark datasets highlight the effectiveness of our framework against the existing UDA approaches relying on source data.
Yuang Liu, Wei Zhang 0056, Jun Wang 0006
CVPR3
2021 Zero-Shot Adversarial Quantization
abstract
Model quantization is a promising approach to compress deep neural networks and accelerate inference, making it possible to be deployed on mobile and edge devices. To retain the high performance of full-precision models, most existing quantization methods focus on fine-tuning quantized model by assuming training datasets are accessible. However, this assumption sometimes is not satisfied in real situations due to data privacy and security issues, thereby making these quantization methods not applicable. To achieve zero-short model quantization without accessing training data, a tiny number of quantization methods adopt either post-training quantization or batch normalization statistics-guided data generation for fine-tuning. However, both of them inevitably suffer from low performance, since the former is a little too empirical and lacks training support for ultra-low precision quantization, while the latter could not fully restore the peculiarities of original data and is often low efficient for diverse data generation. To address the above issues, we propose a zero-shot adversarial quantization (ZAQ) framework, facilitating effective discrepancy estimation and knowledge transfer from a full-precision model to its quantized model. This is achieved by a novel two-level discrepancy modeling to drive a generator to synthesize informative and diverse data examples to optimize the quantized model in an adversarial learning fashion. We conduct extensive experiments on three fundamental vision tasks, demonstrating the superiority of ZAQ over the strong zero-shot baselines and validating the effectiveness of its main components. Code is available at https://git.io/Jqc0y.
Yuang Liu, Wei Zhang 0056, Jun Wang 0006
CVPR3
2021 Empirical or Invariant Risk Minimization? A Sample Complexity Perspective
Kartik Ahuja, Jun Wang 0006, Amit Dhurandhar, Karthikeyan Shanmugam 0001, Kush R. Varshney
ICLR2
2021 Defense against Adversarial Attacks with an Induced Class
abstract
Though deep neural networks have succeeded in various real applications, the prediction performance is significantly degraded when facing adversarial attacks. In this work, we investigate the alternation of the prediction distribution pattern under adversarial attacks and argue that such alternation is the primary reason for performance drop. To this end, we propose a simple yet effective method by introducing an induced class to attract the adversarial attack and thus protect the original classes' prediction order. Experiments on two real-world datasets demonstrate that the proposed method can maintain the prediction performance for both natural and adversarial examples.
Jun Wang 0006, Jian Pu
IJCNN2
2021 Parallel pathway dense neural network with weighted fusion structure for brain tumor segmentation
Fangyan Ye, Yingbin Zheng, Hao Ye 0005, Xiaohao Han, Jun Wang 0006, Jian Pu
Neurocomputing6
2020 Adaptive Structural Fingerprints for Graph Attention Networks
Kai Zhang 0001, Yaokang Zhu, Jun Wang 0006, Jie Zhang 0012
ICLR3
2020 VSB2-Net: Visual-Semantic Bi-Branch Network for Zero-Shot Hashing
abstract
Zero-shot hashing aims at learning hashing model from seen classes and the obtained model is capable of generalizing to unseen classes for image retrieval. Inspired by zero-shot learning, existing zero-shot hashing methods usually transfer the supervised knowledge from seen to unseen classes, by embedding the hamming space to a shared semantic space. However, this makes instances difficult to distinguish due to limited hashing bit numbers, especially for semantically similar unseen classes. We propose a novel inductive zero-shot hashing framework, i.e., VSB2-Net, where both semantic space and visual feature space are embedded to the same hamming space instead. The reconstructive semantic relationships are established in the hamming space, preserving local similarity relationships and explicitly enlarging the discrepancy between semantic hamming vectors. A two-task architecture, comprising of classification module and visual feature reconstruction module, is employed to enhance the generalization and transfer abilities. Extensive evaluation results on several benchmark datasets demonstrate the superiority of our proposed method compared to several state-of-the-art baselines.
Xiangfeng Wang 0001, Bo Jin 0003, Jun Wang 0006, Hongyuan Zha
ICPR5
2020 Adaptive multi-teacher multi-level knowledge distillation
Yuang Liu, Wei Zhang 0056, Jun Wang 0006
Neurocomputing3
2019 The Kelly Growth Optimal Portfolio with Ensemble Learning
abstract
As a competitive alternative to the Markowitz mean-variance portfolio, the Kelly growth optimal portfolio has drawn sufficient attention in investment science. While the growth optimal portfolio is theoretically guaranteed to dominate any other portfolio with probability 1 in the long run, it practically tends to be highly risky in the short term. Moreover, empirical analysis and performance enhancement studies under practical settings are surprisingly short. In particular, how to handle the challenging but realistic condition with insufficient training data has barely been investigated. In order to fill voids, especially grappling with the difficulty from small samples, in this paper, we propose a growth optimal portfolio strategy equipped with ensemble learning. We synergically leverage the bootstrap aggregating algorithm and the random subspace method into portfolio construction to mitigate estimation error. We analyze the behavior and hyperparameter selection of the proposed strategy by simulation, and then corroborate its effectiveness by comparing its out-of-sample performance with those of 10 competing strategies on four datasets. Experimental results lucidly confirm that the new strategy has superiority in extensive evaluation criteria.
Weiwei Shen, Jian Pu, Jun Wang 0006
AAAI4
2019 False positive rate control for positive unlabeled learning
Shuchen Kong, Weiwei Shen, Yingbin Zheng, Jian Pu, Jun Wang 0006
Neurocomputing6
2019 Scaling Up Kernel SVM on Limited Resources: A Low-Rank Linearization Approach
abstract
Kernel support vector machines (SVMs) deliver state-of-the-art results in many real-world nonlinear classification problems, but the computational cost can be quite demanding in order to maintain a large number of support vectors. Linear SVM, on the other hand, is highly scalable to large data but only suited for linearly separable problems. In this paper, we propose a novel approach called low-rank linearized SVM to scale up kernel SVM on limited resources. Our approach transforms a nonlinear SVM to a linear one via an approximate empirical kernel map computed from efficient kernel low-rank decompositions. We theoretically analyze the gap between the solutions of the approximate and optimal rank- k kernel map, which in turn provides guidance on the sampling scheme of the Nyström approximation. Furthermore, we extend it to a semisupervised metric learning scenario in which partially labeled samples can be exploited to further improve the quality of the low-rank embedding. Our approach inherits rich representability of kernel SVM and high efficiency of linear SVM. Experimental results demonstrate that our approach is more robust and achieves a better tradeoff between model representability and scalability against state-of-the-art algorithms for large-scale SVMs.
Liang Lan, Shandian Zhe, Wei Cheng 0002, Jun Wang 0006, Kai Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2018 Tau-FPL: Tolerance-Constrained Learning in Linear Time
Nan Li 0019, Jian Pu, Jun Wang 0006, Junchi Yan, Hongyuan Zha
AAAI4
2018 Hybrid Deep Sequential Modeling for Social Text-Driven Stock Prediction
abstract
In addition to only considering stocks' price series, utilizing short and instant texts from social medias like Twitter has potential to yield better stock market prediction. While some previous approaches have explored this direction, their results are still far from satisfactory due to their reliance on performance of sentiment analysis and limited capabilities of learning direct relations between target stock trends and their daily social texts. To bridge this gap, we propose a novel Cross-modal attention based Hybrid Recurrent Neural Network (CH-RNN), which is inspired by the recent proposed DA-RNN model. Specifically, CH-RNN consists of two essential modules. One adopts DA-RNN to gain stock trend representations for different stocks. The other utilizes recurrent neural network to model daily aggregated social texts. These two modules interact seamlessly by the following two manners: 1) daily representations of target stock trends from the first module are leveraged to select trend-related social texts through a cross-modal attention mechanism, and 2) representations of text sequences and trend series are further integrated. The comprehensive experiments on the real dataset we build demonstrate the effectiveness of CH-RNN and benefit of considering social texts.
Huizhe Wu, Wei Zhang 0056, Weiwei Shen, Jun Wang 0006
CIKM4
2018 Factorization Meets Memory Network: Learning to Predict Activity Popularity
Wen Wang 0016, Wei Zhang 0056, Jun Wang 0006
DASFAA (2)3
2018 Learning Sequential Correlation for User Generated Textual Content Popularity Prediction
abstract
Popularity prediction of user generated textual content is critical for prioritizing information in the web, which alleviates heavy information overload for ordinary readers. Most previous studies model each content instance separately for prediction and thus overlook the sequential correlations between instances of a specific user. In this paper, we go deeper into this problem based on the two observations for each user, i.e., sequential content correlation and sequential popularity correlation. We propose a novel deep sequential model called User Memory-augmented recurrent Attention Network (UMAN). This model encodes the two correlations by updating external user memories which is further leveraged for target text representation learning and popularity prediction. The experimental results on several real-world datasets validate the benefits of considering these correlations and demonstrate UMAN achieves best performance among several strong competitors.
Wen Wang 0016, Wei Zhang 0056, Jun Wang 0006, Junchi Yan, Hongyuan Zha
IJCAI3
2018 User-guided Hierarchical Attention Network for Multi-modal Social Image Popularity Prediction
abstract
Popularity prediction for the growing social images has opened unprecedented opportunities for wide commercial applications, such as precision advertising and recommender system. While a few studies have explored this significant task, little research has addressed its unstructured properties of both visual and textual modalities, and further considered to learn effective representation from multi-modalities for popularity prediction. To this end, we propose a model named User-guided Hierarchical Attention Network (UHAN) with two novel user-guided attention mechanisms to hierarchically attend both visual and textual modalities. It is capable of not only learning effective representation for each modality, but also fusing them to obtain an integrated multi-modal representation under the guidance of user embedding. As no benchmark dataset exists, we extend a publicly available social image dataset by adding the descriptions of images. The comprehensive experiments have demonstrated the rationality of our proposed UHAN and its better performance than several strong alternatives.
Wei Zhang 0056, Wen Wang 0016, Jun Wang 0006, Hongyuan Zha
WWW3
2018 Multiple graph regularized graph transduction via greedy gradient Max-Cut
Yu Xiu, Weiwei Shen, Zhongqun Wang, Sanmin Liu, Jun Wang 0006
Inf. Sci.5
2018 Exploiting Feature and Class Relationships in Video Categorization with Regularized Deep Neural Networks
abstract
In this paper, we study the challenging problem of categorizing videos according to high-level semantics such as the existence of a particular human action or a complex event. Although extensive efforts have been devoted in recent years, most existing works combined multiple video features using simple fusion strategies and neglected the utilization of inter-class semantic relationships. This paper proposes a novel unified framework that jointly exploits the feature relationships and the class relationships for improved categorization performance. Specifically, these two types of relationships are estimated and utilized by imposing regularizations in the learning process of a deep neural network (DNN). Through arming the DNN with better capability of harnessing both the feature and the class relationships, the proposed regularized DNN (rDNN) is more suitable for modeling video semantics. We show that rDNN produces better performance over several state-of-the-art approaches. Competitive results are reported on the well-known Hollywood2 and Columbia Consumer Video benchmarks. In addition, to stimulate future research on large scale video categorization, we collect and release a new benchmark dataset, called FCVID, which contains 91,223 Internet videos and 239 manually annotated categories.
Yu-Gang Jiang 0001, Zuxuan Wu, Jun Wang 0006, Xiangyang Xue 0001, Shih-Fu Chang
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Boosting Alzheimer diagnosis accuracy with the help of incomplete privileged information
abstract
Early and accurate diagnosis of Alzheimer's disease is beneficial to both preserve daily functioning and test possible new treatments. However, current diagnosis depends on dozens of factors, including the family member, past medical problems, tests of memory, blood and urine tests, brain scans and even cerebrospinal fluid specimens. Among them, regular features (e.g., blood and urine tests, brain scans) are simple and accurate. Privileged features (e.g., tests of memory, cerebrospinal fluid specimens) are inconvenient and uncomfortable to acquire, and thus incomplete in most cases. In this work, we propose a two-stage learning framework to predict merely using the regular feature with the aid of incomplete privileged data in the training time. In particular, we first complement the missing data of privileged features by exploring the relationship with regular features and labels. The recovered privileged features, regular features, as well as labels, are combined as a new set of privileged features. Then privileged learning via matching logits is applied to boost the diagnosis accuracy. Our experiments and comparison studies with competing techniques on both synthetic data and real benchmarks have corroborated the effectiveness and superiority of the proposed framework for biomedical applications.
Jian Pu, Jun Wang 0006, Yingbin Zheng, Hao Ye 0005, Weiwei Shen, Hongyuan Zha
BIBM2
2017 Glioma grading based on 3D multimodal convolutional neural network and privileged learning
abstract
Brain tumors, especially high-grade gliomas, are one of the most lethal cancers for humankind today. Early and accurate diagnosis of tumor grading is the key for subsequent therapy and treatment. In the past, conventional computer-aided diagnosis relies on handcrafted features from magnetic resonance images (MRI), which are usually inaccurate and laborious. Recently, deep neural networks have been developed and applied for tumor segmentation and classification. However, most existing methods consider 3D MRI as a series of 2D images and use a simple modality fusion method via feature concatenation. In this paper, we propose an end-to-end 3-dimensional convolutional neural network (3D CNN) with gated multimodal unit (GMU) fusion to integrate the information both in three dimensions and in multiple modalities. Specifically, 3D convolutional kernels are directly applied to the whole MRI images, gathering the abnormalities in sagittal, axial and coronal directions. GMU with hidden states is proposed to fuse the information of multiple MRI modalities in both feature and decision level. Based on these, privilege information extracted by GMU fusion model is utilized to train a novel network called distilled-CNN, which significantly improves the performance of classification using single modality. Empirical studies on BRATS datasets corroborate the effectiveness of the proposed 3D CNN with GMU fusion and distilled-CNN to distinguish benign gliomas and malignant gliomas.
Fangyan Ye, Jian Pu, Jun Wang 0006, Hongyuan Zha
BIBM3
2017 Low-rank decomposition meets kernel learning: A generalized Nyström method
Liang Lan, Kai Zhang 0001, Hancheng Ge, Wei Cheng 0002, Jun Liu 0003, Andreas Rauber, Xiaoli Li 0001, Jun Wang 0006, Hongyuan Zha
Artif. Intell.8
2016 A Bayesian Hashing approach and its application to face recognition
Qi Dai 0001, Jun Wang 0006, Yurong Chen 0001, Yu-Gang Jiang 0001
Neurocomputing3
2016 Multiple task learning with flexible structure regularization
Jian Pu, Jun Wang 0006, Yu-Gang Jiang 0001, Xiangyang Xue 0001
Neurocomputing2
2016 Learning to Hash for Indexing Big Data - A Survey
abstract
The explosive growth in Big Data has attracted much attention in designing efficient indexing and search methods recently. In many critical applications such as large-scale search and pattern matching, finding the nearest neighbors to a query is a fundamental research problem. However, the straightforward solution using exhaustive comparison is infeasible due to the prohibitive computational complexity and memory requirement. In response, approximate nearest neighbor (ANN) search based on hashing techniques has become popular due to its promising performance in both efficiency and accuracy. Prior randomized hashing methods, e.g., locality-sensitive hashing (LSH), explore data-independent hash functions with random projections or permutations. Although having elegant theoretic guarantees on the search quality in certain metric spaces, performance of randomized hashing has been shown insufficient in many real-world applications. As a remedy, new approaches incorporating data-driven learning methods in development of advanced hash functions have emerged. Such learning-to-hash methods exploit information such as data distributions or class labels when optimizing the hash codes or functions. Importantly, the learned hash codes are able to preserve the proximity of neighboring data in the original feature spaces in the hash code spaces. The goal of this paper is to provide readers with systematic understanding of insights, pros, and cons of the emerging techniques. We provide a comprehensive survey of the learning-to-hash framework and representative techniques of various types, including unsupervised, semisupervised, and supervised. In addition, we also summarize recent hashing approaches utilizing the deep learning models. Finally, we discuss the future direction and trends of research in this area.
Jun Wang 0006, Wei Liu 0005, Sanjiv Kumar, Shih-Fu Chang
Proc. IEEE1
2015 Probabilistic Attributed Hashing
abstract
Due to the simplicity and efficiency, many hashing methods have recently been developed for large-scale similarity search. Most of the existing hashing methods focus on mapping low-level features to binary codes, but neglect attributes that are commonly associated with data samples. Attribute data, such as image tag, product brand, and user profile, can represent human recognition better than low-level features. However, attributes have specific characteristics, including high-dimensional, sparse and categorical properties, which is hardly leveraged into the existing hashing learning frameworks. In this paper, we propose a hashing learning framework, Probabilistic Attributed Hashing (PAH), to integrate attributes with low-level features. The connections between attributes and low-level features are built through sharing a common set of latent binary variables, i.e. hash codes, through which attributes and features can complement each other. Finally, we develop an efficient iterative learning algorithm, which is generally feasible for large-scale applications. Extensive experiments and comparison study are conducted on two public datasets, i.e., DBLP and NUS-WIDE. The results clearly demonstrate that the proposed PAH method substantially outperforms the peer methods.
Mingdong Ou, Peng Cui 0001, Jun Wang 0006, Fei Wang 0001, Wenwu Zhu 0001
AAAI3
2015 Multi-View Point Registration via Alternating Optimization
abstract
Multi-view point registration is a relatively less studied problem compared with two-view point registration. Directly applying pairwise registration often leads to matching discrepancy as the mapping between two point sets can be determined either by direct correspondences or by any intermediate point set. Also, the local two-view registration tends to be sensitive to noises. We propose a novel multi-view registration method, where the optimal registration is achieved via an efficient and effective alternating concave minimization process. We further extend our solution to a general case in practice of registration among point sets with different cardinalities. Extensive empirical evaluations of peer methods on both synthetic data and real images suggest our method is robust to large disturbance. In particular, it is shown that our method outperforms peer point matching methods and performs competitively against graph matching approaches. The latter approaches utilize the additional second-order information at the cost of exponentially increased run-time, thus usually being less efficient.
Junchi Yan, Jun Wang 0006, Hongyuan Zha, Xiaokang Yang 0001, Stephen M. Chu
AAAI2
2015 Optimal Bayesian Hashing for Efficient Face Recognition
Qi Dai 0001, Jun Wang 0006, Yurong Chen 0001, Yu-Gang Jiang 0001
IJCAI3
2015 Portfolio Choices with Orthogonal Bandit Learning
Weiwei Shen, Jun Wang 0006, Yu-Gang Jiang 0001, Hongyuan Zha
IJCAI2
2015 Non-transitive Hashing with Latent Similarity Components
abstract
Approximating the semantic similarity between entities in the learned Hamming space is the key for supervised hashing techniques. The semantic similarities between entities are often non-transitive since they could share different latent similarity components. For example, in social networks, we connect with people for various reasons, such as sharing common interests, working in the same company, being alumni and so on. Obviously, these social connections are non-transitive if people are connected due to different reasons. However, existing supervised hashing methods treat the pairwise similarity relationships in a simple and unified way and project data into a single Hamming space, while neglecting that the non-transitive property cannot be ade- quately captured by a single Hamming space. In this paper, we propose a non-transitive hashing method, namely Multi-Component Hashing (MuCH), to identify the latent similarity components to cope with the non-transitive similarity relationships. MuCH generates multiple hash tables with each hash table corresponding to a similarity component, and preserves the non-transitive similarities in different hash table respectively. Moreover, we propose a similarity measure, called Multi-Component Similarity, aggregating Hamming similarities in multiple hash tables to capture the non-transitive property of semantic similarity. We conduct extensive experiments on one synthetic dataset and two public real-world datasets (i.e. DBLP and NUS-WIDE). The results clearly demonstrate that the proposed MuCH method significantly outperforms the state-of-art hashing methods especially on search efficiency.
Mingdong Ou, Peng Cui 0001, Fei Wang 0001, Jun Wang 0006, Wenwu Zhu 0001
KDD4
2015 An Efficient Semi-Supervised Clustering Algorithm with Sequential Constraints
abstract
Semi-supervised clustering leverages side information such as pairwise constraints to guide clustering procedures. Despite promising progress, existing semi-supervised clustering approaches overlook the condition of side information being generated sequentially, which is a natural setting arising in numerous real-world applications such as social network and e-commerce system analysis. Given emerged new constraints, classical semi-supervised clustering algorithms need to re-optimize their objectives over all data samples and constraints in availability, which prevents them from efficiently updating the obtained data partitions. To address this challenge, we propose an efficient dynamic semi-supervised clustering framework that casts the clustering problem into a search problem over a feasible convex set, i.e., a convex hull with its extreme points being an ensemble of m data partitions. According to the principle of ensemble clustering, the optimal partition lies in the convex hull, and can thus be uniquely represented by an m-dimensional probability simplex vector. As such, the dynamic semi-supervised clustering problem is simplified to the problem of updating a probability simplex vector subject to the newly received pairwise constraints. We then develop a computationally efficient updating procedure to update the probability simplex vector in O(m2) time, irrespective of the data size n. Our empirical studies on several real-world benchmark datasets show that the proposed algorithm outperforms the state-of-the-art semi-supervised clustering algorithms with visible performance gain and significantly reduced running time.
Jinfeng Yi, Lijun Zhang 0005, Tianbao Yang, Wei Liu 0005, Jun Wang 0006
KDD5
2015 Consistency-Driven Alternating Optimization for Multigraph Matching: A Unified Approach
abstract
The problem of graph matching (GM) in general is nondeterministic polynomial-complete and many approximate pairwise matching techniques have been proposed. For a general setting in real applications, it typically requires to find the consistent matching across a batch of graphs. Sequentially performing pairwise matching is prone to error propagation along the pairwise matching sequence, and the sequences generated in different pairwise matching orders can lead to contradictory solutions. Motivated by devising a robust and consistent multiple-GM model, we propose a unified alternating optimization framework for multi-GM. In addition, we define and use two metrics related to graphwise and pairwise consistencies. The former is used to find an appropriate reference graph, which induces a set of basis variables and launches the iteration procedure. The latter defines the order in which the considered graphs in the iterations are manipulated. We show two embodiments under the proposed framework that can cope with the nonfactorized and factorized affinity matrix, respectively. Our multi-GM model has two major characters: 1) the affinity information across multiple graphs are explored in each iteration by fixing part of the matching variables via a consistency-driven mechanism and 2) the framework is flexible to incorporate various existing pairwise GM solvers in an out-of-box fashion, and also can proceed with the output of other multi-GM methods. The experimental results on both synthetic data and real images empirically show that the proposed framework performs competitively with the state-of-the-art.
Junchi Yan, Jun Wang 0006, Hongyuan Zha, Xiaokang Yang 0001, Stephen M. Chu
IEEE Trans. Image Process.2
2015 Discriminative Structured Feature Engineering for Macroscale Brain Connectomes
abstract
Neuroimaging techniques can measure structural and functional brain connectivity with unprecedented detail in vivo. This so-called brain connectome can be represented as high dimensional matrices corresponding to edge weights in graphs. After measuring the matrices of two cohorts (i.e., patients and healthy controls), one is often required to formulate computational network models for effective feature engineering to draw discriminative distinctions between the cohorts, as well as estimate the associated statistical significance. We designed a novel method to reveal the intrinsic features of functional matrices of discriminative power for group comparison. More specifically, by encouraging co-selection of edges connected to the same node, we preserved the discriminative edges to maximum extent. To reduce the false positive rate of the extracted discriminative edges, an optimization procedure was developed to evaluate the significance of these edges and remove trivial ones. We validated the proposed method using both synthetic data and real benchmarks, and compared it to ℓ1 regularized logistic regression, univariate t-test and stability selection. The experimental results clearly showed that the proposed approach outperformed the three competing methods under various settings. In addition to increasing the F-measure of feature selection, our approach captured the endogenous, discriminative connectivity patterns consistent with recent findings in biomedical literature. This data-driven method paves a new avenue of enquiry into the inherent nature of network models for functional brain connectomes.
Jian Pu, Jun Wang 0006, Wenwen Yu, Zhuangming Shen, Kristina Zeljic, Bomin Sun, Zheng Wang 0035
IEEE Trans. Medical Imaging2
2014 Privacy and Regression Model Preserved Learning
abstract
Sensitive data such as medical records and business reports usually contains valuable information that can be used to build prediction models. However, designing learning models by directly using sensitive data might result in severe privacy and copyright issues. In this paper, we propose a novel matrix completion based framework that aims to tackle two challenging issues simultaneously: i) handling missing and noisy sensitive data, and ii) preserving the privacy of the sensitive data during the learning process. In particular, the proposed framework is able to mask the sensitive data while ensuring that the transformed data are still usable for training regression models. We show that two key properties, namely model preserving and privacy preserving, are satisfied by the transformed data obtained from the proposed framework. In model preserving, we guarantee that the linear regression model built from the masked data approximates the regression model learned from the original data in a perfect way. In privacy preserving, we ensure that the original sensitive data cannot be recovered since the transformation procedure is irreversible. Given these two characteristics, the transformed data can be safely released to any learners for designing prediction models without revealing any private content. Our empirical studies with a synthesized dataset and multiple sensitive benchmark datasets verify our theoretical claim as well as the effectiveness of the proposed framework.
Jinfeng Yi, Jun Wang 0006, Rong Jin 0001
AAAI2
2014 Which Looks Like Which: Exploring Inter-class Relationships in Fine-Grained Visual Categorization
Jian Pu, Yu-Gang Jiang 0001, Jun Wang 0006, Xiangyang Xue 0001
ECCV (3)3
2014 From Low-Cost Depth Sensors to CAD: Cross-Domain 3D Shape Retrieval via Regression Tree Fields
Yan Wang 0059, Jun Wang 0006, Shih-Fu Chang
ECCV (1)4
2014 A Single-Pass Algorithm for Efficiently Recovering Sparse Cluster Centers of High-dimensional Data
abstract
Learning a statistical model for high-dimensional data is an important topic in machine learning. Although this problem has been well studied in the supervised setting, little is known about its unsupervised counterpart. In this work, we focus on the problem of clustering high-dimensional data with sparse centers. In particular, we address the following open question in unsupervised learning: “is it possible to reliably cluster high-dimensional data when the number of samples is smaller than the data dimensionality?" We develop an efficient clustering algorithm that is able to estimate sparse cluster centers with a single pass over the data. Our theoretical analysis shows that the proposed algorithm is able to accurately recover cluster centers with only O(s\log d) number of samples (data points), provided all the cluster centers are s-sparse vectors in a d dimensional space. Experimental results verify both the effectiveness and efficiency of the proposed clustering algorithm compared to the state-of-the-art algorithms on several benchmark datasets.
Jinfeng Yi, Lijun Zhang 0005, Jun Wang 0006, Rong Jin 0001, Anil K. Jain 0001
ICML3
2014 Predicting employee expertise for talent management in the enterprise
abstract
Strategic planning and talent management in large enterprises composed of knowledge workers requires complete, accurate, and up-to-date representation of the expertise of employees in a form that integrates with business processes. Like other similar organizations operating in dynamic environments, the IBM Corporation strives to maintain such current and correct information, specifically assessments of employees against job roles and skill sets from its expertise taxonomy. In this work, we deploy an analytics-driven solution that infers the expertise of employees through the mining of enterprise and social data that is not specifically generated and collected for expertise inference. We consider job role and specialty prediction and pose them as supervised classification problems. We evaluate a large number of feature sets, predictive models and postprocessing algorithms, and choose a combination for deployment. This expertise analytics system has been deployed for key employee population segments, yielding large reductions in manual effort and the ability to continually and consistently serve up-to-date and accurate data for several business functions. This expertise management system is in the process of being deployed throughout the corporation.
Kush R. Varshney, Vijil Chenthamarakshan, Scott W. Fancher, Jun Wang 0006, DongPing Fang 0001, Aleksandra Mojsilovic
KDD4
2014 Exploring Inter-feature and Inter-class Relationships with Deep Neural Networks for Video Classification
abstract
Videos contain very rich semantics and are intrinsically multimodal. In this paper, we study the challenging task of classifying videos according to their high-level semantics such as human actions or complex events. Although extensive efforts have been paid to study this problem, most existing works combined multiple features using simple fusion strategies and neglected the exploration of inter-class semantic relationships. In this paper, we propose a novel unified framework that jointly learns feature relationships and exploits the class relationships for improved video classification performance. Specifically, these two types of relationships are learned and utilized by rigorously imposing regularizations in a deep neural network (DNN). Such a regularized DNN can be efficiently launched using a GPU implementation with an affordable training cost. Through arming the DNN with better capability of exploring both the inter-feature and the inter-class relationships, the proposed regularized DNN is more suitable for identifying video semantics. With extensive experimental evaluations, we demonstrate that the proposed framework exhibits superior performance over several state-of-the-art approaches. On the well-known Hollywood2 and Columbia Consumer Video benchmarks, we obtain to-date the best reported results: 65.7% and 70.6% respectively in terms of mean average precision.
Zuxuan Wu, Yu-Gang Jiang 0001, Jun Wang 0006, Jian Pu, Xiangyang Xue 0001
ACM Multimedia3
2014 Self-Taught Spectral Clustering via Constraint Augmentation
abstract
Although constrained spectral clustering has been used extensively for the past few years, all work assumes the guidance (constraints) are given by humans. Original formulations of the problem assumed the constraints are given passively whilst later work allowed actively polling an Oracle (human experts). In this paper, for the first time to our knowledge, we explore the problem of augmenting the given constraint set for constrained spectral clustering algorithms. This moves spectral clustering towards the direction of self-teaching as has occurred in the supervised learning literature. We present a formulation for self-taught spectral clustering and show that the self-teaching process can drastically improve performance without further human guidance.
Xiang Wang 0001, Jun Wang 0006, Buyue Qian, Fei Wang 0001, Ian Davidson
SDM2
2013 Learning Hash Codes with Listwise Supervision
abstract
Hashing techniques have been intensively investigated in the design of highly efficient search engines for large-scale computer vision applications. Compared with prior approximate nearest neighbor search approaches like tree-based indexing, hashing-based search schemes have prominent advantages in terms of both storage and computational efficiencies. Moreover, the procedure of devising hash functions can be easily incorporated into sophisticated machine learning tools, leading to data-dependent and task-specific compact hash codes. Therefore, a number of learning paradigms, ranging from unsupervised to supervised, have been applied to compose appropriate hash functions. However, most of the existing hash function learning methods either treat hash function design as a classification problem or generate binary codes to satisfy pair wise supervision, and have not yet directly optimized the search accuracy. In this paper, we propose to leverage list wise supervision into a principled hash function learning framework. In particular, the ranking information is represented by a set of rank triplets that can be used to assess the quality of ranking. Simple linear projection-based hash functions are solved efficiently through maximizing the ranking quality over the training data. We carry out experiments on large image datasets with size up to one million and compare with the state-of-the-art hashing techniques. The extensive results corroborate that our learned hash codes via list wise supervision can provide superior search accuracy without incurring heavy computational overhead.
Jun Wang 0006, Wei Liu 0005, Andy X. Sun, Yu-Gang Jiang 0001
ICCV1
2013 Large-Scale Video Hashing via Structure Learning
abstract
Recently, learning based hashing methods have become popular for indexing large-scale media data. Hashing methods map high-dimensional features to compact binary codes that are efficient to match and robust in preserving original similarity. However, most of the existing hashing methods treat videos as a simple aggregation of independent frames and index each video through combining the indexes of frames. The structure information of videos, e.g., discriminative local visual commonality and temporal consistency, is often neglected in the design of hash functions. In this paper, we propose a supervised method that explores the structure learning techniques to design efficient hash functions. The proposed video hashing method formulates a minimization problem over a structure-regularized empirical loss. In particular, the structure regularization exploits the common local visual patterns occurring in video frames that are associated with the same semantic class, and simultaneously preserves the temporal consistency over successive frames from the same video. We show that the minimization objective can be efficiently solved by an Accelerated Proximal Gradient (APG) method. Extensive experiments on two large video benchmark datasets (up to around 150K video clips with over 12 million frames) show that the proposed method significantly outperforms the state-of-the-art hashing methods.
Guangnan Ye, Dong Liu 0001, Jun Wang 0006, Shih-Fu Chang
ICCV3
2013 Fast Pairwise Query Selection for Large-Scale Active Learning to Rank
abstract
Pair wise learning to rank algorithms (such as Rank SVM) teach a machine how to rank objects given a collection of ordered object pairs. However, their accuracy is highly dependent on the abundance of training data. To address this limitation and reduce annotation efforts, the framework of active pair wise learning to rank was introduced recently. However, in such a framework the number of possible query pairs increases quadratic ally with the number of instances. In this work, we present the first scalable pair wise query selection method using a layered (two-step) hashing framework. The first step relevance hashing aims to retrieve the strongly relevant or highly ranked points, and the second step uncertainty hashing is used to nominate pairs whose ranking is uncertain. The proposed framework aims to efficiently reduce the search space of pair wise queries and can be used with any pair wise learning to rank algorithm with a linear ranking function. We evaluate our approach on large-scale real problems and show it has comparable performance to exhaustive search. The experimental results demonstrate the effectiveness of our approach, and validate the efficiency of hashing in accelerating the search of massive pair wise queries.
Buyue Qian, Xiang Wang 0001, Jun Wang 0006, Nan Cao 0001, Weifeng Zhi, Ian Davidson
ICDM3
2013 Exploring Patient Risk Groups with Incomplete Knowledge
abstract
Patient risk stratification, which aims to stratify a patient cohort into a set of homogeneous groups according to some risk evaluation criteria, is an important task in modern medical informatics. Good risk stratification is the key to good personalized care plan design and delivery. The typical procedure for risk stratification is to first identify a set of risk-relevant medical features (also called risk factors), and then construct a predictive model to estimate the risk scores for individual patients. However, due to the heterogeneity of patients' clinical conditions, the risk factors and their importance vary across different patient groups. Therefore a better approach is to first segment the patient cohort into a set of homogeneous groups with consistent clinical conditions, namely risk groups, and then develop group-specific risk prediction models. In this paper, we propose RISGAL (RISk Group Analysis), a novel semi-supervised learning framework for patient risk group exploration. Our method segments a patient similarity graph into a set of risk groups such that some risk groups are in alignment with (incomplete) prior knowledge from the domain experts while the remaining groups reveal new knowledge from the data. Our method is validated on public benchmark datasets as well as a real electronic medical record database to identify risk groups from a set of potential Congestive Heart Failure (CHF) patients.
Xiang Wang 0001, Fei Wang 0001, Jun Wang 0006, Buyue Qian, Jianying Hu
ICDM3
2013 Multiple Task Learning Using Iteratively Reweighted Least Square
Jian Pu, Yu-Gang Jiang 0001, Jun Wang 0006, Xiangyang Xue 0001
IJCAI3
2013 Comparing apples to oranges: a scalable solution with heterogeneous hashing
abstract
Although hashing techniques have been popular for the large scale similarity search problem, most of the existing methods for designing optimal hash functions focus on homogeneous similarity assessment, i.e., the data entities to be indexed are of the same type. Realizing that heterogeneous entities and relationships are also ubiquitous in the real world applications, there is an emerging need to retrieve and search similar or relevant data entities from multiple heterogeneous domains, e.g., recommending relevant posts and images to a certain Facebook user. In this paper, we address the problem of ``comparing apples to oranges'' under the large scale setting. Specifically, we propose a novel Relation-aware Heterogeneous Hashing (RaHH), which provides a general framework for generating hash codes of data entities sitting in multiple heterogeneous domains. Unlike some existing hashing methods that map heterogeneous data in a common Hamming space, the RaHH approach constructs a Hamming space for each type of data entities, and learns optimal mappings between them simultaneously. This makes the learned hash codes flexibly cope with the characteristics of different data domains. Moreover, the RaHH framework encodes both homogeneous and heterogeneous relationships between the data entities to design hash functions with improved accuracy. To validate the proposed RaHH method, we conduct extensive evaluations on two large datasets; one is crawled from a popular social media sites, Tencent Weibo, and the other is an open dataset of Flickr(NUS-WIDE). The experimental results clearly demonstrate that the RaHH outperforms several state-of-the-art hashing methods with significant performance gains.
Mingdong Ou, Peng Cui 0001, Fei Wang 0001, Jun Wang 0006, Wenwu Zhu 0001, Shiqiang Yang
KDD4
2013 Active Learning to Rank using Pairwise Supervision
abstract
This paper investigates learning a ranking function using pairwise constraints in the context of human-machine interaction. As the performance of a learnt ranking model is predominantly determined by the quality and quantity of training data, in this work we explore an active learning to rank approach. Furthermore, since humans may not be able to confidently provide an order for a pair of similar instances we explore two types of pairwise supervision: (i) a set of “strongly” ordered pairs which contains confidently ranked instances, and (ii) a set of “weakly” ordered pairs which consists of similar or closely ranked instances. Our active knowledge injection is performed by querying domain experts on pairwise orderings, where informative pairs are located by considering both local and global uncertainties. Under this active scheme, querying of pairs which are uninformative or outliers instances would not occur. We evaluate the proposed approach on three real world datasets and compare with representative methods. The promising experimental results demonstrate the superior performance of our approach, and validate the effectiveness of actively using pairwise orderings to improve ranking performance.
Ian Davidson, Buyue Qian, Jun Wang 0006, Xiang Wang 0001
SDM4
2013 Single Network Relational Transductive Learning
abstract
Relational classification on a single connected network has been of particular interest in the machine learning and data mining communities in the last decade or so. This is mainly due to the explosion in popularity of social networking sites such as Facebook, LinkedIn and Google+ amongst others. In statistical relational learning, many techniques have been developed to address this problem, where we have a connected unweighted homogeneous/heterogeneous graph that is partially labeled and the goal is to propagate the labels to the unlabeled nodes. In this paper, we provide a different perspective by enabling the effective use of graph transduction techniques for this problem. We thus exploit the strengths of this class of methods for relational learning problems. We accomplish this by providing a simple procedure for constructing a weight matrix that serves as input to a rich class of graph transduction techniques. Our procedure has multiple desirable properties. For example, the weights it assigns to edges between unlabeled nodes naturally relate to a measure of association commonly used in statistics, namely the Gamma test statistic. We further portray the efficacy of our approach on synthetic as well as real data, by comparing it with state-of-the-art relational learning algorithms, and graph transduction techniques with an adjacency matrix or a real valued weight matrix computed using available attributes as input. In these experiments we see that our approach consistently outperforms other approaches when the graph is sparsely labeled, and remains competitive with the best when the proportion of known labels increases.
Amit Dhurandhar, Jun Wang 0006
J. Artif. Intell. Res.2
2013 Semi-supervised learning using greedy max-cut
Jun Wang 0006, Tony Jebara, Shih-Fu Chang
J. Mach. Learn. Res.1
2013 Query-Adaptive Image Search With Hash Codes
abstract
Scalable image search based on visual similarity has been an active topic of research in recent years. State-of-the-art solutions often use hashing methods to embed high-dimensional image features into Hamming space, where search can be performed in real-time based on Hamming distance of compact hash codes. Unlike traditional metrics (e.g., Euclidean) that offer continuous distances, the Hamming distances are discrete integer values. As a consequence, there are often a large number of images sharing equal Hamming distances to a query, which largely hurts search results where fine-grained ranking is very important. This paper introduces an approach that enables query-adaptive ranking of the returned images with equal Hamming distances to the queries. This is achieved by firstly offline learning bitwise weights of the hash codes for a diverse set of predefined semantic concept classes. We formulate the weight learning process as a quadratic programming problem that minimizes intra-class distance while preserving inter-class relationship captured by original raw image features. Query-adaptive weights are then computed online by evaluating the proximity between a query and the semantic concept classes. With the query-adaptive bitwise weights, returned images can be easily ordered by weighted Hamming distance at a finer-grained hash code level rather than the original Hamming distance level. Experiments on a Flickr image dataset show clear improvements from our proposed approach.
Yu-Gang Jiang 0001, Jun Wang 0006, Xiangyang Xue 0001, Shih-Fu Chang
IEEE Trans. Multim.2
2012 Supervised hashing with kernels
abstract
Recent years have witnessed the growing popularity of hashing in large-scale vision problems. It has been shown that the hashing quality could be boosted by leveraging supervised information into hash function learning. However, the existing supervised methods either lack adequate performance or often incur cumbersome model training. In this paper, we propose a novel kernel-based supervised hashing model which requires a limited amount of supervised information, i.e., similar and dissimilar data pairs, and a feasible training cost in achieving high quality hashing. The idea is to map the data to compact binary codes whose Hamming distances are minimized on similar pairs and simultaneously maximized on dissimilar pairs. Our approach is distinct from prior works by utilizing the equivalence between optimizing the code inner products and the Hamming distances. This enables us to sequentially and efficiently train the hash functions one bit at a time, yielding very short yet discriminative codes. We carry out extensive experiments on two image benchmarks with up to one million samples, demonstrating that our approach significantly outperforms the state-of-the-arts in searching both metric distance neighbors and semantically similar neighbors, with accuracy gains ranging from 13% to 46%.
Wei Liu 0005, Jun Wang 0006, Rongrong Ji, Yu-Gang Jiang 0001, Shih-Fu Chang
CVPR2
2012 Compact Hyperplane Hashing with Bilinear Functions
Wei Liu 0005, Jun Wang 0006, Yadong Mu, Sanjiv Kumar, Shih-Fu Chang
ICML2
2012 Legislative Prediction via Random Walks over a Heterogeneous Graph
abstract
In this article, we propose a random walk-based model to predict legislators' votes on a set of bills. In particular, we first convert roll call data, i.e. the recorded votes and the corresponding deliberative bodies, to a heterogeneous graph, where both the legislators and bills are treated as vertices. Three types of weighted edges are then computed accordingly, representing legislators' social and political relations, bills' semantic similarity, and legislator-bill vote relations. Through performing two-stage random walks over this heterogeneous graph, we can estimate legislative votes on past and future bills. We apply this proposed method on real legislative roll call data of the United States Congress and compare to state-of-the-art approaches. The experimental results demonstrate the superior performance and unique prediction power of the proposed model.
Jun Wang 0006, Kush R. Varshney, Aleksandra Mojsilovic
SDM1
2012 Semi-Supervised Hashing for Large-Scale Search
abstract
Hashing-based approximate nearest neighbor (ANN) search in huge databases has become popular due to its computational and memory efficiency. The popular hashing methods, e.g., Locality Sensitive Hashing and Spectral Hashing, construct hash functions based on random or principal projections. The resulting hashes are either not very accurate or are inefficient. Moreover, these methods are designed for a given metric similarity. On the contrary, semantic similarity is usually given in terms of pairwise labels of samples. There exist supervised hashing methods that can handle such semantic similarity, but they are prone to overfitting when labeled data are small or noisy. In this work, we propose a semi-supervised hashing (SSH) framework that minimizes empirical error over the labeled set and an information theoretic regularizer over both labeled and unlabeled sets. Based on this framework, we present three different semi-supervised hashing methods, including orthogonal hashing, nonorthogonal hashing, and sequential hashing. Particularly, the sequential hashing method generates robust codes in which each hash function is designed to correct the errors made by the previous ones. We further show that the sequential learning paradigm can be extended to unsupervised domains where no labeled pairs are available. Extensive experiments on four large datasets (up to 80 million samples) demonstrate the superior performance of the proposed SSH methods over state-of-the-art supervised and unsupervised hashing techniques.
Jun Wang 0006, Sanjiv Kumar, Shih-Fu Chang
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Robust and Scalable Graph-Based Semisupervised Learning
abstract
Graph-based semisupervised learning (GSSL) provides a promising paradigm for modeling the manifold structures that may exist in massive data sources in high-dimensional spaces. It has been shown effective in propagating a limited amount of initial labels to a large amount of unlabeled data, matching the needs of many emerging applications such as image annotation and information retrieval. In this paper, we provide reviews of several classical GSSL methods and a few promising methods in handling challenging issues often encountered in web-scale applications. First, to successfully incorporate the contaminated noisy labels associated with web data, label diagnosis and tuning techniques applied to GSSL are surveyed. Second, to support scalability to the gigantic scale (millions or billions of samples), recent solutions based on anchor graphs are reviewed. To help researchers pursue new ideas in this area, we also summarize a few popular data sets and software tools publicly available. Important open issues are discussed at the end to stimulate future research.
Wei Liu 0005, Jun Wang 0006, Shih-Fu Chang
Proc. IEEE2
2012 Fast Semantic Diffusion for Large-Scale Context-Based Image and Video Annotation
abstract
Exploring context information for visual recognition has recently received significant research attention. This paper proposes a novel and highly efficient approach, which is named semantic diffusion, to utilize semantic context for large-scale image and video annotation. Starting from the initial annotation of a large number of semantic concepts (categories), obtained by either machine learning or manual tagging, the proposed approach refines the results using a graph diffusion technique, which recovers the consistency and smoothness of the annotations over a semantic graph. Different from the existing graph-based learning methods that model relations among data samples, the semantic graph captures context by treating the concepts as nodes and the concept affinities as the weights of edges. In particular, our approach is capable of simultaneously improving annotation accuracy and adapting the concept affinities to new test data. The adaptation provides a means to handle domain change between training and test data, which often occurs in practice. Extensive experiments are conducted to improve concept annotation results using Flickr images and TV program videos. Results show consistent and significant performance gain (10 +% on both image and video data sets). Source codes of the proposed algorithms are available online.
Yu-Gang Jiang 0001, Qi Dai 0001, Jun Wang 0006, Chong-Wah Ngo, Xiangyang Xue 0001, Shih-Fu Chang
IEEE Trans. Image Process.3
2011 Hashing with Graphs
Wei Liu 0005, Jun Wang 0006, Sanjiv Kumar, Shih-Fu Chang
ICML2
2011 Lost in binarization: query-adaptive ranking for similar image search with compact codes
abstract
With the proliferation of images on the Web, fast search of visually similar images has attracted significant attention. State-of-the-art techniques often embed high-dimensional visual features into low-dimensional Hamming space, where search can be performed in real-time based on Hamming distance of compact binary codes. Unlike traditional metrics (e.g., Euclidean) of raw image features that produce continuous distance, the Hamming distances are discrete integer values. In practice, there are often a large number of images sharing equal Hamming distances to a query, resulting in a critical issue for image search where ranking is very important. In this paper, we propose a novel approach that facilitates query-adaptive ranking for the images with equal Hamming distance. We achieve this goal by firstly offline learning bit weights of the binary codes for a diverse set of predefined semantic concept classes. The weight learning process is formulated as a quadratic programming problem that minimizes intra-class distance while preserving interclass relationship in the original raw image feature space. Query-adaptive weights are then rapidly computed by evaluating the proximity between a query and the concept categories. With the adaptive bit weights, the returned images can be ordered by weighted Hamming distance at a finer-grained binary code level rather than at the original integer Hamming distance level. Experimental results on a Flickr image dataset show clear improvements from our query-adaptive ranking approach.
Yu-Gang Jiang 0001, Jun Wang 0006, Shih-Fu Chang
ICMR2
2010 Semi-supervised hashing for scalable image retrieval
abstract
Large scale image search has recently attracted considerable attention due to easy availability of huge amounts of data. Several hashing methods have been proposed to allow approximate but highly efficient search. Unsupervised hashing methods show good performance with metric distances but, in image search, semantic similarity is usually given in terms of labeled pairs of images. There exist supervised hashing methods that can handle such semantic similarity but they are prone to overfitting when labeled data is small or noisy. Moreover, these methods are usually very slow to train. In this work, we propose a semi-supervised hashing method that is formulated as minimizing empirical error on the labeled data while maximizing variance and independence of hash bits over the labeled and unlabeled data. The proposed method can handle both metric as well as semantic similarity. The experimental results on two large datasets (up to one million samples) demonstrate its superior performance over state-of-the-art supervised and unsupervised methods.
Jun Wang 0006, Ondrej Kumar, Shih-Fu Chang
CVPR1
2010 Sequential Projection Learning for Hashing with Compact Codes
Jun Wang 0006, Sanjiv Kumar, Shih-Fu Chang
ICML1
2010 In a Blink of an Eye and a Switch of a Transistor: Cortically Coupled Computer Vision
abstract
Our society's information technology advancements have resulted in the increasingly problematic issue of information overload—i.e., we have more access to information than we can possibly process. This is nowhere more apparent than in the volume of imagery and video that we can access on a daily basis—for the general public, availability of YouTube video and Google Images, or for the image analysis professional tasked with searching security video or satellite reconnaissance. Which images to look at and how to ensure we see the images that are of most interest to us, begs the question of whether there are smart ways to triage this volume of imagery. Over the past decade, computer vision research has focused on the issue of ranking and indexing imagery. However, computer vision is limited in its ability to identify interesting imagery, particularly as “interesting” might be defined by an individual. In this paper we describe our efforts in developing brain–computer interfaces (BCIs) which synergistically integrate computer vision and human vision so as to construct a system for image triage. Our approach exploits machine learning for real-time decoding of brain signals which are recorded noninvasively via electroencephalography (EEG). The signals we decode are specific for events related to imagery attracting a user's attention. We describe two architectures we have developed for this type of cortically coupled computer vision and discuss potential applications and challenges for the future.
Paul Sajda, Eric Pohlmeyer, Jun Wang 0006, Lucas C. Parra, Christoforos Christoforou, Jacek Dmochowski, Barbara Hanna, Claus Bahlmann, Maneesh Kumar Singh 0001, Shih-Fu Chang
Proc. IEEE3
2009 Label diagnosis through self tuning forweb image search
abstract
Semi-supervised learning (SSL) relies on partial supervision information for prediction, where only a small set of samples are associated with labels. Performance of SSL is significantly degraded if the given labels are not reliable. Such problems arise in realistic applications such as web image search using noisy textual tags. This paper proposes a novel and efficient graph based SSL method with the unique capacity of pruning contradictory labels and inferring new labels through a bidirectional and alternating optimization process. The objective is to automatically identify the most suitable samples for manipulation, labeling or unlabeling, and meanwhile estimate a smooth classification function over a weighted graph. Different from other graph based SSL approaches, the proposed method employs a bivariate objective function and iteratively modifies label variables on both labeled and unlabeled samples. Starting from such a SSL setting, we present a relearning framework to improve the performance of base learner, particularly for the application of web image search. Besides the toy demonstration on artificial data, we evaluated the proposed method on Flickr image search with unreliable textual labels. Experimental results confirm the significant improvements of the method over the baseline text based search engine and the state-of-the-art SSL methods.
Jun Wang 0006, Yu-Gang Jiang 0001, Shih-Fu Chang
CVPR1
2009 Domain adaptive semantic diffusion for large scale context-based video annotation
abstract
Learning to cope with domain change has been known as a challenging problem in many real-world applications. This paper proposes a novel and efficient approach, named domain adaptive semantic diffusion (DASD), to exploit semantic context while considering the domain-shift-of-context for large scale video concept annotation. Starting with a large set of concept detectors, the proposed DASD refines the initial annotation results using graph diffusion technique, which preserves the consistency and smoothness of the annotation over a semantic graph. Different from the existing graph learning methods which capture relations among data samples, the semantic graph treats concepts as nodes and the concept affinities as the weights of edges. Particularly, the DASD approach is capable of simultaneously improving the annotation results and adapting the concept affinities to new test data. The adaptation provides a means to handle domain change between training and test data, which occurs very often in video annotation task. We conduct extensive experiments to improve annotation results of 374 concepts over 340 hours of videos from TRECVID 2005-2007 data sets. Results show consistent and significant performance gain over various baselines. In addition, the proposed approach is very efficient, completing DASD over 374 concepts within just 2 milliseconds for each video shot on a regular PC.
Yu-Gang Jiang 0001, Jun Wang 0006, Shih-Fu Chang, Chong-Wah Ngo
ICCV2
2009 Graph construction and b-matching for semi-supervised learning
abstract
Graph based semi-supervised learning (SSL) methods play an increasingly important role in practical machine learning systems. A crucial step in graph based SSL methods is the conversion of data into a weighted graph. However, most of the SSL literature focuses on developing label inference algorithms without extensively studying the graph building method and its effect on performance. This article provides an empirical study of leading semi-supervised methods under a wide range of graph construction algorithms. These SSL inference algorithms include the Local and Global Consistency (LGC) method, the Gaussian Random Field (GRF) method, the Graph Transduction via Alternating Minimization (GTAM) method as well as other techniques. Several approaches for graph construction, sparsification and weighting are explored including the popular k-nearest neighbors method (kNN) and the b-matching method. As opposed to the greedily constructed kNN graph, the b-matched graph ensures each node in the graph has the same number of edges and produces a balanced or regular graph. Experimental results on both artificial data and real benchmark datasets indicate that b-matching produces more robust graphs and therefore provides significantly better prediction accuracy without any significant change in computation time.
Tony Jebara, Jun Wang 0006, Shih-Fu Chang
ICML2
2009 Brain state decoding for rapid image retrieval
abstract
Human visual perception is able to recognize a wide range of targets under challenging conditions, but has limited throughput. Machine vision and automatic content analytics can process images at a high speed, but suffers from inadequate recognition accuracy for general target classes. In this paper, we propose a new paradigm to explore and combine the strengths of both systems. A single trial EEG-based brain machine interface (BCI) subsystem is used to detect objects of interest of arbitrary classes from an initial subset of images. The EEG detection outcomes are used as input to a graph-based pattern mining subsystem to identify, refine, and propagate the labels to retrieve relevant images from a much larger pool. The combined strategy is unique in its generality, robustness, and high throughput. It has great potential for advancing the state of the art in media retrieval applications. We have evaluated and demonstrated significant performance gains of the proposed system with multiple and diverse image classes over several data sets, including those from Internet (Caltech 101) and remote sensing images. In this paper, we will also present insights learned from the experiments and discuss future research directions.
Jun Wang 0006, Eric Pohlmeyer, Barbara Hanna, Yu-Gang Jiang 0001, Paul Sajda, Shih-Fu Chang
ACM Multimedia1
2009 An image score inference system for RNAi genome-wide screening based on fuzzy mixture regression modeling
Jun Wang 0006, Xiaobo Zhou 0001, Fuhai Li 0001, Pamela Bradley, Shih-Fu Chang, Norbert Perrimon, Stephen T. C. Wong
J. Biomed. Informatics1
2008 Active microscopic cellular image annotation by superposable graph transduction with imbalanced labels
abstract
Systematic content screening of cell phenotypes in microscopic images has been shown promising in gene function understanding and drug design. However, manual annotation of cells and images in genome-wide studies is cost prohibitive. In this paper, we propose a highly efficient active annotation framework, in which a small amount of expert input is leveraged to rapidly and effectively infer the labels over the remaining unlabeled data. We formulate this as a graph based transductive learning problem and develop a novel method for label propagation. Specifically, a label regularizer method is proposed to handle the important label imbalance issue, typically seen in the cellular image screening applications. We also design a new scheme which breaks the graph into linear superposition of contributions from individual labeled samples. We take advantage of such a superposable representation to achieve fast annotation in an interactive setting. Extensive evaluations over toy data and realistic cellular images confirm the superiority of the proposed method over existing alternatives.
Jun Wang 0006, Shih-Fu Chang, Xiaobo Zhou 0001, Stephen T. C. Wong
CVPR1
2008 Graph transduction via alternating minimization
abstract
Graph transduction methods label input data by learning a classification function that is regularized to exhibit smoothness along a graph over labeled and unlabeled samples. In practice, these algorithms are sensitive to the initial set of labels provided by the user. For instance, classification accuracy drops if the training set contains weak labels, if imbalances exist across label classes or if the labeled portion of the data is not chosen at random. This paper introduces a propagation algorithm that more reliably minimizes a cost function over both a function on the graph and a binary label matrix. The cost function generalizes prior work in graph transduction and also introduces node normalization terms for resilience to label imbalances. We demonstrate that global minimization of the function is intractable but instead provide an alternating minimization scheme that incrementally adjusts the function and the labels towards a reliable local minimum. Unlike prior methods, the resulting propagation of labels does not prematurely commit to an erroneous labeling and obtains more consistent labels. Experiments are shown for synthetic and real classification tasks including digit and text recognition. A substantial improvement in accuracy compared to state of the art semi-supervised methods is achieved. The advantage are even more dramatic when labeled instances are limited.
Jun Wang 0006, Tony Jebara, Shih-Fu Chang
ICML1