Jian Chen 0011

dblp:49/6002-11 · DBLP profile ↗
← Back
93ranked-venue papers
14as first author
52since 2021 · last 2026
0000-0003-4769-1526ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 2 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 4 first-author · 19 since 2021Databases, data management, data science and information retrieval · 24 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training
abstract
Recent advancements in text-to-image (T2I) generation have led to the emergence of highly expressive models such as diffusion transformers (DiTs), exemplified by FLUX. However, their massive parameter sizes lead to slow inference, high memory usage, and poor deployability. Existing acceleration methods (e.g., single-step distillation and attention pruning) often suffer from significant performance degradation and incur substantial training costs. To address these limitations, we propose FastFLUX, an architecture-level pruning framework designed to enhance the inference efficiency of FLUX. At its core is the Block-wise Replacement with Linear Layers (BRLL) method, which replaces structurally complex residual branches in ResBlocks with lightweight linear layers while preserving the original shortcut connections for stability. Furthermore, we introduce Sandwich Training (ST), a localized fine-tuning strategy that leverages LoRA to supervise neighboring blocks, mitigating performance drops caused by structural replacement. Experiments show that our FastFLUX maintains high image quality under both qualitative and quantitative evaluations, while significantly improving inference speed, even with 20% of the hierarchy pruned.
Fuhan Cai, Jie Li 0022, Wenbo Li 0002, Jian Chen 0011, Xiangzhong Fang
AAAI5
2026 A Better Start: Sensitivity-Aware Warm-Up for Robust and Efficient Fine-Tuning
abstract
As an essential component of fine-tuning, warm-up plays a crucial role in promoting stability and generalization. Many studies have examined its underlying mechanisms from different aspects. However, most of the studies focus on incorporating these insights into optimizers to reduce the reliance on warm-up. Little attention has been paid to addressing the inherent limitations of the warm-up itself, which restricts its effectiveness. In this work, we revisit warm-up from a loss landscape perspective and identify several limitations with existing warm-up, including: (1) susceptibility to nearby suboptimal traps, (2) sensitivity to hyperparameters and random seeds, and (3) inefficiency during the early stages of training. To overcome these limitations, we propose Sensitivity-Aware Warm-Up (SAWU), a lightweight and adaptive strategy that dynamically leverages learning sensitivity during warm-up to guide updates toward better and more stable basins. In addition, SAWU also introduces an adaptive scheduling mechanism and phase transition strategy across warm-up, stable, and decay phases to further enhance robustness and efficiency. Extensive experiments on various downstream tasks show that SAWU significantly outperforms the vanilla method (e.g., average 3.43% improvement on RoBerta). Moreover, SAWU can be easily combined with various optimizers and remains effective even when warm-up-based methods fail (e.g, it lifts RAdam from 49.46% to 91.78% on qnli. Thanks to its lightweight nature, SAWU introduces minimal overhead and even reduces training time by over 5% compared to other methods.
Yile Chen 0004, Zeyi Wen, Jian Chen 0011, Jin Huang 0007
AAAI3
2026 NaVLA$^2$: A Vision-Language-Audio-Action Model for Multimodal Instruction Navigation
abstract
Embodied navigation is a fundamental capability for intelligent agents, yet remains challenging in partially observable environments where navigation instructions can be difficult to interpret. However, existing tasks only provide unimodal instructions, which are ambiguous in complex multimodal environments with multiple similar objects, and may result in misinterpretation and navigation failure. To overcome these limitations, we introduce MINav, a novel task where the navigation path is precisely described by a multimodal instruction. The instruction provides multimodal cues, including object categories, RGB images, language descriptions, and auditory descriptions, which help the agent to disambiguate and ground objects in the environment and navigate effectively. We further construct a large-scale dataset of 43.9K navigation episodes using a two-stage pipeline that first annotates multimodal references of objects and then synthesizes diverse multimodal instructions. We find that existing methods struggle on MINav task, indicating substantial room for improvement in agents' multimodal grounding. To address this, we propose NaVLA^2, a vision-language-audio-action model that additionally integrates spatial audio and employs a CoThinkAct module to jointly generate high-level reasoning and consistent low-level actions. Experimental results demonstrate that NaVLA^2 significantly outperforms competitive baselines on MINav benchmark. We hope that our proposed MINav and NaVLA^2 will facilitate future research toward agents with stronger multimodal understanding and grounding capabilities for navigation.
Jugang Fan, Peihao Chen, Jian Chen 0011, Mingkui Tan
AAAI5
2026 Emotion and Intention Guided Multi-Modal Learning for Sticker Response Selection
abstract
Stickers are widely used in online communication to convey emotions and implicit intentions. The Sticker Response Selection (SRS) task aims to select the most contextually appropriate sticker based on the dialogue. However, existing methods typically rely on semantic matching and model emotional and intentional cues separately, which can lead to mismatches when emotions and intentions are misaligned. To address this issue, we propose Emotion and Intention Guided Multi-Modal Learning (EIGML). This framework is the first to jointly model emotion and intention, effectively reducing the bias caused by isolated modeling and significantly improving selection accuracy. Specifically, we introduce Dual-Level Contrastive Framework to perform both intra-modality and inter-modality alignment, ensuring consistent representation of emotional and intentional features within and across modalities. In addition, we design an Intention-Emotion Guided Multi-Modal Fusion module that integrates emotional and intentional information progressively through three components: Emotion-Guided Intention Knowledge Selection, Intention-Emotion Guided Attention Fusion, and Similarity-Adjusted Matching Mechanism. This design injects rich, effective information into the model and enables a deeper understanding of the dialogue, ultimately enhancing sticker selection performance. Experimental results on two public datasets show that EIGML outperforms state-of-the-art baselines, achieving higher accuracy and a better understanding of emotional and intentional features.
Yuxuan Hu 0005, Jian Chen 0011, Yuhao Wang 0006, Zixuan Li 0001, Pengyue Jia, Wei Wang 0077, Chengming Li 0004, Xiangyu Zhao 0001
AAAI2
2026 Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video Models
abstract
Vision Transformers (ViTs) have achieved strong performance in video action recognition, but their high computational cost limits their practicality. Lightweight CNNs are more efficient but suffer from accuracy gaps. Cross-Architecture Knowledge Distillation (CAKD) addresses this by transferring knowledge from ViTs to CNNs, yet existing methods often struggle with architectural mismatch and overlook the value of stronger homogeneous CNN teachers. To tackle these challenges, we propose a Dual-Teacher Knowledge Distillation framework that leverages both a heterogeneous ViT teacher and a homogeneous CNN teacher to collaboratively guide a lightweight CNN student. We introduce two key components: (1) Discrepancy-Aware Teacher Weighting, which dynamically fuses the predictions from ViT and CNN teachers by assigning adaptive weights based on teacher confidence and prediction discrepancy with the student, enabling more informative and effective supervision; and (2) a Structure Discrepancy-Aware Distillation strategy, where the student learns the residual features between ViT and CNN teachers via a lightweight auxiliary branch, focusing on transferable architectural differences without mimicking all of ViT’s high-dimensional patterns. Extensive experiments on benchmarks including HMDB51, EPIC-KITCHENS-100, and Kinetics-400, demonstrate that our method consistently outperforms state-of-the-art distillation approaches, achieving notable performance improvements with a maximum accuracy gain of 5.95% on HMDB51.
Hongsen Ye, Changxin Huang, Xiping Hu, Jian Chen 0011, Runhao Zeng
AAAI5
2026 Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing
abstract
Wenhao Yuan, Chenchen Lin, Jian Chen, Jinfeng Xu, Xuehe Wang, Edith Cheuk-Han Ngai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wenhao Yuan 0005, Chenchen Lin, Jian Chen 0011, Jinfeng Xu 0003, Xuehe Wang, Edith C. H. Ngai
ACL (1)3
2026 Sustainable and Responsible ECG-Based AI Diagnostics: Masked Frequency Reconstruction with Peak-Aware Transformers
Wei Wang 0077, Jian Chen 0011, Junxin Chen 0001, Zeling Xu, Yuntao Zou, Henry H. Y. Tong
WWW2
2026 Leveraging pre-trained models for kernel machines
Zeyi Wen, Jian Chen 0011, Jin Huang 0007
Pattern Recognit.3
2026 EaSFE: Scalable and Efficient Feature Engineering for Boosting Machine Learning Performance
abstract
Feature engineering plays a critical role in machine learning (ML), but existing methods often struggle with high computational cost and limited scalability when applied to large-scale and sparse datasets. In this article, we propose EaSFE, an efficient and scalable feature engineering framework that unifies feature generation, filtering, and evaluation in an end-to-end manner. EaSFE is designed to efficiently construct and select informative features while explicitly considering computational and memory constraints. To achieve scalability, EaSFE incorporates parallel and distributed execution mechanisms, as well as a chunk-based data processing strategy that enables memory-efficient feature engineering on large datasets. In addition, EaSFE adopts tailored storage and execution strategies to handle high-dimensional sparse data effectively. Extensive experiments on multiple real-world datasets demonstrate that EaSFE consistently improves predictive performance (e.g., 5% accuracy improvement in poker ) while substantially enhancing efficiency (i.e., over 10x speedup) compared to existing feature engineering methods. In addition, EaSFE is demonstrated to scale to large and sparse datasets, successfully handling datasets with over 119 million training instances and 54 million features.
Jian Chen 0011, Yile Chen 0004, Zhenya Zheng, Zeyi Wen, Jin Huang 0007
ACM Trans. Knowl. Discov. Data1
2025 Learning to Generate Gradients for Test-Time Adaptation via Test-Time Training Layers
abstract
Test-time adaptation (TTA) aims to fine-tune a trained model online using unlabeled testing data to adapt to new environments or out-of-distribution data, demonstrating broad application potential in real-world scenarios. However, in this optimization process, unsupervised learning objectives like entropy minimization frequently encounter noisy learning signals. These signals produce unreliable gradients, which hinder the model’s ability to converge to an optimal solution quickly and introduce significant instability into the optimization process. In this paper, we seek to resolve these issues from the perspective of optimizer design. Unlike prior TTA using manually designed optimizers like SGD, we employ a learning-to-optimize approach to automatically learn an optimizer, called Meta Gradient Generator (MGG). Specifically, we aim for MGG to effectively utilize historical gradient information during the online optimization process to optimize the current model. To this end, in MGG, we design a lightweight and efficient sequence modeling layer -- gradient memory layer. It exploits a self-supervised reconstruction loss to compress historical gradient information into network parameters, thereby enabling better memorization ability over a long-term adaptation process. We only need a small number of unlabeled samples to pre-train MGG, and then the trained MGG can be deployed to process unseen samples. Promising results on ImageNet-C/R/Sketch/A indicate that our method surpasses current state-of-the-art methods with fewer updates, less data, and significantly shorter adaptation times. Compared with a previous SOTA SAR, we achieve 7.4% accuracy improvement and 4.2x faster adaptation speed on ImageNet-C.
Shuaicheng Niu, Ronghao Zhang, Yaofo Chen, Runhao Zeng, Jian Chen 0011, Xiping Hu
AAAI6
2025 Restabilizing Diffusion Models with Predictive Noise Fusion Strategy for Image Super-Resolution
abstract
Diffusion models are prominent in image generation for producing detailed and realistic images from Gaussian noises. However, they often encounter instability issues in image restoration tasks, e.g., super-resolution. Existing methods typically rely on multiple runs to find an initial noise that produces a reasonably restored image. Unfortunately, these methods are computationally expensive and time-consuming without guaranteeing stable and consistent performance. To address these challenges, we propose a novel Predictive Noise Fusion Strategy (PNFS) that predicts pixel-wise errors in the restored image and combines different noises to generate a more effective noise. Extensive experiments show that PNFS significantly improves the stability and performance of diffusion models in super-resolution, both quantitatively and qualitatively. Furthermore, PNFS can be flexibly integrated into various diffusion models to enhance their stability.
Luoqian Jiang, Bingna Xu, Haolin Pan, Jiezhang Cao, Wenbo Li 0002, Jian Chen 0011
AAAI7
2025 Gap Preserving Distillation by Building Bidirectional Mappings with A Dynamic Teacher
abstract
Knowledge distillation aims to transfer knowledge from a large teacher model to a compact student counterpart, often coming with a significant performance gap between them. Interestingly, we find that a too-large performance gap can hamper the training process. To alleviate this, we propose a **Gap Preserving Distillation (GPD)** method that trains an additional dynamic teacher model from scratch along with the student to maintain a reasonable performance gap. To further strengthen distillation, we develop a hard strategy by enforcing both models to share parameters. Besides, we also build the soft bidirectional mappings between them through ***Inverse Reparameterization (IR)*** and ***Channel-Branch Reparameterization (CBR)***. IR initializes a larger dynamic teacher with approximately the same accuracy as the student to avoid a too large gap in early stage of training. CBR enables direct extraction of an effective student model from the dynamic teacher without post-training. In experiments, GPD significantly outperforms existing distillation methods on top of both CNNs and transformers, achieving up to 1.58\% accuracy improvement. Interestingly, GPD also generalizes well to the scenarios without a pre-trained teacher, including training from scratch and fine-tuning, yielding a large improvement of 1.80\% and 0.89\% on ResNet18, respectively.
Shulian Zhang, Haolin Pan, Jing Liu 0048, Yulun Zhang 0001, Jian Chen 0011
ICLR6
2025 DERI: Cross-Modal ECG Representation Learning with Deep ECG-Report Interaction
abstract
Electrocardiogram (ECG) is widely used to diagnose cardiac conditions via deep learning methods. Although existing self-supervised learning (SSL) methods have achieved great performance in learning representation for ECG-based cardiac conditions classification, the clinical semantics can not be effectively captured. To overcome this limitation, we proposed to learn cross-modal ECG representations that contain more clinical semantics via a novel framework with \textbf{D}eep \textbf{E}CG-\textbf{R}eport \textbf{I}nteraction (\textbf{DERI}). Specifically, we design a novel framework combining multiple alignments and mutual feature reconstructions to learn effective representation of the ECG with the clinical report, which fuses the clinical semantics of the report. An RME-Module inspired by masked modeling is proposed to improve the ECG representation learning. Furthermore, we extend ECG representation learning to report generation with a language model, which is significant for evaluating clinical semantics in the learned representations and even clinical applications. Comprehensive experiments with various settings are conducted on various datasets to show the superior performance of our DERI. Our code is released on https://github.com/cccccj-03/DERI.
Jian Chen 0011, Xiaoru Dong, Wei Wang 0077, Shaorui Zhou, Lequan Yu, Xiping Hu
IJCAI1
2025 ECG2TOK: ECG Pre-Training with Self-Distillation Semantic Tokenizers
abstract
Self-supervised learning (SSL) has garnered increasing attention in electrocardiogram (ECG) analysis for its effectiveness in resource-limited settings. Existing state-of-the-art SSL methods rely on time-frequency detail reconstruction, but due to the inherent redundancy of ECG signals and individual variability, these approaches often yield suboptimal performance. In contrast, discrete label prediction becomes a superior pre-training objective by encouraging models to efficiently abstract ECG high-level semantics. However, the continuity and significant variability of ECG signals pose a challenge in generating semantically discrete labels. To address this issue, we propose an ECG pretraining framework with a self-distillation semantic tokenizer (ECG2TOK), which maps continuous ECG signals into discrete labels for self-supervised training. Specifically, the tokenizer extracts semantically aware embeddings of ECG by self-distillation and performs online clustering to generate semantically rich discrete labels. Subsequently, the SSL model is trained in conjunction with masking strategies and discrete label prediction to facilitate the abstraction of high-level semantic representations. We evaluate ECG2TOK in six downstream tasks, demonstrating that ECG2TOK efficiently achieves state-of-the-art performance and up to a 30.73% AUC increase in low-resource scenarios. Moreover, visualization experiments demonstrate that the discrete labels generated by ECG2TOK exhibit consistent semantics closely associated with clinical features. Our code is available on https://github.com/YXYanova/ECG2TOK.
Xiaoyan Yuan, Wei Wang 0077, Han Liu 0008, Jian Chen 0011, Xiping Hu
IJCAI4
2025 MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion Recognition
Jian Chen 0011, Yuxuan Hu 0005, Haifeng Lu, Wei Wang 0077, Min Yang 0007, Chengming Li 0004, Xiping Hu
ACM Multimedia1
2025 Federated Rényi Fair Inference in Federated Heterogeneous System
abstract
Federated Learning (FL) is a prominent distributed learning approach that addresses two major challenges: statistical heterogeneity (i.e., non-identically distributed data) and system heterogeneity (i.e., variability in the communication and computation on each client). As FL is commonly applied in sectors such as commercial and financial, group disparities may emerge and cause harm. However, current fairness algorithms assume homogeneous data, which does not align with the FL context. The main challenge is estimating global fairness measures (e.g., Rényi or Pearson correlation) in an asynchronous, heterogeneous system. To address this, we propose the FedRényi algorithm, which regularizes fairness by Rényi correlation. For statistical heterogeneity, FedRényi aggregates local fairness statistics to estimate the global Rényi correlation with an estimation error bound of $O(1/\sqrt{n})$, where $n$ is the total number of data samples. This theoretical result improves significantly over the prior result $O(1/\sqrt{K})$ with $K$ clients. We further prove that FedRényi converges at the same rate as in the homogeneous setting. For system heterogeneity, FedRényi approximates missing client updates through weighted averaging over a nearest neighbor region, ensuring a non-expansive approximation error under non-convex conditions. Extensive experiments demonstrate that FedRényi achieves a promising fairness-accuracy trade-off, with at least 2\% improvement over baselines.
Zhiyong Ma, Yuanjie Shi, Yan Yan 0006, Jian Chen 0011
UAI4
2025 Vehicle Dynamics and Interaction for Trajectory Prediction and Traffic Control
abstract
Trajectory prediction is a crucial challenge in autonomous vehicle motion planning and decision-making techniques. However, existing methods face limitations in accurately capturing vehicle dynamics and interactions. To address this issue, this article proposes a novel approach to extracting vehicle velocity and acceleration, enabling the learning of vehicle dynamics and encoding them as auxiliary information. The VDI-LSTM model is designed, incorporating graph convolution and attention mechanisms to capture vehicle interactions using trajectory data and dynamic information. Specifically, a dynamics encoder is designed to capture the dynamic information, a dynamic graph is employed to represent vehicle interactions, and an attention mechanism is introduced to enhance the performance of LSTM and graph convolution. To demonstrate the effectiveness of our model, extensive experiments are conducted, including comparisons with several baselines and ablation studies on real-world highway datasets. Experimental results show that VDI-LSTM outperforms other baselines compared, which obtains a 3% improvement on the average RMSE indicator over the five prediction steps.
Jian Chen 0011, Shaorui Zhou, Wei Wang 0077, Yuzhu Hu, Jianqing Li 0001, Ben-Guo He, Junxin Chen 0001, Marwan Omar, Ali Kashif Bashir, Xiping Hu
ACM Trans. Auton. Adapt. Syst.1
2025 iCTS: Iterative and Hierarchical Clock Tree Synthesis With Skew-Latency-Load Tree
abstract
The advancement of modern clock tree synthesis (CTS) encounters a bottleneck, primarily due to the difficulty in achieving multiobjective co-optimization among complex design processes. To concurrently optimize skew, latency, and load capacitance, we propose an iterative and hierarchical CTS framework, which is composed of clustering, topology generation and routing, buffering, and optimization. First, we introduce a capacitance-based metric to achieve adaptive balanced clustering and optimize the cluster results through simulated annealing. Second, to construct a clock tree with lower latency, load capacitance, and skew, we introduce the skew-latency-load tree (SLLT), which combines the advantages of bound skew tree and Steiner shallow-light tree, and we propose an effective SLLT construction algorithm. Third, to further optimize CTS result by buffering, we introduce the critical wirelength evaluation (CWE) to evaluate the capability of each buffer, and propose the insertion delay estimation (IDE) to reduce the evaluation bias during buffering, then design the iterative skew convergence algorithm (ISCA) to achieve complete convergence of skew. We validate our solution using 28 nm process technology. Compared to our method, the commercial tool increases skew, latency, and clock capacitance by 39.5%, 13.0%, and 18.5%, respectively, while the OpenROAD by 101.6%, 50.7%, and 25.5%, respectively.
Zhipeng Huang 0009, Bei Yu 0001, Wenxing Zhu, Jian Chen 0011, Zhixue He
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation
abstract
Test-time adaptation (TTA) aims to boost the generalization capability of a trained model by conducting self-/unsupervised learning during the testing phase. While most existing TTA methods for video primarily utilize visual supervisory signals, they often overlook the potential contribution of inherent audio data. To address this gap, we propose a novel approach that incorporates audio information into video TTA. Our method capitalizes on the rich semantic content of audio to generate audio-assisted pseudo-labels, a new concept in the context of video TTA. Specifically, we propose an audio-to-video label mapping method by first employing pre-trained audio models to classify audio signals extracted from videos and then mapping the audio-based predictions to video label spaces through large language models, thereby establishing a connection between the audio categories and video labels. To effectively leverage the generated pseudo-labels, we present a flexible adaptation cycle that determines the optimal number of adaptation iterations for each sample, based on changes in loss and consistency across different views. This enables a customized adaptation process for each sample. Experimental results on two widely used datasets (UCF101-C and Kinetics-Sounds-C), as well as on two newly constructed audio-video TTA datasets (AVE-C and AVMIT-C) with various corruption types, demonstrate the superiority of our approach. Our method consistently improves adaptation performance across different video classification models and represents a significant step forward in integrating audio information into video TTA. The code and datasets will be made publicly available.
Runhao Zeng, Ronghao Zhang, Shuaicheng Niu, Jian Chen 0011, Xiping Hu, Victor C. M. Leung
IEEE Trans. Circuits Syst. Video Technol.5
2025 Towards Recommendation on Good Quality Data Science Solutions
abstract
Data science aims to solve real-world problems with the knowledge derived from data. Successfully tackling a data science problem requires practitioners to choose an appropriate solution, which potentially comprises various components such as pre-processing techniques, learning algorithms, hyper-parameters, and so on. Therefore, a problem-driven recommendation for the promising solution is invaluable, as it facilitates efficient and convenient problem-solving. However, existing solution recommendation approaches confront notable challenges when dealing with limited and sparse prior experience in practical applications. Learning from such prior easily leads to overfitting and poor generalization in solution recommendations. To address this issue, we propose a novel solution recommendation method that can predict a good-quality data science solution, including the pre-processing, the learning algorithm, and hyper-parameters, for a given problem. The foundation of our method is a carefully designed ranking model that exploits a weight-sharing structure and a newly proposed loss. The ranking model focuses on incorporating relative ranking information into the predicted performance score of each solution. With these techniques, our method can recommend the solution with the highest score and effectively mitigate the limitations of using sparse prior experience. Our experiments demonstrate the superiority of our method in predicting solutions with higher accuracy and rank, even trained on highly sparse historical performance records. It also reduces recommendation time significantly compared to the baselines, offering remarkable efficiency and convenience for practitioners.
Jian Chen 0011, Yile Chen 0004, Zeyi Wen, Jin Huang 0007
ACM Trans. Knowl. Discov. Data1
2025 CMFAN: Cross-Modal Feature Alignment Network for Few-Shot Single-View 3D Reconstruction
abstract
Few-shot single-view 3D reconstruction learns to reconstruct the novel category objects based on a query image and a few support shapes. However, since the query image and the support shapes are of different modalities, there is an inherent feature misalignment problem damaging the reconstruction. Previous works in the literature do not consider this problem. To this end, we propose the cross-modal feature alignment network (CMFAN) with two novel techniques. One is a strategy for model pretraining, namely, cross-modal contrastive learning (CMCL), here the 2D images and 3D shapes of the same objects compose the positives, and those from different objects form the negatives. With CMCL, the model learns to embed the 2D and 3D modalities of the same object into a tight area in the feature space and push away those from different objects, thus effectively aligning the global cross-modal features. The other is cross-modal feature fusion (CMFF), which further aligns and fuses the local features. Specifically, it first re-represents the local features with the cross-attention operation, making the local features share more information. Then, CMFF generates a descriptor for the support features and attaches it to each local feature vector of the query image with dense concatenation. Moreover, CMFF can be applied to multilevel local features and brings further advantages. We conduct extensive experiments to evaluate the effectiveness of our designs, and CMFAN sets new state-of-the-art performance in all of the 1-/10-/25-shot tasks of ShapeNet and ModelNet datasets.
Lvlong Lai, Jian Chen 0011, Zehong Zhang, Guosheng Lin, Qingyao Wu
IEEE Trans. Neural Networks Learn. Syst.2
2024 Enhancing the Performance of Bandit-based Hyperparameter Optimization
abstract
Bandit-based methods are commonly used for hyperparameter optimization (HPO), which is significant in data analytics. When confronted with numerous configurations and high-dimensional large problems, existing bandit-based methods face challenges of high evaluation cost and poor optimization performance. To address these challenges, we introduce an improved bandit-based approach that exhibits enhanced evaluation ability and is suitable for situations with limited resources. Specifically, our method first effectively utilizes the feature and label information to conduct representative groups for further evaluation. After that, two kinds of folds (i.e., general folds and special folds) are constructed to facilitate better evaluation of the configuration in the cross-validation process. Additionally, we incorporate variance and subset size into the evaluation metric to comprehensively evaluate the configuration. We integrate our proposed method into three commonly used bandit-based methods, and experimental results on multiple datasets show that our method has advantages in stability with accuracy improvement of 1 % to 15 % on the datasets tested. In addition, since our method can avoid configurations that are low-quality but time-consuming to evaluate, it is always more efficient than the existing bandit-based methods, and can even reduce the execution time by half in some datasets. Sometimes it takes a little more time, but the improvement in accuracy can be significant.
Yile Chen 0004, Zeyi Wen, Jian Chen 0011, Jin Huang 0003
ICDE3
2024 TGCA-PVT: Topic-Guided Context-Aware Pyramid Vision Transformer for Sticker Emotion Recognition
abstract
Online chatting has become an essential aspect of our daily interactions, with stickers emerging as a prevalent tool for conveying emotions more vividly than plain text. While conventional image emotion recognition focuses on global features, sticker emotion recognition necessitates incorporating both global and local features, along with additional modalities like text. To address this, we introduce a topic ID-guided transformer method to facilitate a more nuanced analysis of the stickers. Since each sticker will have a topic, and the same topic will have the same object, we introduce a topic ID as a flag to group images by theme. Our approach encompasses a novel topic-guided context-aware module and a topic-guided attention mechanism, enabling the extraction of comprehensive topic context features from stickers sharing the same topic ID, significantly enhancing emotion recognition accuracy. Moreover, we integrate a frequency linear attention module to leverage frequency domain information to capture better the object information of the stickers and a locally enhanced re-attention mechanism for improved local feature extraction. Extensive experiments and ablation studies on the large-scale sticker emotion dataset SER30k validate the efficacy of our method. Experimental results show that our proposed method obtains the best accuracy on both single-modal and multi-modal sticker emotion recognition.
Jian Chen 0011, Wei Wang 0077, Yuzhu Hu, Junxin Chen 0001, Han Liu 0008, Xiping Hu
ACM Multimedia1
2024 Graph-Enhanced Low-Resource ECG Representation Learning for Emotion Recognition Based on Wearable Internet of Things
abstract
Internet of Things (IoT) devices like wearable devices have enabled quick monitoring of electrocardiogram (ECG) signals with lower resources than multielectrode ECG devices, opening up development opportunities for sustainable ECG-based emotion recognition. However, existing methods that rely on predesigned features extracted from single-lead ECG signals cannot automatically extract effective features from the original ECG signal collected by IoT devices. To address this limitation, we propose a novel approach leveraging signal transformation and graph representation learning for ECG-based emotion recognition. The signal graph learning process can be divided into local subgraph learning for ECG representation learning and signal enhancement graph to derive the graph-enhanced representation. We employ a designed loss function by calculating cosine similarity to extract an effective representation of the original signal from the transformed signal in the local subgraph learning. Additionally, we utilize a graph convolution model based on the signal enhancement graph to obtain a graph-enhanced representation of the ECG signal. The method incorporates six signal transformations and constructs a self-signal transformation graph. For emotion recognition, we design a classification network comprising convolutional neural networks (CNNs) and long short-term memory (LSTM) networks. Experiments on public data sets show the superiority of our method among other baselines. Ablation studies are conducted to verify the performance.
Jian Chen 0011, Yuzhu Hu, Lalit Garg, G. Thippa Reddy, Gautam Srivastava 0001, Wei Wang 0077
IEEE Internet Things J.1
2024 Boosting semi-supervised learning with Contrastive Complementary Labeling
Qinyi Deng, Zhibang Yang, Haolin Pan, Jian Chen 0011
Neural Networks5
2024 Towards Lightweight Super-Resolution With Dual Regression Learning
abstract
Deep neural networks have exhibited remarkable performance in image super-resolution (SR) tasks by learning a mapping from low-resolution (LR) images to high-resolution (HR) images. However, the SR problem is typically an ill-posed problem and existing methods would come with several limitations. First, the possible mapping space of SR can be extremely large since there may exist many different HR images that can be super-resolved from the same LR image. As a result, it is hard to directly learn a promising SR mapping from such a large space. Second, it is often inevitable to develop very large models with extremely high computational cost to yield promising SR performance. In practice, one can use model compression techniques to obtain compact models by reducing model redundancy. Nevertheless, it is hard for existing model compression methods to accurately identify the redundant components due to the extremely large SR mapping space. To alleviate the first challenge, we propose a dual regression learning scheme to reduce the space of possible SR mappings. Specifically, in addition to the mapping from LR to HR images, we learn an additional dual regression mapping to estimate the downsampling kernel and reconstruct LR images. In this way, the dual mapping acts as a constraint to reduce the space of possible mappings. To address the second challenge, we propose a dual regression compression (DRC) method to reduce model redundancy in both layer-level and channel-level based on channel pruning. Specifically, we first develop a channel number search method that minimizes the dual regression loss to determine the redundancy of each layer. Given the searched channel numbers, we further exploit the dual regression manner to evaluate the importance of channels and prune the redundant ones. Extensive experiments show the effectiveness of our method in obtaining accurate and efficient SR models.
Mingkui Tan, Zeshuai Deng, Jingdong Wang 0001, Qi Chen 0014, Jiezhang Cao, Yanwu Xu 0001, Jian Chen 0011
IEEE Trans. Pattern Anal. Mach. Intell.8
2024 Enhanced Long-Tailed Recognition With Contrastive CutMix Augmentation
abstract
Real-world data often follows a long-tailed distribution, where a few head classes occupy most of the data and a large number of tail classes only contain very limited samples. In practice, deep models often show poor generalization performance on tail classes due to the imbalanced distribution. To tackle this, data augmentation has become an effective way by synthesizing new samples for tail classes. Among them, one popular way is to use CutMix that explicitly mixups the images of tail classes and the others, while constructing the labels according to the ratio of areas cropped from two images. However, the area-based labels entirely ignore the inherent semantic information of the augmented samples, often leading to misleading training signals. To address this issue, we propose a Contrastive CutMix (ConCutMix) that constructs augmented samples with semantically consistent labels to boost the performance of long-tailed recognition. Specifically, we compute the similarities between samples in the semantic space learned by contrastive learning, and use them to rectify the area-based labels. Experiments show that our ConCutMix significantly improves the accuracy on tail classes as well as the overall performance. For example, based on ResNeXt-50, we improve the overall accuracy on ImageNet-LT by 3.0% thanks to the significant improvement of 3.3% on tail classes. We highlight that the improvement also generalizes well to other benchmarks and models. Our code and pretrained models are available at https://github.com/PanHaulin/ConCutMix.
Haolin Pan, Mianjie Yu, Jian Chen 0011
IEEE Trans. Image Process.4
2024 An Ensemble Classification Model for Depression Based on Wearable Device Sleep Data
abstract
Depression is one of the most common mental disorders, with sleep disturbances as typical symptoms. With the popularity of wearable devices increasing in recent years, more and more people wear portable devices to track sleep quality. Based on this, we believe that depression detection through wearable sleep data is more intelligent and economical. However, the majority of wearable devices face the problem of missing data during the data collection process. Otherwise, most existing studies of depression identification focus on the utilization of complex data, making it difficult to generalize and susceptible to noise interference. To address these issues, we propose a systematic ensemble classification model for depression (ECD). For the missing data problem of wearable devices, we design an improved GAIN method to further control the generation range of interpolated values, which can achieve a more reasonable treatment of missing values. Compared with the original GAIN approach, the improved method shows a 28.56% improvement when using MAE as the metric. For depression recognition, we use ensemble learning to construct a depression classification model which combines five classification models, including SVM, KNN, LR, CBR, and DT. Ensemble learning can improve the model's robustness and generalization. The voting mechanism is used in several places to improve noise immunity. The final classification model performed great on the dataset, with a precision of 92.55% and a recall of 91.89%. These results illustrate how efficient this method is in automatically detecting depression.
Yuzhu Hu, Jian Chen 0011, Junxin Chen 0001, Wei Wang 0077, Xiping Hu
IEEE J. Biomed. Health Informatics2
2024 CMNet: Component-Aware Matching Network for Few-Shot Point Cloud Classification
abstract
Few-shot point cloud classification is currently an under-explored problem which aims to learn a point cloud classifier for novel categories given a few annotated training data. Most existing methods achieve classification by matching a query point cloud to the most similar support category at the global representation level. However, due to the complicated structure of the point cloud and scarce available training data, the global representations of the point clouds and categories are of low quality, limiting the matching accuracy. Therefore, in this paper, we propose the Component-Aware Matching Network (CMNet) that matches the point clouds at the component level in addition to the global level. Specifically, we construct a component set for each query point cloud and support category and develop a metric to measure the similarity between two component sets. The final prediction is the weighted sum of the global and component matching probabilities. Besides, we carefully devise a component matching pretraining scheme for CMNet to enhance its ability to extract component features, further improving its performance. To evaluate the effectiveness of our design, we conduct comprehensive experiments on three benchmarks, namely ModelNet-FS, ShapeNet-FS and Shrec-FS. As a result, CMNet consistently outperforms the existing methods with significant margins in all the experiments of the three benchmarks and sets new state-of-the-art performance.
Lvlong Lai, Jian Chen 0011, Guosheng Lin, Qingyao Wu
IEEE Trans. Multim.2
2024 Zero-Shot Single-View Point Cloud Reconstruction via Cross-Category Knowledge Transferring
abstract
Single-view point cloud reconstruction aims to generate a 3D point cloud of an object given one 2D image taken from an arbitrary viewpoint. Most previous works assume that all the test categories have been present to the model during training. However, it is impossible to know all the test categories that the model will meet in advance. And we discover these methods can not deal with novel categories well. Therefore, in this article, we investigate a more realistic and challenging setting of single-view point cloud reconstruction, zero-shot, where the model's performance on novel categories is pursued. Towards this task, we propose the Cross-Category Knowledge Transferring Network (CCKTN), which maintains a knowledge bank to mine transferable knowledge from known categories to help reconstruct novel categories. Additionally, we conduct auxiliary learning for the point cloud reconstruction model with the point cloud autoencoder via sharing the same knowledge bank. This design enables the knowledge bank to collect more fruitful 3D knowledge of point clouds. Moreover, we devise a diversity loss regularization for the knowledge vectors to guarantee their diversities, further enhancing CCKTN's performance. Comprehensive experiments conducted on ShapeNet and ModelNet datasets show CCKTN's superiority towards existing methods and demonstrate CCKTN's effectiveness for reconstructing novel category objects.
Lvlong Lai, Jian Chen 0011, Qingyao Wu
IEEE Trans. Multim.2
2023 Dynamic Vehicle Graph Interaction for Trajectory Prediction Based on Video Signals
abstract
The roadside video surveillance signal can help people achieve vehicle tracking and trajectory generation. Using these trajectories can learn the future motion of vehicles. Existing prediction methods can not analyze the interaction between vehicles well. To this end, we design a dynamic vehicle graph to represent the dynamic interaction between vehicles for trajectory prediction. A graph convolution module with an attention mechanism is used to extract the feature of the dynamic vehicle graph. Experiments comparing with baseline methods are conducted on a real-world dataset to demonstrate our model’s improvement in forecast accuracy.
Jian Chen 0011, Wei Wang 0077, Junxin Chen 0001
ICASSP1
2023 Downscaled Representation Matters: Improving Image Rescaling with Collaborative Downscaled Images
abstract
Deep networks have achieved great success in image rescaling (IR) task that seeks to learn the optimal downscaled representations, i.e., low-resolution (LR) images, to reconstruct the original high-resolution (HR) images. Compared with super-resolution methods that consider a fixed downscaling scheme, e.g., bicubic, IR often achieves significantly better reconstruction performance thanks to the learned downscaled representations. This highlights the importance of a good downscaled representation. Existing IR methods mainly learn the downscaled representation by jointly optimizing the downscaling and upscaling models. Unlike them, we seek to improve the downscaled representation through a different and more direct way – directly optimizing the downscaled image itself instead of the down-/upscaling models. Consequently, we propose a Hierarchical Collaborative Downscaling (HCD) method that performs gradient descent w.r.t. the reconstruction loss in both HR and LR domains to improve the downscaled representations, so as to boost IR performance. Extensive experiments show that our HCD significantly improves the reconstruction performance both quantitatively and qualitatively. Particularly, we improve over popular IR methods by >0.57 dB PSNR on Set5. Moreover, we also highlight the flexibility of our HCD since it can generalize well across diverse image rescaling models. The code is available at https://github.com/xubingna/HCD.
Bingna Xu, Luoqian Jiang, Mianjie Yu, Jian Chen 0011
ICCV5
2023 A Tensor-based Markov Chain Model for Heterogeneous Information Network Collective Classification : Extended abstract
abstract
Heterogeneous Information Network(HIN) collective classification aims to classify one type of node, which is associated with multiple types of nodes through multiple types of relations. Previous studies have revealed that exploiting the relative importance of relation types is quite useful for improving node classification performance. We propose a Tensor-based Markov chain (T-Mark) model to improve the nodes classification accuracy by predicting the labels for unlabeled nodes and the importance ranking of relationship types automatically and simultaneously. Specifically, we build two tensor equations according to the HIN structure and content similarities among nodes of both labeled and unlabeled data. Consequently, We solve the semi-supervised T-Mark model by using an iterative process until obtaining two stationary distributions for labels and relation types. Experimental results on several real-world datasets demonstrate the effectiveness of T-Mark.
Chao Han 0002, Jian Chen 0011, Mingkui Tan, Michael Kwok-Po Ng, Qingyao Wu
ICDE2
2023 Exploring Motion Cues for Video Test-Time Adaptation
abstract
Test-time adaptation (TTA) aims at boosting the generalization capability of a trained model by conducting self-/un-supervised learning during testing in real-world applications. Though TTA on image-based tasks has seen significant progress, TTA techniques for video remain scarce. Naively introducing image-based TTA methods into video tasks may achieve limited performance, since these methods do not consider the special nature of video tasks, e.g., the motion information. In this paper, we propose leveraging motion cues in videos to design a new test-time learning scheme for video classification. We extract spatial appearance and dynamic motion clip features using two sampling rates (i.e., slow and fast) and propose a fast-to-slow unidirectional alignment scheme to align fast motion and slow appearance features, thereby enhancing the motion encoding ability. Additionally, we propose a slow-fast dual contrastive learning strategy to learn a joint feature space for fastly and slowly sampled clips, guiding the model to extract discriminative video features. Lastly, we introduce a stochastic pseudo-negative sampling scheme to provide better adaptation supervision by selecting a more reliable pseudo-negative label compared to the pseudo-positive label used in prior TTA methods. This technique reduces the adaptation difficulty often caused by poor performance on out-of-distribution test data before adaptation. Our approach significantly improves performance on various video classification backbones, as demonstrated through extensive experiments on two benchmark datasets.
Runhao Zeng, Huixuan Xu 0003, Shuaicheng Niu, Jian Chen 0011
ACM Multimedia5
2023 FGPrompt: Fine-grained Goal Prompting for Image-goal Navigation
abstract
Learning to navigate to an image-specified goal is an important but challenging task for autonomous systems like household robots. The agent is required to well understand and reason the location of the navigation goal from a picture shot in the goal position. Existing methods try to solve this problem by learning a navigation policy, which captures semantic features of the goal image and observation image independently and lastly fuses them for predicting a sequence of navigation actions. However, these methods suffer from two major limitations. 1) They may miss detailed information in the goal image, and thus fail to reason the goal location. 2) More critically, it is hard to focus on the goal-relevant regions in the observation image, because they attempt to understand observation without goal conditioning. In this paper, we aim to overcome these limitations by designing a Fine-grained Goal Prompting (\sexyname) method for image-goal navigation. In particular, we leverage fine-grained and high-resolution feature maps in the goal image as prompts to perform conditioned embedding, which preserves detailed information in the goal image and guides the observation encoder to pay attention to goal-relevant regions. Compared with existing methods on the image-goal navigation benchmark, our method brings significant performance improvement on 3 benchmark datasets (\textit{i.e.,} Gibson, MP3D, and HM3D). Especially on Gibson, we surpass the state-of-the-art success rate by 8\% with only 1/50 model size.
Peihao Chen, Jugang Fan, Jian Chen 0011, Thomas H. Li, Mingkui Tan
NeurIPS4
2023 Discrete limited attentional collaborative filtering for fast social recommendation
Zhibin Hu, Xuebin Zhou, Zhiwei He 0002, Zehang Yang, Jian Chen 0011, Jin Huang 0007
Eng. Appl. Artif. Intell.5
2023 Decentralized Gradient-Quantization Based Matrix Factorization for Fast Privacy-Preserving Point-of-Interest Recommendation
abstract
With the rapidly growing of location-based social networks, point-of-interest (POI) recommendation has been attracting tremendous attentions. Previous works for POI recommendation usually use matrix factorization (MF)-based methods, which achieve promising performance. However, existing MF-based methods suffer from two critical limitations: (1) Privacy issues: all users’ sensitive data are collected to the centralized server which may leak on either the server side or during transmission. (2) Poor resource utilization and training efficiency: training on centralized server with potentially huge low-rank matrices is computational inefficient. In this paper, we propose a novel decentralized gradient-quantization based matrix factorization (DGMF) framework to address the above limitations in POI recommendation. Compared with the centralized MF methods which store all sensitive data and low-rank matrices during model training, DGMF treats each user’s device (e.g., phone) as an independent learner and keeps the sensitive data on each user’s end. Furthermore, a privacy-preserving and communication-efficient mechanism with gradient-quantization technique is presented to train the proposed model, which aims to handle the privacy problem and reduces the communication cost in the decentralized setting. Theoretical guarantees of the proposed algorithm and experimental studies on real-world datasets demonstrate the effectiveness of the proposed algorithm.
Xuebin Zhou, Zhibin Hu, Jian Chen 0011
J. Artif. Intell. Res.4
2023 Improving fine-tuning of self-supervised models with Contrastive Initialization
Haolin Pan, Qinyi Deng, Haomin Yang, Jian Chen 0011
Neural Networks5
2023 Efficient Decomposition Selection for Multi-class Classification
abstract
Choosing a decomposition method for multi-class classification is an important trade-off between efficiency and predictive accuracy. Trying all the decomposition methods to find the best one is too time-consuming for many applications, while choosing the wrong one may result in large loss on predictive accuracy. In this paper, we propose an automatic decomposition method selection approach called “D-Chooser”, which is lightweight and can choose the best decomposition method accurately. D-Chooser is equipped with our proposed difficulty index which consists of sub-metrics including distribution divergence, overlapping regions, unevenness degree and relative size of the solution space. The difficulty index has two intriguing properties: 1) fast to compute and 2) measuring multi-class problems comprehensively. Extensive experiments on real-world multi-class problems show that D-Chooser achieves an accuracy of 80.56% in choosing the best decomposition method. It can choose the best method in just a few seconds, while existing approaches verify the effectiveness of a decomposition method often takes a few hours. We also provide case studies on Kaggle competitions and the results confirm that D-Chooser is able to choose a better decomposition method than the winning solutions.
Zeyi Wen, Bingsheng He, Jian Chen 0011
IEEE Trans. Knowl. Data Eng.4
2022 Efficient Second-Order Optimization for Neural Networks with Kernel Machines
abstract
Second-order optimization has been recently explored in neural network training. However, the recomputation of the Hessian matrix in the second-order optimization posts much extra computation and memory burden in the training. There have been some attempts to address this issue by approximation on the Hessian matrix, which unfortunately degrades the performance of the neural models. In order to tackle this issue, we propose Kernel Stochastic Gradient Descent (Kernel SGD) which solves the optimization problem in a space transformed by the Hessian matrix of the kernel machine. Kernel SGD eliminates the Hessian matrix recomputation in the training and requires a much smaller memory cost which can be controlled via the mini-batch size. We show that Kernel SGD optimization is theoretically guaranteed to converge. Our experimental results on tabular, image and text data confirm that Kernel SGD converges up to 30 times faster than the existing second-order optimization techniques, and achieves the highest test accuracy on all the tasks tested. Kernel SGD even outperforms the first-order optimization baselines in some problems tested in our experiments.
Yile Chen 0004, Jian Chen 0011, Zeyi Wen, Jin Huang 0003
CIKM3
2022 Tackling background ambiguities in multi-class few-shot point cloud semantic segmentation
Lvlong Lai, Jian Chen 0011, Chi Zhang 0007, Zehong Zhang, Guosheng Lin, Qingyao Wu
Knowl. Based Syst.2
2022 Towards Accurate and Compact Architectures via Neural Architecture Transformer
abstract
Designing effective architectures is one of the key factors behind the success of deep neural networks. Existing deep architectures are either manually designed or automatically searched by some Neural Architecture Search (NAS) methods. However, even a well-designed/searched architecture may still contain many nonsignificant or redundant modules/operations (e.g., some intermediate convolution or pooling layers). Such redundancy may not only incur substantial memory consumption and computational cost but also deteriorate the performance. Thus, it is necessary to optimize the operations inside an architecture to improve the performance without introducing extra computational cost. To this end, we have proposed a Neural Architecture Transformer (NAT) method which casts the optimization problem into a Markov Decision Process (MDP) and seeks to replace the redundant operations with more efficient operations, such as skip or null connection. Note that NAT only considers a small number of possible replacements/transitions and thus comes with a limited search space. As a result, such a small search space may hamper the performance of architecture optimization. To address this issue, we propose a Neural Architecture Transformer++ (NAT++) method which further enlarges the set of candidate transitions to improve the performance of architecture optimization. Specifically, we present a two-level transition rule to obtain valid transitions, i.e., allowing operations to have more efficient types (e.g., convolution → separable convolution) or smaller kernel sizes (e.g., 5×5 → 3×3). Note that different operations may have different valid transitions. We further propose a Binary-Masked Softmax (BMSoftmax) layer to omit the possible invalid transitions. Last, based on the MDP formulation, we apply policy gradient to learn an optimal policy, which will be used to infer the optimized architectures. Extensive experiments show that the transformed architectures significantly outperform both their original counterparts and the architectures optimized by existing methods.
Mingkui Tan, Qi Chen 0014, Jian Chen 0011, Peilin Zhao, Junzhou Huang
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 CAM-based non-local attention network for weakly supervised fire detection
Lvlong Lai, Jian Chen 0011, Qingyao Wu
Serv. Oriented Comput. Appl.3
2022 Traffic Node Importance Evaluation Based on Clustering in Represented Transportation Networks
abstract
With the ability to acquire and process large-scale traffic big data, a cooperative intelligent transport system can be realized. Identifying important nodes in a traffic network contributes to better traffic control, which plays a more critical role in improving the traffic efficiency of a cooperative intelligent transport system. However, existing traffic node importance evaluation methods rely on manually designed metrics such as betweenness, degree, which may lead to biased results. Meanwhile, the traditional method of iteratively deleting nodes is unsuitable for a large-scale traffic network. In this paper, we propose a novel traffic node importance evaluation method based on clustering in represented transportation network. Specifically, the proposed method first construct a length-weighted network based on the geographic road network. Then, it learns the low-dimensional embeddings of nodes employing network representation learning. Finally, it clusters each nodes with a machine learning method and identifies the critical nodes through vehicle flow of different nodes. Experimental results on the real-word dataset show that our proposed method has excellent performance compared to baseline methods.
Xinlong Huang, Jian Chen 0011, Wei Wang 0077, Xiping Hu
IEEE Trans. Intell. Transp. Syst.2
2022 A Tensor-Based Markov Chain Model for Heterogeneous Information Network Collective Classification
abstract
Heterogeneous Information Network (HIN) collecitve classification studies the problem of predicting labels for one type of nodes in a HIN which contains multiple types of nodes multiple types of links among them. Previous studies have revealed that exploiting relative importance of links is quite useful to improve node classification performance as connected nodes tend to have similar labels. Most existing approaches exploit the relative importance of links either by directly counting the number of connections among nodes or by learning the weight of each type of link from labeled data only. However, these approaches either neglect the importance of types of links to the class labels or may lead to overfitting problem. We propose aTensor-basedMarkov chain (T-Mark) approach, which is able to automatically and simultaneously predict the labels for unlabeled nodes and give the relative importance of types of links that actually improve the classification accuracy. Specifically, we build two tensor equations by using the HIN and features of nodes from both labeled and unlabeled data. A Markov chain-based model is proposed and it is solved by an iterative process to obtain the stationary distributions. Theoretical analyses of the existence and uniqueness of such probability distributions are given. Extensive experimental results demonstrate that T-Mark is able to achieve superior performance in the comparison and obtain reasonable relative importance of links.
Chao Han 0002, Jian Chen 0011, Mingkui Tan, Michael Kwok-Po Ng, Qingyao Wu
IEEE Trans. Knowl. Data Eng.2
2022 Parallel and Distributed Structured SVM Training
abstract
Structured Support Vector Machines (structured SVMs) are a fundamental machine learning algorithm, and have solid theoretical foundation and high effectiveness in applications such as natural language parsing and computer vision. However, training structured SVMs is very time-consuming, due to the large number of constraints and inferior convergence rates, especially for large training data sets. The high cost of training structured SVMs has hindered its adoption to new applications. In this article, we aim to improve the efficiency of structured SVMs by proposing a parallel and distributed solution (namelyFastSSVM) for training structured SVMs building on top of MPI and OpenMP. FastSSVM exploits a series of optimizations (e.g., optimizations on data storage and synchronization) to efficiently use the resources of the nodes in a cluster and the cores of the nodes. Moreover, FastSSVM tackles the large constraint set problem by batch processing and addresses the slow convergence challenge by adapting stop conditions based on the improvement of each iteration. We theoretically prove that our solution is guaranteed to converge to a global optimum. A comprehensive experimental study shows that FastSSVM can achieve at least four times speedup over the existing solutions, and in some cases can achieve two to three orders of magnitude speedup.
Jiantong Jiang, Zeyi Wen, Zeke Wang, Bingsheng He, Jian Chen 0011
IEEE Trans. Parallel Distributed Syst.5
2021 Enhancing SVMs with Problem Context Aware Pipeline
abstract
In recent years, many data mining practitioners have treated deep neural networks (DNNs) as a standard recipe of creating the state-of-the-art solutions. As a result, models like Support Vector Machines (SVMs) have been overlooked. While the results from DNNs are encouraging, DNNs also come with their huge number of parameters in the model and overheads in long training/inference time. SVMs have excellent properties such as convexity, good generality and efficiency. In this paper, we propose techniques to enhance SVMs with an automatic pipeline which exploits the context of the learning problem. The pipeline consists of several components including data aware subproblem construction, feature customization, data balancing among subproblems with augmentation, and kernel hyper-parameter tuner. Comprehensive experiments show that our proposed solution is more efficient, while producing better results than the other SVM based approaches. Additionally, we conduct a case study of our proposed solution on a popular sentiment analysis problem---the aspect term sentiment analysis (ATSA) task. The study shows that our SVM based solution can achieve competitive predictive accuracy to DNN (and even majority of the BERT) based approaches. Furthermore, our solution is about 40 times faster in inference and has 100 times fewer parameters than the models using BERT. Our findings can encourage more research work on conventional machine learning techniques which may be a good alternative for smaller model size and faster training/inference.
Zeyi Wen, Zhishang Zhou, Hanfeng Liu, Bingsheng He, Xia Li 0007, Jian Chen 0011
KDD6
2021 StackRec: Efficient Training of Very Deep Sequential Recommender Models by Iterative Stacking
abstract
Deep learning has brought great progress for the sequential recommendation (SR) tasks. With advanced network architectures, sequential recommender models can be stacked with many hidden layers, e.g., up to 100 layers on real-world recommendation datasets. Training such a deep network is difficult because it can be computationally very expensive and takes much longer time, especially in situations where there are tens of billions of user-item interactions. To deal with such a challenge, we present StackRec, a simple, yet very effective and efficient training framework for deep SR models by iterative layer stacking. Specifically, we first offer an important insight that hidden layers/blocks in a well-trained deep SR model have very similar distributions. Enlightened by this, we propose the stacking operation on the pre-trained layers/blocks to transfer knowledge from a shallower model to a deep model, then we perform iterative stacking so as to yield a much deeper but easier-to-train SR model. We validate the performance of StackRec by instantiating it with four state-of-the-art SR models in three practical scenarios with real-world datasets. Extensive experiments show that StackRec achieves not only comparable performance, but also substantial acceleration in training time, compared to SR models that are trained from scratch. Codes are available at https://github.com/wangjiachun0426/StackRec.
Jiachun Wang, Fajie Yuan, Jian Chen 0011, Qingyao Wu, Min Yang 0007, Guoxiao Zhang
SIGIR3
2021 Neural graph personalized ranking for Top-N Recommendation
Zhibin Hu, Jiachun Wang, Yan Yan 0006, Peilin Zhao, Jian Chen 0011, Jin Huang 0003
Knowl. Based Syst.5
2021 Content-aware convolutional neural networks
Yaofo Chen, Mingkui Tan, Kui Jia, Jian Chen 0011, Jingdong Wang 0001
Neural Networks5
2021 Deep Learning for Image Super-Resolution: A Survey
abstract
Image Super-Resolution (SR) is an important class of image processing techniqueso enhance the resolution of images and videos in computer vision. Recent years have witnessed remarkable progress of image super-resolution using deep learning techniques. This article aims to provide a comprehensive survey on recent advances of image super-resolution using deep learning approaches. In general, we can roughly group the existing studies of SR techniques into three major categories: supervised SR, unsupervised SR, and domain-specific SR. In addition, we also cover some other important issues, such as publicly available benchmark datasets and performance evaluation metrics. Finally, we conclude this survey by highlighting several future directions and open issues which should be further addressed by the community in the future.
Zhihao Wang 0004, Jian Chen 0011, Steven C. H. Hoi
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Conditional Automated Channel Pruning for Deep Neural Networks
abstract
Channel pruning has become one of the predominant compression methods to deploy deep models on resource-constrained devices. Most channel pruning methods often use a fixed compression rate for all the layers of the model, which, however, may not be optimal. To address this issue, given a specific target compression rate, one can search for the optimal compression rate for each layer via some automated methods. Nevertheless, when we consider multiple compression rates, these methods have to repeat the channel pruning process multiple times, once for each rate, which can be unnecessary and inefficient. To tackle the problem, we propose a Conditional Automated Channel Pruning (CACP) method which simultaneously produces compressed models under different compression rates through a single channel pruning process. Specifically, CACP takes a set of compression rates and the original model as its input, and outputs the feasible compressed models that satisfy the considered compression rates. To learn CACP, we cast the layer-by-layer channel pruning process into a Markov decision process (MDP), in which we seek to solve a series of decision-making problems. Based on MDP, we develop a reinforcement learning (RL) framework with deep deterministic policy gradient (DDPG) to learn the optimal policy. To satisfy the constraint items in the optimization problem, we design a constraint-guaranteed method, which guides the agent to search for compressed models that satisfy the computational constraints by limiting the action space. Extensive experiments on CIFAR-10 and ImageNet datasets demonstrate the superiority of our method over existing methods.
Yixin Liu 0002, Luoqian Jiang, Jian Chen 0011
IEEE Signal Process. Lett.5
2020 Closed-Loop Matters: Dual Regression Networks for Single Image Super-Resolution
abstract
Deep neural networks have exhibited promising performance in image super-resolution (SR) by learning a nonlinear mapping function from low-resolution (LR) images to high-resolution (HR) images. However, there are two underlying limitations to existing SR methods. First, learning the mapping function from LR to HR images is typically an ill-posed problem, because there exist infinite HR images that can be downsampled to the same LR image. As a result, the space of the possible functions can be extremely large, which makes it hard to find a good solution. Second, the paired LR-HR data may be unavailable in real-world applications and the underlying degradation method is often unknown. For such a more general case, existing SR models often incur the adaptation problem and yield poor performance. To address the above issues, we propose a dual regression scheme by introducing an additional constraint on LR data to reduce the space of the possible functions. Specifically, besides the mapping from LR to HR images, we learn an additional dual regression mapping estimates the down-sampling kernel and reconstruct LR images, which forms a closed-loop to provide additional supervision. More critically, since the dual regression process does not depend on HR images, we can directly learn from LR images. In this sense, we can easily adapt SR models to real-world data, e.g., raw video frames from YouTube. Extensive experiments with paired training data and unpaired real-world data demonstrate our superiority over existing methods.
Jian Chen 0011, Jingdong Wang 0001, Qi Chen 0014, Jiezhang Cao, Zeshuai Deng, Yanwu Xu 0001, Mingkui Tan
CVPR2
2020 Fg2seq: Effectively Encoding Knowledge for End-To-End Task-Oriented Dialog
abstract
End-to-end Task-oriented spoken dialog systems typically require modeling two types of inputs, namely, the dialog history which is a sequence of utterances and the knowledge base (KB) associated with the dialog history. While modeling these inputs, current state-of-the-art models typically ignore the rich structure in the knowledge graph or its intrinsic association with the dialog history. In this paper, we propose a Flow-to-Graph seq2seq model (FG2Seq) which can effectively encode knowledge by considering inherent structural information of the knowledge graph and latent semantic information from dialog history. Experiments on two publicly available task oriented dialog datasets show that our proposed FG2Seq achieves robust performance on generating appropriate system responses and outperforms the baseline systems.
Zhenhao He, Qingyao Wu, Jian Chen 0011
ICASSP4
2020 Breaking the Curse of Space Explosion: Towards Efficient NAS with Curriculum Search
abstract
Neural architecture search (NAS) has become an important approach to automatically find effective architectures. To cover all possible good architectures, we need to search in an extremely large search space with billions of candidate architectures. More critically, given a large search space, we may face a very challenging issue of space explosion. However, due to the limitation of computational resources, we can only sample a very small proportion of the architectures, which provides insufficient information for the training. As a result, existing methods may often produce sub-optimal architectures. To alleviate this issue, we propose a curriculum search method that starts from a small search space and gradually incorporates the learned knowledge to guide the search in a large space. With the proposed search strategy, our Curriculum Neural Architecture Search (CNAS) method significantly improves the search efficiency and finds better architectures than existing NAS methods. Extensive experiments on CIFAR-10 and ImageNet demonstrate the effectiveness of the proposed method.
Yaofo Chen, Peilin Zhao, Jian Chen 0011, Junzhou Huang, Mingkui Tan
ICML5
2020 PointDrop: Improving Object Detection from Sparse Point Clouds via Adversarial Data Augmentation
abstract
Current 3D object detection methods achieve accurate and efficient results on the standard point cloud dataset. However, in real-world applications, the point cloud samples obtained in the real-time running may be much sparser due to various reasons (occlusion, low reflectivity of objects and fewer laser beams) and existing methods do not consider the limitations of their models on sparse point clouds. To improve the robustness of an object detector to sparser point clouds, we propose PointDrop, which learns to drop the features of some key points in the point clouds to generate challenging sparse samples for data augmentation. Moreover, PointDrop is able to adjust the difficulty of the generated samples based on the capacity of the detector and thus progressively improve the performance of the detector. We create two sparse point clouds datasets from the KITTI dataset to evaluate our method, and the experimental results show that PointDrop significantly improves the robustness of the detector to sparse point clouds.
Jian Chen 0011
ICPR2
2020 Task-Oriented Dialog Generation with Enhanced Entity Representation
Zhenhao He, Jiachun Wang, Jian Chen 0011
INTERSPEECH3
2020 ThunderGBM: Fast GBDTs and Random Forests on GPUs
abstract
Gradient Boosting Decision Trees (GBDTs) and Random Forests (RFs) have been used in many real-world applications. They are often a standard recipe for building state-of-the-art solutions to machine learning and data mining problems. However, training and prediction are very expensive computationally for large and high dimensional problems. This article presents an efficient and open source software toolkit called ThunderGBM which exploits the high-performance Graphics Processing Units (GPUs) for GBDTs and RFs. ThunderGBM supports classification, regression and ranking, and can run on single or multiple GPUs of a machine. Our experimental results show that ThunderGBM outperforms the existing libraries while producing similar models, and can handle high dimensional problems where existing GPU-based libraries fail. Documentation, examples, and more details about ThunderGBM are available at https://github.com/xtra-computing/thundergbm.
Zeyi Wen, Hanfeng Liu, Jiashuai Shi, Qinbin Li, Bingsheng He, Jian Chen 0011
J. Mach. Learn. Res.6
2020 Multi-way backpropagation for training compact deep neural networks
Jian Chen 0011, Anton van den Hengel, Qinfeng Shi, Mingkui Tan
Neural Networks2
2020 Hierarchical Neural Architecture Search for Single Image Super-Resolution
abstract
Deep neural networks have exhibited promising performance in image super-resolution (SR). Most SR models follow a hierarchical architecture that contains both the cell-level design of computational blocks and the network-level design of the positions of upsampling blocks. However, designing SR models heavily relies on human expertise and is very labor-intensive. More critically, these SR models often contain a huge number of parameters and may not meet the requirements of computation resources in real-world applications. To address the above issues, we propose a Hierarchical Neural Architecture Search (HNAS) method to automatically design promising architectures with different requirements of computation cost. To this end, we design a hierarchical SR search space and propose a hierarchical controller for architecture search. Such a hierarchical controller is able to simultaneously find promising cell-level blocks and network-level positions of upsampling layers. Moreover, to design compact architectures with promising performance, we build a joint reward by considering both the performance and computation cost to guide the search process. Extensive experiments on five benchmark datasets demonstrate the superiority of our method over existing methods.
Yongsheng Luo, Zhenhao He, Jin Huang 0003, Jian Chen 0011
IEEE Signal Process. Lett.5
2020 Scripted Video Generation With a Bottom-Up Generative Adversarial Network
abstract
Generating videos given a text description (such as a script) is non-trivial due to the intrinsic complexity of image frames and the structure of videos. Although Generative Adversarial Networks (GANs) have been successfully applied to generate images conditioned on a natural language description, it is still very challenging to generate realistic videos in which the frames are required to follow both spatial and temporal coherence. In this paper, we propose a novel Bottom-up GAN (BoGAN) method for generating videos given a text description. To ensure the coherence of the generated frames and also make the whole video match the language descriptions semantically, we design a bottom-up optimisation mechanism to train BoGAN. Specifically, we devise a region-level loss via attention mechanism to preserve the local semantic alignment and draw details in different sub-regions of video conditioned on words which are most relevant to them. Moreover, to guarantee the matching between text and frame, we introduce a frame-level discriminator, which can also maintain the fidelity of each frame and the coherence across frames. Last, to ensure the global semantic alignment between whole video and given text, we apply a video-level discriminator. We evaluate the effectiveness of the proposed BoGAN on two synthetic datasets (i.e., SBMG and TBMG) and two real-world datasets (i.e., MSVD and KTH).
Qi Chen 0014, Qi Wu 0001, Jian Chen 0011, Qingyao Wu, Anton van den Hengel, Mingkui Tan
IEEE Trans. Image Process.3
2019 Efficient Multi-Class Probabilistic SVMs on GPUs
abstract
Multi-class SVMs with the probabilistic output (MP-SVMs) are important techniques in pattern recognition. Two key challenges for efficient GPU accelerations for MP-SVM are: (i) many kernel values are repeatedly computed as a binary SVM classifier is trained iteratively, resulting in repeated accesses to the high latency GPU memory; (ii) performing training or estimating probability in parallel requires a much larger memory footprint than the GPU memory. To overcome the challenges, we propose GMP-SVM to reduce high latency memory accesses and memory consumption through batch processing, computation/data reusing and sharing. Experimental results show that our solution (available in https://github.com/Xtra-Computing/thundersvm) outperforms LibSVM by 100 times while retaining the same accuracy.
Zeyi Wen, Jiashuai Shi, Bingsheng He, Jian Chen 0011
ICDE4
2019 Multi-Level Visual-Semantic Alignments with Relation-Wise Dual Attention Network for Image and Text Matching
abstract
Image-text matching is central to visual-semantic cross-modal retrieval and has been attracting extensive attention recently. Previous studies have been devoted to finding the latent correspondence between image regions and words, e.g., connecting key words to specific regions of salient objects. However, existing methods are usually committed to handle concrete objects, rather than abstract ones, e.g., a description of some action, which in fact are also ubiquitous in description texts of real-world. The main challenge in dealing with abstract objects is that there is no explicit connections between them, unlike their concrete counterparts. One therefore has to alternatively find the implicit and intrinsic connections between them. In this paper, we propose a relation-wise dual attention network (RDAN) for image-text matching. Specifically, we maintain an over-complete set that contains pairs of regions and words. Then built upon this set, we encode the local correlations and the global dependencies between regions and words by training a visual-semantic network. Then a dual pathway attention network is presented to infer the visual-semantic alignments and image-text similarity. Extensive experiments validate the efficacy of our method, by achieving the state-of-the-art performance on several public benchmark datasets.
Zhibin Hu, Yongsheng Luo, Jiong Lin, Yan Yan 0006, Jian Chen 0011
IJCAI5
2019 NAT: Neural Architecture Transformer for Accurate and Compact Architectures
abstract
Designing effective architectures is one of the key factors behind the success of deep neural networks. Existing deep architectures are either manually designed or automatically searched by some Neural Architecture Search (NAS) methods. However, even a well-searched architecture may still contain many non-significant or redundant modules or operations (e.g., convolution or pooling), which may not only incur substantial memory consumption and computation cost but also deteriorate the performance. Thus, it is necessary to optimize the operations inside an architecture to improve the performance without introducing extra computation cost. Unfortunately, such a constrained optimization problem is NP-hard. To make the problem feasible, we cast the optimization problem into a Markov decision process (MDP) and seek to learn a Neural Architecture Transformer (NAT) to replace the redundant operations with the more computationally efficient ones (e.g., skip connection or directly removing the connection). Based on MDP, we learn NAT by exploiting reinforcement learning to obtain the optimization policies w.r.t. different architectures. To verify the effectiveness of the proposed strategies, we apply NAT on both hand-crafted architectures and NAS based architectures. Extensive experiments on two benchmark datasets, i.e., CIFAR-10 and ImageNet, demonstrate that the transformed architecture by NAT significantly outperforms both its original form and those architectures optimized by existing methods.
Mingkui Tan, Qi Chen 0014, Jian Chen 0011, Peilin Zhao, Junzhou Huang
NeurIPS5
2019 Efficient Multi-Class Probabilistic SVMs on GPUs
abstract
Recently, many researchers have been working on improving other traditional machine learning algorithms (besides deep learning) using high-performance hardware such as Graphics Processing Units (GPUs). The recent success of machine learning is not only due to more effective algorithms, but also more efficient systems and implementations. In this paper, we propose a novel and efficient solution to multi-class SVMs with probabilistic output (MP-SVMs) accelerated by GPUs. MP-SVMs are an important technique for many pattern recognition applications. However, MP-SVMs are very time-consuming to use, because using an MP-SVM classifier requires training many binary SVMs and performing probability estimation by combining results of all the binary SVMs. GPUs have much higher computation capability than CPUs and are potentially excellent hardware to accelerate MP-SVMs. Still, two key challenges for efficient GPU accelerations for MP-SVM are: (i) many kernel values are repeatedly computed as a binary SVM classifier is trained iteratively, resulting in repeated accesses to the high latency GPU memory; (ii) performing training or estimating probability in a highly parallel way requires a much larger memory footprint than the GPU memory. To overcome the challenges, we propose a solution called GMP-SVM which exploits two-level (i.e., binary SVM level and MP-SVM level) optimization for training MP-SVMs and high parallelism for estimating probability. GMP-SVM reduces high latency memory accesses and memory consumption through batch processing, kernel value reusing and sharing, and support vector sharing. Experimental results show that GMP-SVM outperforms the GPU baseline by two to five times, and LibSVM with OpenMP by an order of magnitude. Also, GMP-SVM produces the same SVM classifier as LibSVM.
Zeyi Wen, Jiashuai Shi, Bingsheng He, Jian Chen 0011
IEEE Trans. Knowl. Data Eng.4
2019 Auto-Embedding Generative Adversarial Networks For High Resolution Image Synthesis
abstract
Generating images via a generative adversarial network (GAN) has attracted much attention recently. However, most of the existing GAN-based methods can only produce lowresolution images of limited quality. Directly generating highresolution images using GANs is nontrivial, and often produces problematic images with incomplete objects. To address this issue, we develop a novel GAN called auto-embedding generative adversarial network, which simultaneously encodes the global structure features and captures the fine-grained details. In our network, we use an autoencoder to learn the intrinsic high-level structure of real images and design a novel denoiser network to provide photo-realistic details for the generated images. In the experiments, we are able to produce 512 × 512 images of promising quality directly from the input noise. The resultant images exhibit better perceptual photo-realism, that is, with sharper structure and richer details, than other baselines on several datasets, including Oxford-102 Flowers, Caltech-UCSD Birds (CUB), High-Quality Large-scale CelebFaces Attributes (CelebAHQ), Large-scale Scene Understanding (LSUN), and ImageNet.
Qi Chen 0014, Jian Chen 0011, Qingyao Wu, Qinfeng Shi, Mingkui Tan
IEEE Trans. Multim.3
2019 Exploiting GPUs for Efficient Gradient Boosting Decision Tree Training
abstract
In this paper, we present a novel parallel implementation for training Gradient Boosting Decision Trees (GBDTs) on Graphics Processing Units (GPUs). Thanks to the excellent results on classification/regression and the open sourced libraries such as XGBoost, GBDTs have become very popular in recent years and won many awards in machine learning and data mining competitions. Although GPUs have demonstrated their success in accelerating many machine learning applications, it is challenging to develop an efficient GPU-based GBDT algorithm. The key challenges include irregular memory accesses, many sorting operations with small inputs and varying data parallel granularities in tree construction. To tackle these challenges on GPUs, we propose various novel techniques including (i) Run-length Encoding compression and thread/block workload dynamic allocation, (ii) data partitioning based on stable sort, and fast and memory efficient attribute ID lookup in node splitting, (iii) finding approximate split points using two-stage histogram building, (iv) building histograms with the aware of sparsity and exploiting histogram subtraction to reduce histogram building workload, (v) reusing intermediate training results for efficient gradient computation, and (vi) exploiting multiple GPUs to handle larger data sets efficiently. Our experimental results show that our algorithm named ThunderGBM can be 10x times faster than the state-of-the-art libraries (i.e., XGBoost, LightGBM and CatBoost) running on a relatively high-end workstation of 20 CPU cores. In comparison with the libraries on GPUs, ThunderGBM can handle higher dimensional problems which the libraries become extremely slow or simply fail. For the data sets the existing libraries on GPUs can handle, ThunderGBM achieves up to 10 times speedup on the same hardware, which demonstrates the significance of our GPU optimizations. Moreover, the models trained by ThunderGBM are identical to those trained by XGBoost, and have similar quality as those trained by LightGBM and CatBoost.
Zeyi Wen, Jiashuai Shi, Bingsheng He, Jian Chen 0011, Kotagiri Ramamohanarao, Qinbin Li
IEEE Trans. Parallel Distributed Syst.4
2018 Selecting Proper Multi-Class SVM Training Methods
abstract
Support Vector Machines (SVMs) are excellent candidate solutions to solving multi-class problems, and multi-class SVMs can be trained by several different methods. Different training methods commonly produce SVMs with different effectiveness, and no multi-class SVM training method always outperforms other multi-class SVM training methods on all problems. This raises difficulty for practitioners to choose the best training method for a given problem. In this work, we propose a Multi-class Method Selection (MMS) approach to help users select the most appropriate method among one-versus-one (OVO), one-versus-all (OVA) and structural SVMs (SSVMs) for a given problem. Our key idea is to select the training method based on the distribution of training data and the similarity between different classes. Using the distribution and class similarity, we estimate the unclassifiable rate of each multi-class SVM training method, and select the training method with the minimum unclassifiable rate. Our initial findings show: (i) SSVMs with linear kernel perform worse than OVO and OVA; (ii) MMS often produces SVM classifiers that can confidently classify unseen instances.
Zeyi Wen, Jian Chen 0011, Jin Huang 0003
AAAI3
2018 Double Forward Propagation for Memorized Batch Normalization
abstract
Batch Normalization (BN) has been a standard component in designing deep neural networks (DNNs). Although the standard BN can significantly accelerate the training of DNNs and improve the generalization performance, it has several underlying limitations which may hamper the performance in both training and inference. In the training stage, BN relies on estimating the mean and variance of data using a single mini-batch. Consequently, BN can be unstable when the batch size is very small or the data is poorly sampled. In the inference stage, BN often uses the so called moving mean and moving variance instead of batch statistics, i.e., the training and inference rules in BN are not consistent. Regarding these issues, we propose a memorized batch normalization (MBN), which considers multiple recent batches to obtain more accurate and robust statistics. Note that after the SGD update for each batch, the model parameters will change, and the features will change accordingly, leading to the Distribution Shift before and after the update for the considered batch. To alleviate this issue, we present a simple Double-Forward scheme in MBN which can further improve the performance. Compared to related methods, the proposed MBN exhibits consistent behaviors in both training and inference. Empirical results show that the MBN based models trained with the Double-Forward scheme greatly reduce the sensitivity of data and significantly improve the generalization performance.
Qingyao Wu, Chaorui Deng, Jian Chen 0011, Mingkui Tan
AAAI4
2018 Efficient Support Vector Machine Training Algorithm on GPUs
abstract
Support Vector Machines (SVMs) are popular for many machine learning tasks. With rapid growth of dataset size, the high cost of training limits the wide use of SVMs. Several SVM implementations on GPUs have been proposed to accelerate SVMs. However, they support only classification (SVC) or regression (SVR). In this work, we propose a simple and effective SVM training algorithm on GPUs which can be used for SVC, SVR and one-class SVM. Initial experiments show that our implementation outperforms existing ones. We are in the process of encapsulating our algorithm into an easy-to-use library which has Python, R and MATLAB interfaces.
Jiashuai Shi, Zeyi Wen, Bingsheng He, Jian Chen 0011
AAAI4
2018 ThunderSVM: A Fast SVM Library on GPUs and CPUs
abstract
Support Vector Machines (SVMs) are classic supervised learning models for classification, regression and distribution estimation. A survey conducted by Kaggle in 2017 shows that 26% of the data mining and machine learning practitioners are users of SVMs. However, SVM training and prediction are very expensive computationally for large and complex problems. This paper presents an efficient and open source SVM software toolkit called ThunderSVM which exploits the high-performance of Graphics Processing Units (GPUs) and multi-core CPUs. ThunderSVM supports all the functionalities–including classification (SVC), regression (SVR) and one-class SVMs–of LibSVM and uses identical command line options, such that existing LibSVM users can easily apply our toolkit. ThunderSVM can be used through multiple language interfaces including C/C++, Python, R and MATLAB. Our experimental results show that ThunderSVM is generally an order of magnitude faster than LibSVM while producing identical SVMs. In addition to the high efficiency, we design our convex optimization solver in a general way such that SVC, SVR, and one-class SVMs share the same solver for the ease of maintenance. Documentation, examples, and more about ThunderSVM are available at https://github.com/zeyiwen/thundersvm
Zeyi Wen, Jiashuai Shi, Qinbin Li, Bingsheng He, Jian Chen 0011
J. Mach. Learn. Res.5
2017 Improving Efficiency of SVM k-Fold Cross-Validation by Alpha Seeding
abstract
The k-fold cross-validation is commonly used to evaluate the effectiveness of SVMs with the selected hyper-parameters. It is known that the SVM k-fold cross-validation is expensive, since it requires training k SVMs. However, little work has explored reusing the h-th SVM for training the (h+1)-th SVM for improving the efficiency of k-fold cross-validation. In this paper, we propose three algorithms that reuse the h-th SVM for improving the efficiency of training the (h+1)-th SVM. Our key idea is to efficiently identify the support vectors and to accurately estimate their associated weights (also called alpha values) of the next SVM by using the previous SVM. Our experimental results show that our algorithms are several times faster than the k-fold cross-validation which does not make use of the previously trained SVM. Moreover, our algorithms produce the same results (hence same accuracy) as the k-fold cross-validation which does not make use of the previously trained SVM.
Zeyi Wen, Bin Li 0073, Kotagiri Ramamohanarao, Jian Chen 0011, Rui Zhang 0003
AAAI4
2017 Tensor Based Relations Ranking for Multi-relational Collective Classification
abstract
In this paper, we study relations ranking and object classification for multi-relational data where objects are interconnected by multiple relations. The relations among objects should be exploited for achieving a good classification. While most existing approaches exploit either by directly counting the number of connections among objects or by learning the weight of each relation from labeled data only. In this paper, we propose an algorithm, TensorRRCC, which is able to determine the ranking of relations and the labels of objects simultaneously. Our basic idea is that highly ranked relations within a class should play more important roles in object classification, and class membership information is important for determining a ranking quality over the relations w.r.t. a specific learning task. TensorRRCC implements the idea by modeling a Markov chain on transition probability graphs from connection and feature information with both labeled and unlabeled objects and propagates the ranking scores of relations and relevant classes of objects. An iterative progress is proposed to solve a set of tensor equations to obtain the stationary distribution of relations and objects. We compared our algorithm with current collective classification algorithms on two real-world data sets and the experimental results show the superiority of our method.
Chao Han 0002, Qingyao Wu, Michael Kwok-Po Ng, Jiezhang Cao, Mingkui Tan, Jian Chen 0011
ICDM6
2017 Age classification with deep learning face representation
Jin Huang 0007, Bin Li 0073, Jia Zhu 0003, Jian Chen 0011
Multim. Tools Appl.4
2016 Joint Classification with Heterogeneous Labels Using Random Walk with Dynamic Label Propagation
Yongxin Liao, Shenxi Yuan, Jian Chen 0011, Qingyao Wu, Bin Li 0073
PAKDD (1)3
2016 Online Feature Selection of Class Imbalance via PA Algorithm
Chao Han 0002, Yun-Kun Tan, Jin-Hui Zhu, Jian Chen 0011, Qingyao Wu
J. Comput. Sci. Technol.5
2016 ML-FOREST: A Multi-Label Tree Ensemble Method for Multi-Label Classification
abstract
Multi-label classification deals with the problem where each example is associated with multiple class labels. Since the labels are often dependent to other labels, exploiting label dependencies can significantly improve the multi-label classification performance. The label dependency in existing studies is often given as prior knowledge or learned from the labels only. However, in many real applications, such prior knowledge may not be available, or labeled information might be very limited. In this paper, we propose a new algorithm, called Ml-Forest , to learn an ensemble of hierarchical multi-label classifier trees to reveal the intrinsic label dependencies. In Ml-Forest, we construct a set of hierarchical trees, and develop a label transfer mechanism to identify the multiple relevant labels in a hierarchical way. In general, the relevant labels at higher levels of the trees capture more discriminable label concepts, and they will be transferred into lower level children nodes that are harder to discriminate. The relevant labels in the hierarchy are then aggregated to compute label dependency and make the final prediction. Our empirical study shows encouraging results of the proposed algorithm in comparison with the state-of-the-art multi-label classification algorithms under Friedman test and post-hoc Nemenyi test.
Qingyao Wu, Mingkui Tan, Hengjie Song, Jian Chen 0011, Michael Kwok-Po Ng
IEEE Trans. Knowl. Data Eng.4
2016 Heads-Join: Efficient Earth Mover's Distance Similarity Joins on Hadoop
abstract
The Earth Mover's Distance (EMD) similarity join has a number of important applications such as near duplicate image retrieval and distributed based pattern analysis. However, the computational cost of EMD is super cubic and consequently the EMD similarity join operation is prohibitive for datasets of even medium size. We propose to employ the Hadoop platform to speed up the operation. Simply porting the state-of-the-art metric distance similarity join algorithms to Hadoop results in inefficiency because they involve excessive distance computations and are vulnerable to skewed data distributions. We propose a novel framework, named HEADS-JOIN, which transforms data into the space of EMD lower bounds and performs pruning and partitioning at a low cost because computing these EMD lower bounds has constant or linear complexity. We investigate both range and top-k joins, and design efficient algorithms on three popular Hadoop computation paradigms, i.e., MapReduce, Bulk Synchronous Parallel, and Spark. We conduct extensive experiments on both real and synthetic datasets. The results show that HEADS-JOIN outperforms the state-of-the-art metric similarity join technique, i.e., Quickjoin, by up to an order of magnitude and scales out well.
Jin Huang 0003, Rui Zhang 0003, Rajkumar Buyya, Jian Chen 0011, Yongwei Wu 0001
IEEE Trans. Parallel Distributed Syst.4
2015 A privacy-enhancing model for location-based personalized recommendations
Jin Huang 0007, Jianzhong Qi 0001, Yabo Xu, Jian Chen 0011
Distributed Parallel Databases4
2015 Zip: An Algorithm Based on Loser Tree for Common Contacts Searching in Large Graphs
Jin Huang 0007, Jia Zhu 0003, Jian Chen 0011, Rui Ding 0007
J. Comput. Sci. Technol.5
2015 A Structure Learning Algorithm for Bayesian Network Using Prior Knowledge
Jungang Xu, Jian Chen 0011, Chao Han 0002
J. Comput. Sci. Technol.3
2015 Analysis and evaluation of the top-k most influential location selection query
Jian Chen 0011, Jin Huang 0003, Zeyi Wen, Zhen He 0002, Kerry L. Taylor, Rui Zhang 0003
Knowl. Inf. Syst.1
2014 MELODY-JOIN: Efficient Earth Mover's Distance similarity joins using MapReduce
abstract
The Earth Mover's Distance (EMD) similarity join retrieves pairs of records with EMD below a given threshold. It has a number of important applications such as near duplicate image retrieval and pattern analysis in probabilistic datasets. However, the computational cost of EMD is super cubic to the number of bins in the histograms used to represent the data objects. Consequently, the EMD similarity join operation is prohibitive for large datasets. This is the first paper that specifically addresses the EMD similarity join and we propose to use MapReduce to approach this problem. The MapReduce algorithms designed for generic metric distance similarity joins are inefficient for the EMD similarity join because they involve a large number of distance computations and have unbalanced workloads on reducers when dealing with skewed datasets. We propose a novel framework, named Melody-Join, which transforms data into the space of EMD lower bounds and performs pruning and partitioning at a low cost because computing these EMD lower bounds has a constant complexity. Furthermore, we address two key problems, the limited pruning power and the unbalanced workloads, by enhancing each phase in the Melody-Join framework. We conduct extensive experiments on real datasets. The results show that Melody-Join outperforms the state-of-the-art technique by an order of magnitude, scales up better on large datasets than the state-of-the-art technique, and scales out well on distributed machines.
Jin Huang 0003, Rui Zhang 0003, Rajkumar Buyya, Jian Chen 0011
ICDE4
2013 Recommendations for two-way selections using skyline view queries
Jian Chen 0011, Jin Huang 0007, Bin Jiang 0009, Jian Pei 0001, Jian Yin 0001
Knowl. Inf. Syst.1
2013 Skyline distance: a measure of multidimensional competence
Jin Huang 0007, Bin Jiang 0009, Jian Pei 0001, Jian Chen 0011, Yong Tang 0001
Knowl. Inf. Syst.4
2012 Integrating Tags and Ratings Into User Profiling for Personalized Search in Collaborative Tagging Systems
abstract
Recently, some systems allow users to rate and annotate resources, e.g., Movie Lens, and we consider that it provides a way to identify favor tags and annoying tags of a user by integrating user's rating and tags. In this paper, we reveal and elaborate on the limitations of current works on user profiling for personalized search in collaborative tagging systems. Then we propose a new multi-level user profiling model by integrating tags and ratings to achieve personalized search, which can reflect not only the user's favor but also a user's nuisances. To the best of our knowledge, this is the first effort to integrate the ratings and tags to model multi-level user profiles for personalized search.
Yi Cai 0001, Jian Chen 0011, Yifeng Shao, Ho-fung Leung, Huaqing Min
Web Intelligence3
2011 Top-k most influential locations selection
abstract
We propose and study a new type of facility location selection query, the top-k most influential location selection query. Given a set M of customers and a set F of existing facilities, this query finds k locations from a set C of candidate locations with the largest influence values, where the influence of a candidate location c (c in C) is defined as the number of customers in M who are the reverse nearest neighbors of c. We first present a naive algorithm to process the query. However, the algorithm is computationally expensive and not scalable to large datasets. This motivates us to explore more efficient solutions. We propose two branch and bound algorithms, the Estimation Expanding Pruning (EEP) algorithm and the Bounding Influence Pruning (BIP) algorithm. These algorithms exploit various geometric properties to prune the search space, and thus achieve much better performance than that of the naive algorithm. Specifically, the EEP algorithm estimates the distances to the nearest existing facilities for the customers and the numbers of influenced customers for the candidate locations, and then gradually refines the estimation until the answer set is found, during which distance metric based pruning techniques are used to improve the refinement efficiency. BIP only estimates the numbers of influenced customers for the candidate locations. But it uses the existing facilities to limit the space for searching the influenced customers and achieve a better estimation, which results in an even more efficient algorithm. Extensive experiments conducted on both real and synthetic datasets validate the efficiency of the algorithms.
Jin Huang 0003, Zeyi Wen, Jianzhong Qi 0001, Rui Zhang 0003, Jian Chen 0011, Zhen He 0002
CIKM5
2010 Towards Progressive and Load Balancing Distributed Computation: A Case Study on Skyline Analysis
Jin Huang 0007, Jian Chen 0011, Jian Pei 0001, Jian Yin 0001
J. Comput. Sci. Technol.3
2008 Face Recognition Using Clustering Based Optimal Linear Discriminant Analysis
Wenxin Yang, Shuqin Rao, Jina Wang, Jian Yin 0001, Jian Chen 0011
ADMA5
2005 Mining Correlated Rules for Associative Classification
Jian Chen 0011, Jian Yin 0001, Jin Huang 0007
ADMA1
2005 Associative Classification in Text Categorization
Jian Chen 0011, Jian Yin 0001, Jun Zhang 0003, Jin Huang 0007
ICIC (1)1
2003 Improving end-to-end quality of services in 3G wireless networks by wireless early regulation of real-time flows
abstract
In this paper, we propose to adapt the early regulation of unresponsive flows (ERUF) to third generation wireless networks employing link layer retransmissions. Wireless channel degradations may result in backlog of downlink packets at the link layer queue, causing real-time packets with hard delivery deadlines to expire and be dropped at the receiver. We propose to regulate these congested flows by dropping expiring packets at the ingress edge node of the general packet radio service (GPRS)/universal mobile telecommunication services (UMTS) core network to release shared network resources for other flows. Based on an analysis of the characteristics of the radio link control (RLC) layer of the GPRS/UMTS network, we develop a new set of mechanisms based on active queue management to achieve this goal. We present simulation results to show that this new wireless early regulation of unresponsive flows (WERUF) scheme can significantly improve the overall end-to-end quality-of-service of all traffic flows.
Jian Chen 0011, Victor C. M. Leung
PIMRC1
2003 Applying active queue management to link layer buffers for real-time traffic over third generation wireless networks
abstract
Wireless channels have the characteristic that the link quality varies with propagation conditions. For real-time flows with hard time deadlines, link layer retransmission over the wireless network due to fluctuations in link quality may result in many packets being dropped due to deadline expiry. The expired packets waste network resources and lead to long queuing delay for subsequent packets. In this paper, we propose to use active queue management to limit the transmission queue length and hence queuing delay, thus eliminating expiration packet drops. This allows the buffer and wireless bandwidth that would otherwise be wasted by expiring packets to be released earlier for other packets. We apply this mechanism to the radio link control layer in third generation wireless networks. The effectiveness of the proposed mechanism is verified by simulations.
Jian Chen 0011, Victor C. M. Leung
WCNC1