VLDB 2026 Research / reviewers in the wild / expert
Chao Li 0028
dblp:66/190-28
· DBLP profile ↗
36ranked-venue papers
1as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 13 since 2021Computer networks · 9 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic ConsistencyabstractCross-modal Knowledge Distillation has demonstrated promising performance on paired modalities with strong semantic connections, referred to as Symmetric Cross-modal Knowledge Distillation (SCKD). However, implementing SCKD becomes exceedingly constrained in real-world scenarios due to the limited availability of paired modalities. To this end, we investigate a general and effective knowledge learning concept under weak semantic consistency, dubbed Asymmetric Cross-modal Knowledge Distillation (ACKD), aiming to bridge modalities with limited semantic overlap. Nevertheless, the shift from strong to weak semantic consistency improves flexibility but exacerbates challenges in knowledge transmission costs, which we rigorously verified based on optimal transport theory. To mitigate the issue, we further propose a framework, namely SemBridge, integrating a Student-Friendly Matching module and a Semantic-aware Knowledge Alignment module. The former leverages self-supervised learning to acquire semantic-based knowledge and provide personalized instruction for each student sample by dynamically selecting the relevant teacher samples. The latter seeks the optimal transport path by employing Lagrangian optimization. To facilitate the research, we curate a benchmark dataset derived from two modalities, namely Multi-Spectral (MS) and asymmetric RGB images, tailored for remote sensing scene classification. Comprehensive experiments exhibit that our framework achieves state-of-the-art performance compared with 7 existing approaches on 6 different model architectures across various datasets. Riling Wei, Kelu Yao, Chuanguang Yang, Jin Wang 0039, Zhuoyan Gao, Chao Li 0028 |
AAAI | 6 |
| 2026 | Space Computing Constellation: System Architecture, Implementations, and ChallengesabstractLow Earth Orbit (LEO) satellite constellations have experienced rapid growth in recent years, driven by their potential to deliver global, high-bandwidth Internet services with low latency. Beyond connectivity, LEO constellations also offer promising opportunities to enable in-orbit processing of space-native data to support a wide range of emerging space applications. In this context, the concept of space computing has been proposed, a paradigm that seamlessly integrates networking and computing to provide computing-as-a-service anytime and anywhere in space. However, the inherent characteristics of satellite constellations, such as dynamic network topologies, constrained system resources, and the harsh space environment, pose significant challenges in achieving this vision. This paper outlines the system architecture and the key enabling technologies for space computing, including spaceborne computers, laser communications, spaceborne router, distributed operating systems, and onboard AI. We also present the implementation of an open space computing platform, the 3-Body computing constellation, along with the in-orbit experimental results that demonstrate the advantages of multi-satellite distributed computing. Furthermore, we outline future research directions essential for advancing toward a truly interconnected, autonomous, and intelligent space computing system. Hua Wang 0011, Kelu Yao, Luqi Gong, Yichao Jin 0001, Yuan Liu 0030, Junxiao Xue, Zhiguo Wan, Chao Li 0028, Zhifeng Zhao |
IEEE Internet Things J. | 12 |
| 2025 | Forensics-Bench: A Comprehensive Forgery Detection Benchmark Suite for Large Vision Language ModelsabstractRecently, the rapid development of AIGC has significantly boosted the diversities of fake media spread in the Internet, posing unprecedented threats to social security, politics, law, and etc. To detect the ever-increasingly diverse malicious fake media in the new era of AIGC, recent studies have proposed to exploit Large Vision Language Models (LVLMs) to design robust forgery detectors due to their impressive performance on a wide range of multimodal tasks. However, it still lacks a comprehensive benchmark designed to comprehensively assess LVLMs' discerning capabilities on forgery media. To fill this gap, we present Forensics-Bench, a new forgery detection evaluation benchmark suite to assess LVLMs across massive forgery detection tasks, requiring comprehensive recognition, location and reasoning capabilities on diverse forgeries. Forensics-Bench comprises 63, 292 meticulously curated multi-choice visual questions, covering 112 unique forgery detection types from 5 perspectives: forgery semantics, forgery modalities, forgery tasks, forgery types and forgery models. We conduct thorough evaluations on 22 open-sourced LVLMs and 3 proprietary models GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet, highlighting the significant challenges of comprehensive forgery detection posed by Forensics-Bench. We anticipate that Forensics-Bench will motivate the community to advance the frontier of LVLMs, striving for all-around forgery detectors in the era of AIGC. The deliverables will be updated here. Jin Wang 0039, Chenghui Lv, Shichao Dong 0001, Kelu Yao, Chao Li 0028, Wenqi Shao |
CVPR | 7 |
| 2025 | ECG-guided individual identification via PPGabstractPhotoplethsmography (PPG)-based individual identification aiming at recognizing humans via intrinsic cardiovascular activities has raised extensive attention due to its high security and resistance to mimicry. However, this kind of technology witnesses unpromising results due to the limitation of low information density. To this end, electrocardiogram (ECG) signals have been introduced as a novel modality to enhance the density of input information. Specifically, a novel cross-modal knowledge distillation framework is implemented to propagate discriminate knowledge from ECG modality to PPG modality without incurring additional computational demands at the inference phase. Furthermore, to ensure efficient knowledge propagation, Contrastive Language–Image Pre-training (CLIP)-based knowledge alignment and cross-knowledge assessment modules are proposed respectively. Comprehensive experiments are conducted and results show our framework outperforms the baseline model with the improvement of 2.8% and 3.0% in terms of overall accuracy on seen- and unseen individual recognitions. Riling Wei, Kelu Yao, Chuanguang Yang, Chao Li 0028 |
ICASSP | 6 |
| 2025 | DPNet: Dynamic Pooling Network for Accurate and Efficient Size-Aware Tiny Object DetectionabstractIn unmanned aerial systems, especially in complex environments, accurately detecting tiny objects is crucial. Resizing images is a common strategy to improve detection accuracy, particularly for small objects. However, simply enlarging images significantly increases computational costs and the number of negative samples, severely degrading detection performance and limiting its applicability. This paper proposes a Dynamic Pooling Network (DPNet) for tiny object detection to mitigate these issues. DPNet employs a flexible down-sampling strategy by introducing a factor (df) to relax the fixed down-sampling process of the feature map to an adjustable one. Furthermore, we design a lightweight predictor to predict df for each input image, which will be used to decrease the resolution of feature map in backbone. Thus, we achieve input-aware down-sampling. We design an Adaptive Normalization Module (ANM) to make a unified detector well compatible with different dfs. At the same time, we also design a guidance loss to supervise the predictor’s training. DPNet realizes the dynamic allocation of computing resources to trade off detection accuracy and efficiency through this. Experiments on the TinyCOCO and TinyPerson datasets show that our DPNet can save over 35% and 25% GFLOPs, respectively, while maintaining comparable detection performance.The code will be made publicly available. Luqi Gong, Yikun Chen, Tianliang Yao, Chao Li 0028, Shuai Zhao 0001, Guangjie Han |
IEEE Internet Things J. | 5 |
| 2025 | Bridging HSI and LiDAR Data With Frequency-Domain Hierarchical Fusion for Enhanced ClassificationabstractRemote sensing data from hyperspectral imaging (HSI) and LiDAR provide complementary perspectives for terrain and object analysis. However, existing methods for multimodal data fusion primarily focus on spatial-domain feature alignment, often overlooking the potential of frequency-domain information to enhance classification accuracy. To bridge this gap, we introduce the Frequency-Domain Hierarchical Perception Fusion Network (FHPF-Net), a novel framework for precise classification of remote sensing images. This network leverages both spatial and frequency-domain information, and provides a new perspective for heterogeneous data integration. To extract and utilize frequency-domain features, we propose the HighLow Spectral Separation and Mining (HLSSM) module, which isolates high-frequency details such as edges and textures from low-frequency structural patterns in HSI and LiDAR data. This separation facilitates targeted feature extraction while preserving crucial contextual information. Additionally, we introduce the Hierarchical Superimposed Multi-domain Information Fusion (HSMIF) module, which employs a multi-level fusion strategy to integrate spatial and frequency-domain features, ensuring consistency and complementarity between the two data sources. Finally, we introduce a Learnable Voting Pre-label Fusion (LVPF) strategy to effectively integrate multi-branch outputs, enhancing classification performance and model robustness. The proposed FHPF-Net effectively captures diverse responses across heterogeneous data types, enabling robust classification in complex environments. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods. Luqi Gong, Yilang Li, Fanda Fan, Shuai Zhao 0001, Chao Li 0028 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Diagnosing the Compositional Knowledge of Vision Language Models from a Game-Theoretic ViewabstractCompositional reasoning capabilities are usually considered as fundamental skills to characterize human perception. Recent studies show that current Vision Language Models (VLMs) surprisingly lack sufficient knowledge with respect to such capabilities. To this end, we propose to thoroughly diagnose the composition representations encoded by VLMs, systematically revealing the potential cause for this weakness. Specifically, we propose evaluation methods from a novel game-theoretic view to assess the vulnerability of VLMs on different aspects of compositional understanding, e.g., relations and attributes. Extensive experimental results demonstrate and validate several insights to understand the incapabilities of VLMs on compositional reasoning, which provide useful and reliable guidance for future studies. The deliverables will be updated here. Jin Wang 0039, Shichao Dong 0001, Yapeng Zhu, Kelu Yao, Chao Li 0028 |
ICML | 6 |
| 2024 | Aligning Knowledge Graph with Visual Perception for Object-goal NavigationabstractObject-goal navigation is a challenging task that requires guiding an agent to specific objects based on first-person visual observations. The ability of agent to comprehend its surroundings plays a crucial role in achieving successful object finding. However, existing knowledge-graph-based navigators often rely on discrete categorical one-hot vectors and vote counting strategy to construct graph representation of the scenes, which results in misalignment with visual images. To provide more accurate and coherent scene descriptions and address this misalignment issue, we propose the Aligning Knowledge Graph with Visual Perception (AKGVP) method for object-goal navigation. Technically, our approach introduces continuous modeling of the hierarchical scene architecture and leverages visual-language pre-training to align natural language description with visual perception. The integration of a continuous knowledge graph architecture and multimodal feature alignment empowers the navigator with a remarkable zero-shot navigation capability. We extensively evaluate our method using the AI2-THOR simulator and conduct a series of experiments to demonstrate the effectiveness and efficiency of our navigator. Nuo Xu 0006, Wen Wang 0017, Zheyuan Lin, Wei Song 0008, Chunlong Zhang, Jason Gu, Chao Li 0028 |
ICRA | 9 |
| 2024 | Low-Latency Deep Learning Inference Schedule on Multi-Core MCUabstractEmerging Artificial Internet-of-Things (AIoT) services based on Microcontroller Units (MCU) heavily harness Deep Learning (DL) to improve user experiences. Such DL-assisted services depend on fast Neural Network (NN) execution for high responsiveness, demanding tiny IoT devices to minimize the NN execution latency by efficiently utilizing their underlying hardware resources. However, existing inference frameworks cannot achieve satisfied real-time performance for Multi-core MCU (MMCU), now a mainstream platform for AIoT. We mention that two improvements can be made to speed up inference: 1) select appropriate data layout for each operator in NNs, and 2) exploit the capability of MMCU by properly partitioning each operator into multiple cores.In this paper, we propose a novel idea of joint data layout and Intra-Operator Parallelism (IOP) for low-latency DL inference on MMCU. We formulate the problem based on a self-built latency predictor, which predicts the execution latency of each operator within NNs on a given MMCU. An algorithm with time complexity being ${\mathcal{O}}\left({|E|M{N^3}}\right)$ is proposed to find an optimal scheduling plan, where N and M refer to the maximum number of possible layouts and IOP strategies for an operator. Our experimental evaluation demonstrates that our scheduling plan can achieve a speedup of 1.52×−3.37× for CMSIS-NN, a state-of-the-art edge inference software stack. Besides, compared with the state-of-the-art IOP execution system, our scheduling plan achieves a speedup of approximately 1.67×. Chaonong Xu, Chao Li 0028, Weiming Kong |
IJCNN | 3 |
| 2024 | Complexity and algorithm of setting optimal location for data sink in real-time NOMA-based IIoTs
Chaonong Xu, Chao Li 0028 |
Comput. Commun. | 4 |
| 2024 | Cooperative Partial Task Offloading and Resource Allocation for IIoT Based on Decentralized Multiagent Deep Reinforcement LearningabstractEdge computing has become increasingly important to fulfill the diversified Quality-of-Service (QoS) or Quality-of-Experience (QoE) demands for Industrial Internet of Things (IIoT) applications, such as machine condition monitoring, fault diagnosis, intelligent production scheduling, and production quality control. Due to the heterogeneity of IIoT systems, it is of urgent necessity to concentrate on the cloud–edge–end cooperative partial task offloading and resource allocation (CPTORA) problem for realizing workload balancing, efficient resource utilization, and better QoS/QoE of IIoT applications. However, the challenge lies in how to make real-time, accurate, decentralized task offloading (TO) and resource allocation (RA) decisions for dynamic and device-intensive IIoT. Therefore, this work examines the CPTORA problem for IIoT, aiming at minimizing its long-run overall delay and energy costs. To lower the problem complexity, this problem is decomposed into the TO subproblem and the RA subproblem. Then, an improved soft actor–critic-based decentralized multiagent deep reinforcement learning (MADRL) algorithm is proposed to address the TO subproblem, where each IIoT device can learn its globally optimal policy and make its decisions independently. This algorithm innovatively combines the divergence regularization, the distributional reinforcement learning, and the value function decomposition methods to improve convergence speed and accuracy of the existing MADRL methods. After receiving the TO decisions of every IIoT device, every edge server employs the Lagrange multiplier method and Karush–Kuhn–Tucker condition to solve its RA subproblem. The experimental results show that the proposed algorithm decreases the overall delay and energy costs more effectively, compared to the other state-of-the-art MADRL approaches. Fan Zhang 0014, Guangjie Han, Li Liu 0022, Yu Zhang 0001, Yan Peng 0001, Chao Li 0028 |
IEEE Internet Things J. | 6 |
| 2023 | Towards Understanding the Generalization of Deepfake Detectors from a Game-Theoretical ViewabstractThis paper aims to explain the generalization of deep-fake detectors from the novel perspective of multi-order interactions among visual concepts. Specifically, we propose three hypotheses: 1. Deepfake detectors encode multi-order interactions among visual concepts, in which the low-order interactions usually have substantially negative contributions to deepfake detection. 2. Deepfake detectors with better generalization abilities tend to encode low-order interactions with fewer negative contributions. 3. Generalized deepfake detectors usually weaken the negative contributions of low-order interactions by suppressing their strength. Accordingly, we design several mathematical metrics to evaluate the effect of low-order interaction for deepfake detectors. Extensive comparative experiments are conducted, which verify the soundness of our hypotheses. Based on the analyses, we further propose a generic method, which directly reduces the toxic effects of low-order interactions to improve the generalization of deepfake detectors to some extent. Kelu Yao, Jin Wang 0039, Boyu Diao, Chao Li 0028 |
ICCV | 4 |
| 2023 | Hybrid Parallel Inference for Large Model on Heterogeneous Clusters for High ThroughputabstractIn high-throughput intelligent computing scenarios, multi-device parallelism strategies based on data parallelism or pipeline parallelism have been extensively utilized to accelerate large deep neural network model inference. Data parallelism offers nearly linear improvement in inference speed, but it is limited by the memory capacity of a single device which constrains the model size. On the other hand, pipeline parallelism can support larger models, but the total communication of the activations among devices is high, which limits the improvement of the inference speed. To address the demand for efficient model inference in high-throughput heterogeneous scenarios, we proposes a hybrid parallelism strategy that combines data parallelism and pipeline parallelism. The strategy involves grouping heterogeneous device clusters and then employing inter-group data parallelism along with intra-group pipeline parallelism. Moreover, we propose an algorithm to find an optimal hybrid parallel inference strategy with maximum throughput. The control variables of the strategy includes the number of groups, group-device assignments and model partition ratios. Our experimental evaluation demonstrates that compared to PipeEdge, a pipeline parallel inference framework for heterogeneous cluster, our strategy can achieve 1.7× −3.4× acceleration in an 8-device heterogeneous cluster without loss of accuracy. Chaonong Xu, Weiming Kong, Chao Li 0028, Luqi Gong |
ICPADS | 5 |
| 2023 | Atmospheric Phase Screen Reconstruction in SAR Interforometry Using ACGAN NetworkabstractAtmospheric phase screen (APS) is the dominant error source of InSAR. The spatial-temporal variations of APS in interferograms can lead to incorrect interpretation of phase and inaccurate extraction of surface deformation. In recent years, deep learning denoising models have been applied to study the features of APS in InSAR interferograms and extract deformation information in existing researches. However, the characteristics of APS are diverse, and it is difficult to extract this knowledge using a deep learning model based on limited interferograms that contain APS. Moreover, existing methods mainly use synthetic data to train the model and remove APS. In order to provide sufficient APS data for training deep learning networks, this paper proposes a new method for generating atmospheric phase samples. This method is based on the auxiliary classifier generative adversarial network (ACGAN) to generate more APS samples from available interferograms, fully learning the atmospheric delay errors related to terrain, atmospheric turbulence, and heavy rainfall, and generating atmospheric phase screen interferogram samples with multiple features. The technique is applied to 235 interferograms collected over Hangzhou on ascending orbit number 108 between January 12, 2020, and December 3, 2022, to generate different types of atmospheric sample screen interferogram samples. The network has shown great potential for the removal of atmospheric phase in InSAR research. Jing Wang 0057, Chao Li 0028, Chao Wang 0004, Hong Zhang 0001 |
IGARSS | 2 |
| 2023 | Underwater Pollution Tracking Based on Software-Defined Multi-Tier Edge Computing in 6G-Based Underwater Wireless NetworksabstractThe forthcoming 6G networks are expected to provide a vision of overlapping aerial-ground-underwater wireless networks. Meanwhile, the rapid development of the Internet of Underwater Things (IoUTs) brings forth many categories of Autonomous Underwater Vehicle (AUV)-assisted Underwater Wireless Networks (UWNs). In this paper, we argue that the AUV-assisted UWNs can be intelligently utilized to track underwater pollution. To perform smart underwater pollution tracking, we propose the paradigm of AUV flock-based networking system and Software-Defined Networking (SDN)-enabled AUV flock Networking System (SDN-AUVNS). We introduce the concept of Mobile Edge Computing (MEC) into the control of SDN-AUVNS and propose the upgrade of the control plane of the SDN-AUVNS to with the multi-tier edge computing ability. By the proposed system architecture, we adopt the artificial potential field theory to construct the network controlling model. And we present the underwater tracking model for SDN-AUVNS, especially for the underwater pollution equipotential line of a particular concentration. Furthermore, to provide accurate path planning for the equipotential line tracking, we utilize the linearizability mechanism to optimize and revise the control input for the SDN-AUVNS. Lastly, we give a fast united control algorithm that can intelligently schedule the SDN-AUVNS to track underwater pollution equipotential lines. In particular, we propose a smart approach with the name of ’Inverse Distance Weighting’ to optimize the detection sample of the SDN-AUVNS. Evaluation results indicate that our proposal is able to track/survey the equipotential lines within a satisfactory error. Chuan Lin 0001, Guangjie Han, Jinfang Jiang, Chao Li 0028, Syed Bilal Hussain Shah |
IEEE J. Sel. Areas Commun. | 4 |
| 2023 | Base and Meta: A New Perspective on Few-Shot SegmentationabstractDespite the progress made by few-shot segmentation (FSS) in low-data regimes, the generalization capability of most previous works could be fragile when countering hard query samples with seen-class objects. This paper proposes a fresh and powerful scheme to tackle such an intractable bias problem, dubbed base and meta (BAM). Concretely, we apply an auxiliary branch (base learner) to the conventional FSS framework (meta learner) to explicitly identify base-class objects, i.e., the regions that do not need to be segmented. Then, the coarse results output by these two learners in parallel are adaptively integrated to derive accurate segmentation predictions. Considering the sensitivity of meta learner, we further introduce adjustment factors to estimate the scene differences between support and query image pairs from both style and appearance perspectives, so as to facilitate the model ensemble forecasting. The remarkable performance gains on standard benchmarks (PASCAL-5$^{i}$, COCO-20$^{i}$, and FSS-1000) manifest the effectiveness, and surprisingly, our versatile scheme sets new state-of-the-arts even with two plain learners. Furthermore, in light of its unique nature, we also discuss several more practical but challenging extensions, including generalized FSS, 3D point cloud FSS, class-agnostic FSS, cross-domain FSS, weak-label FSS, and zero-shot segmentation. Our source code is available athttps://github.com/chunbolang/BAM. Chunbo Lang, Gong Cheng 0003, Binfei Tu, Chao Li 0028, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Retain and Recover: Delving Into Information Loss for Few-Shot SegmentationabstractBenefiting from advances in few-shot learning techniques, their application to dense prediction tasks (e.g., segmentation) has also made great strides in the past few years. However, most existing few-shot segmentation (FSS) approaches follow a similar pipeline to that of few-shot classification, where some core components are directly exploited regardless of various properties between tasks. We note that such an ill-conceived framework introduces unnecessary information loss, which is clearly unacceptable given the already very limited training sample. To this end, we delve into the typical types of information loss and provide a reasonably effective way, namely Retain And REcover (RARE). The main focus of this paper can be summarized as follows: (i) the loss of spatial information due to global pooling; (ii) the loss of boundary information due to mask interpolation; (iii) the degradation of representational power due to sample averaging. Accordingly, we propose a series of strategies to retain/recover the avoidable/unavoidable information, such as unidirectional pooling, error-prone region focusing, and adaptive integration. Extensive experiments on two popular benchmarks (i.e., PASCAL-5iand COCO-20i) demonstrate the effectiveness of our scheme, which is not restricted to a particular baseline approach. The ultimate goal of our work is to address different information loss problems within a unified framework, and it also exhibits superior performance compared to other methods with similar motivations. The source code will be made available at https://github.com/chunbolang/RARE. Chunbo Lang, Gong Cheng 0003, Binfei Tu, Chao Li 0028, Junwei Han 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Towards Robust Knowledge Graph Embedding via Multi-Task Reinforcement LearningabstractNowadays, Knowledge graphs (KGs) have been playing a pivotal role in AI-related applications. Despite the large sizes, existing KGs are far from complete and comprehensive. In order to continuously enrich KGs, automatic knowledge construction and update mechanisms are usually utilized, which inevitably bring in plenty of noise. However, most existing knowledge graph embedding (KGE) methods assume that all the triple facts in KGs are correct, and project both entities and relations into a low-dimensional space without considering noise and knowledge conflicts. This will lead to low-quality and unreliable representations of KGs. To this end, in this paper, we propose a general multi-task reinforcement learning framework, which can greatly alleviate the noisy data problem. In our framework, we exploit reinforcement learning for choosing high-quality knowledge triples while filtering out the noisy ones. Also, in order to take full advantage of the correlations among semantically similar relations, the triple selection processes of similar relations are trained in a collective way with multi-task learning. Moreover, we extend popular KGE models TransE, DistMult, ConvE and RotatE with the proposed framework. Finally, the experimental validation shows that our approach is able to enhance existing KGE models and can provide more robust representations of KGs in noisy scenarios. Zhao Zhang 0011, Fuzhen Zhuang, Hengshu Zhu, Chao Li 0028, Hui Xiong 0001, Qing He 0003, Yongjun Xu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Topic-aware Intention Network for Explainable Recommendation with Knowledge EnhancementabstractRecently, recommender systems based on knowledge graphs (KGs) have become a popular research direction. Graph neural network (GNN) is the key technology of KG-based recommendation systems. However, existing GNNs have a significant flaw: They cannot explicitly model users’ intent in recommendations. Intent plays an essential role in users’ behaviors. For example, users may first generate an intent to purchase a certain group of items and then select a specific item from the group based on their preferences. Therefore, explicitly modeling intent has a positive significance for improving recommendation performance and providing explanations for recommendations. In this article, we propose a new model called Topic-aware Intention Network (TIN) for explainable recommendations with KGs. TIN models user representations from both preference and intent views. Specifically, we design a relational attention graph neural network to selectively aggregate information in KG to learn user preferences, and we propose a knowledge-enhanced topic model to learn user intent, which is viewed as topics hidden in user behavior sequences. Finally, we obtain the user representation by fusing user preference and intent through an attention network. The experimental results show that our proposed model outperforms the state-of-the-art methods and can generate reasonable explanations for the recommendation results. Zhao Zhang 0011, Fuzhen Zhuang, Yongjun Xu 0001, Chao Li 0028 |
ACM Trans. Inf. Syst. | 5 |
| 2022 | Interpretable Generative Adversarial NetworksabstractLearning a disentangled representation is still a challenge in the field of the interpretability of generative adversarial networks (GANs). This paper proposes a generic method to modify a traditional GAN into an interpretable GAN, which ensures that filters in an intermediate layer of the generator encode disentangled localized visual concepts. Each filter in the layer is supposed to consistently generate image regions corresponding to the same visual concept when generating different images. The interpretable GAN learns to automatically discover meaningful visual concepts without any annotations of visual concepts. The interpretable GAN enables people to modify a specific visual concept on generated images by manipulating feature maps of the corresponding filters in the layer. Our method can be broadly applied to different types of GANs. Experiments have demonstrated the effectiveness of our method. Chao Li 0028, Kelu Yao, Jin Wang 0039, Boyu Diao, Yongjun Xu 0001, Quanshi Zhang |
AAAI | 1 |
| 2022 | Data Augmentation for Few-Shot Knowledge Graph Completion from Hierarchical PerspectiveabstractFew-shot knowledge graph completion (FKGC) has become a new research focus in the field of knowledge graphs in recent years, which aims to predict the missing links for relations that only have a few associative triples. Existing models attempt to solve the problem via learning entity and relation representations. However, the limited training data severely hinders the performance of existing models. To this end, we propose to solve the FKGC problem with the data augmentation technique. Specifically, we perform data augmentation from two perspectives, i.e., inter-task view and intra-task view. The former generates new tasks for FKGC, while the latter enriches the support or query set for an individual task. It is worth noting that the proposed framework can be applied to a number of existing FKGC models. Experimental evaluation on two public datasets indicates our model is capable of achieving substantial improvements over baselines. Yuanzhou Yao, Zhao Zhang 0011, Yongjun Xu 0001, Chao Li 0028 |
COLING | 4 |
| 2022 | Towards Understanding the Effect of Node Features on the Predictions of Graph Neural Networks
Zhao Zhang 0011, Boyu Diao, Yongjun Xu 0001, Chao Li 0028 |
ICANN (2) | 5 |
| 2022 | Pruning Filter via Gaussian Distribution Feature for Deep Neural Networks AccelerationabstractDeep learning has achieved impressive results in many areas, but the deployment of edge intelligent devices is still very slow. To solve this problem, we propose a novel compression and acceleration method based on data distribution characteristics for deep neural networks, namely Pruning Filter via Gaussian Distribution Feature (PFGDF). Compared with previous advanced pruning methods, PFGDF compresses the model by filters with insignificance in distribution, regardless of the contribution and sensitivity information of the convolution filter. PFGDF is significantly different from weight sparsification pruning because it does not require the special accelerated library to process the sparse weight matrix and introduces no more extra parameters. The pruning process of PFGDF is automated. Furthermore, the model compressed by PFGDF can restore the same performance as the uncompressed model. We evaluate PFGDF through extensive experiments, on CIFAR-10, PFGDF compresses the convolution filter on VGG-16 by 66.62% with more than 90% parameter reduced, while the inference time is accelerated by 83.73% on Huawei MATE 10. Jianrong Xu, Boyu Diao, Bifeng Cui, Chao Li 0028, Hailong Hong |
IJCNN | 5 |
| 2022 | Complexity of minimum uplink power scheduling with delay bound for Backbone-assisted PD-NOMA wireless networks
Chaonong Xu, Yutong Zhu, Chao Li 0028 |
Comput. Networks | 4 |
| 2022 | Intelligent Blockchain-Enabled Adaptive Collaborative Resource Scheduling in Large-Scale Industrial Internet of ThingsabstractWith the explosive growth of devices and tasks deployed in the industrial Internet of Things (IIoT), the lack of interconnection and collaboration between devices leads to poor timeliness and security in IIoT resource scheduling. This article focuses on the issue of adaptive scheduling of resources in large-scale IIoT. First, a collaborative terminal-edge IIoT architecture is designed, which introduces blockchain and AI technology to support dynamic resource scheduling in untrustworthy environments. Then, a smart contract-based multidimensional resource transaction model is developed to improve the efficiency and security of resource scheduling by establishing a credit-based consensus mechanism. Distributed transaction learning resource scheduling algorithm is further proposed to implement resource-adaptive scheduling between devices in IIoT. Extensive simulation experiments are conducted to evaluate the proposed method with respect to several performance aspects covering the scheduling decision delay, transaction generation ratio, and security. The obtained results demonstrate that the comprehensive scheduling performance of the proposed method outperforms other existing algorithms. Guangjie Han, Chao Li 0028 |
IEEE Trans. Ind. Informatics | 5 |
| 2021 | Towards Compressing Efficient Generative Adversarial Networks for Image Translation via Pruning and Distilling
Luqi Gong, Chao Li 0028, Hailong Hong, Hui Zhu 0002, Tangwen Qian, Yongjun Xu 0001 |
ICANN (2) | 2 |
| 2021 | Channel Pruning via Multi-Criteria based on Weight DependencyabstractChannel pruning has demonstrated its effectiveness in compressing ConvNets. In many related arts, the importance of an output feature map is only determined by its associated filter. However, these methods ignore a small part of weights in the next layer which disappears as the feature map is removed. They ignore the phenomenon of weight dependency. Besides, many pruning methods use only one criterion for evaluation and find a sweet spot of pruning structure and accuracy in a trial-and-error fashion, which can be time-consuming. In this paper, we proposed a channel pruning algorithm via multi-criteria based on weight dependency, CPMC, which can compress a pre-trained model directly. CPMC defines channel importance in three aspects, including its associated weight value, computational cost, and parameter quantity. According to the phenomenon of weight dependency, CPMC gets channel importance by assessing its associated filter and the corresponding partial weights in the next layer. Then CPMC uses global normalization to achieve cross-layer comparison. Finally, CPMC removes less important channels by global ranking. CPMC can compress various CNN models, including VGGNet, ResNet, and DenseNet on various image classification datasets. Extensive experiments have shown CPMC outperforms the others significantly. Yangchun Yan, Rongzuo Guo, Chao Li 0028, Yongjun Xu 0001 |
IJCNN | 3 |
| 2021 | Optimal data sink location for real-time NOMA-based Industrial IoTsabstractReal-time performance is one of the most vital metrics for applications in Industrial Internet of Things (IIoTs), and the relative geographic relationship between data sink and wireless sensors has great influence on the real-time performance. Since the locations of wireless sensors are in generally fixed in IIoTs, setting reasonable location for data sink is an efficient way for improving the real-time performance. In this paper, we investigate Non-Orthogonal Multiple Access (NOMA) based IIoTs, and consider how to minimize average access delay by setting suitable location for data sink. We formulate the problem and present an algorithm by mapping the problem into the classic minimum chain covering problem, and make the problem algorithm-tractable. Simulation results reveal that due to the full exploitation of NOMA parallelism, average access delay decreases more than 60% for some typical settings, and it can even reach 70% for the linear network topology. Chaonong Xu, Chao Li 0028 |
IPCCC | 3 |
| 2020 | Gated Convolutional Networks with Hybrid Connectivity for Image ClassificationabstractWe propose a simple yet effective method to reduce the redundancy of DenseNet by substantially decreasing the number of stacked modules by replacing the original bottleneck by our SMG module, which is augmented by local residual. Furthermore, SMG module is equipped with an efficient two-stage pipeline, which aims to DenseNet-like architectures that need to integrate all previous outputs, i.e., squeezing the incoming informative but redundant features gradually by hierarchical convolutions as a hourglass shape and then exciting it by multi-kernel depthwise convolutions, the output of which would be compact and hold more informative multi-scale features. We further develop a forget and an update gate by introducing the popular attention modules to implement the effective fusion instead of a simple addition between reused and new features. Due to the Hybrid Connectivity (nested combination of global dense and local residual) and Gated mechanisms, we called our network as the HCGNet. Experimental results on CIFAR and ImageNet datasets show that HCGNet is more prominently efficient than DenseNet, and can also significantly outperform state-of-the-art networks with less complexity. Moreover, HCGNet also shows the remarkable interpretability and robustness by network dissection and adversarial defense, respectively. On MS-COCO, HCGNet can consistently learn better features than popular backbones. Chuanguang Yang, Zhulin An, Hui Zhu 0002, Kun Zhang 0045, Kaiqiang Xu, Chao Li 0028, Yongjun Xu 0001 |
AAAI | 7 |
| 2020 | A deep neural network compression algorithm based on knowledge transfer for edge devices
Yanming Chen 0002, Chao Li 0028, Luqi Gong, Yiwen Zhang 0001, Weisong Shi |
Comput. Commun. | 2 |
| 2020 | Reliable uplink transmissions for NOMA-based Industrial Wireless Networks with guaranteed real-time performance
Chaonong Xu, Jianxiong Wu, Chao Li 0028 |
Comput. Commun. | 3 |
| 2019 | Multi-objective Pruning for CNNs Using Genetic Algorithm
Chuanguang Yang, Zhulin An, Chao Li 0028, Boyu Diao, Yongjun Xu 0001 |
ICANN (2) | 3 |
| 2016 | A Reliable Depth-Based Routing Protocol with Network Coding for Underwater Sensor NetworksabstractWith the rapid development of marine technology, underwater sensor networks (UWSNs) are gradually evolving from research to practice in recent years. Practicability and reliability are two major concerns for routing protocols in UWSNs. As localization is not necessary in depth-based routing protocol (DBR), it has an outstanding practicability than other geographic routing protocols. However, the reliability is not well ensured. In this paper, we propose an innovative depth-based routing with network coding improving routing reliability while preserving the intrinsic distributed manner of DBR and introducing little time delay and energy cost. Moreover, a simple analytical performance model where ideal MAC is assumed is proposed to derive the analytical delivery ratio for our DBR-NC and DBR protocols. This analytical model is validated by simulation results. The extensive simulation results show that the proposed DBR-NC protocol outperforms (over 15%) the state of art DBR protocols in terms of packet delivery ratio. We also show that our DBR-NC will not introduce much extra delay and energy consumptions. Boyu Diao, Yongjun Xu 0001, Qi Wang 0025, Zhao Chen 0007, Chao Li 0028, Zhulin An, Guangjie Han |
ICPADS | 5 |
| 2016 | Detection of Co-salient Objects by Looking Deep and Wide
Dingwen Zhang, Junwei Han 0001, Chao Li 0028, Jingdong Wang 0001, Xuelong Li 0001 |
Int. J. Comput. Vis. | 3 |
| 2015 | Co-saliency detection via looking deep and wideabstractWith the goal of effectively identifying common and salient objects in a group of relevant images, co-saliency detection has become essential for many applications such as video foreground extraction, surveillance, image retrieval, and image annotation. In this paper, we propose a unified co-saliency detection framework by introducing two novel insights: 1) looking deep to transfer higher-level representations by using the convolutional neural network with additional adaptive layers could better reflect the properties of the co-salient objects, especially their consistency among the image group; 2) looking wide to take advantage of the visually similar neighbors beyond a certain image group could effectively suppress the influence of the common background regions when formulating the intra-group consistency. In the proposed framework, the wide and deep information are explored for the object proposal windows extracted in each image, and the co-saliency scores are calculated by integrating the intra-image contrast and intra-group consistency via a principled Bayesian formulation. Finally the window-level co-saliency scores are converted to the superpixel-level co-saliency maps through a foreground region agreement strategy. Comprehensive experiments on two benchmark datasets have demonstrated the consistent performance gain of the proposed approach. Dingwen Zhang, Junwei Han 0001, Chao Li 0028, Jingdong Wang 0001 |
CVPR | 3 |
| 2015 | A Self-Paced Multiple-Instance Learning Framework for Co-Saliency DetectionabstractAs an interesting and emerging topic, co-saliency detection aims at simultaneously extracting common salient objects in a group of images. Traditional co-saliency detection approaches rely heavily on human knowledge for designing hand-crafted metrics to explore the intrinsic patterns underlying co-salient objects. Such strategies, however, always suffer from poor generalization capability to flexibly adapt various scenarios in real applications, especially due to their lack of insightful understanding of the biological mechanisms of human visual co-attention. To alleviate this problem, we propose a novel framework for this task, by naturally reformulating it as a multiple-instance learning (MIL) problem and further integrating it into a self-paced learning (SPL) regime. The proposed framework on one hand is capable of fitting insightful metric measurements and discovering common patterns under co-salient regions in a self-learning way by MIL, and on the other hand tends to promise the learning reliability and stability by simulating the human learning process through SPL. Experiments on benchmark datasets have demonstrated the effectiveness of the proposed framework as compared with the state-of-the-arts. Dingwen Zhang, Deyu Meng, Chao Li 0028, Lu Jiang 0004, Qian Zhao 0002, Junwei Han 0001 |
ICCV | 3 |