EDBT 2026 Demo / reviewers in the wild / expert
Hengyue Pan
dblp:163/2103
· DBLP profile ↗
22ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-2999-7401ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 13 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 7 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual cuts for efficiency: Streamlining NAS with train-time pruning of supernet and search space
Di Niu 0001, Hengyue Pan, Jingfei Jiang, Jinwei Xu |
Knowl. Based Syst. | 2 |
| 2025 | Zero-resource Hallucination Detection for Text Generation via Graph-based Contextual Knowledge Triples ModelingabstractLLMs obtain remarkable performance but suffer from hallucinations. Most research on detecting hallucination focuses on questions with short and concrete correct answers that are easy to check faithfulness. Hallucination detections for text generation with open-ended answers are more hard. Some researchers use external knowledge to detect hallucinations in generated texts, but external resources for specific scenarios are hard to access. Recent studies on detecting hallucinations in long texts without external resources conduct consistency comparison among multiple sampled outputs. To handle long texts, researchers split long texts into multiple facts and individually compare the consistency of each pair of facts. However, these methods (1) hardly achieve alignment among multiple facts; (2) overlook dependencies between multiple contextual facts. In this paper, we propose a graph-based context-aware (GCA) hallucination detection method for text generations, which aligns facts and considers the dependencies between contextual facts in consistency comparison. Particularly, to align multiple facts, we conduct a triple-oriented response segmentation to extract multiple knowledge triples. To model dependencies among contextual triples (facts), we construct contextual triples into a graph and enhance triples’ interactions via message passing and aggregating via RGCN. To avoid the omission of knowledge triples in long texts, we conduct an LLM-based reverse verification by reconstructing the knowledge triples. Experiments show that our model enhances hallucination detection and excels all baselines. Xinyue Fang, Zhen Huang 0006, Zhiliang Tian, Minghui Fang 0002, Ziyi Pan, Quntian Fang, Zhihua Wen, Hengyue Pan |
AAAI | 8 |
| 2025 | DiffRS: An Extensible Diffusion Model for Remote Sensing Image GenerationabstractRemote sensing image generation is of great value for virtual environment creation and adversarial learning for fake news detection. It could also address the learning sample shortage in the region of interest. However, most current image generation methods are limited to producing images of fixed sizes, few studies on extensible natural image generation largely focus on the stitching of random contents, lacking effective exploration of contextual information, which weakens the coherence of the extended images. To address this problem, we propose an extensible generation method for remote sensing images with the model DiffRS. This approach allows for sequential extension of arbitrary sizes by exploring the generated neighboring regions. The method is particularly suitable for scenes like remote sensing images where a generation block could cover multiple independent targets, rather than natural image tasks which may stitch across regions to form a completely target. Compared to the state-of-the-art extensible generation methods, DiffRS could improve the large scale image generation with better structure consistency, richer details and higher realism. Experiments showed that DiffRS could improve the FID score by 4.6% and 3.2% respectively in comparison with the MultiDiffusion and Mixture of Diffuser models. Xin Niu 0002, Jingfei Jiang, Hengyue Pan |
ICASSP | 4 |
| 2024 | Auto-Prox: Training-Free Vision Transformer Architecture Search via Automatic Proxy DiscoveryabstractThe substantial success of Vision Transformer (ViT) in computer vision tasks is largely attributed to the architecture design. This underscores the necessity of efficient architecture search for designing better ViTs automatically. As training-based architecture search methods are computationally intensive, there’s a growing interest in training-free methods that use zero-cost proxies to score ViTs. However, existing training-free approaches require expert knowledge to manually design specific zero-cost proxies. Moreover, these zero-cost proxies exhibit limitations to generalize across diverse domains. In this paper, we introduce Auto-Prox, an automatic proxy discovery framework, to address the problem. First, we build the ViT-Bench-101, which involves different ViT candidates and their actual performance on multiple datasets. Utilizing ViT-Bench-101, we can evaluate zero-cost proxies based on their score-accuracy correlation. Then, we represent zero-cost proxies with computation graphs and organize the zero-cost proxy search space with ViT statistics and primitive operations. To discover generic zero-cost proxies, we propose a joint correlation metric to evolve and mutate different zero-cost proxy candidates. We introduce an elitism-preserve strategy for search efficiency to achieve a better trade-off between exploitation and exploration. Based on the discovered zero-cost proxy, we conduct a ViT architecture search in a training-free manner. Extensive experiments demonstrate that our method generalizes well to different datasets and achieves state-of-the-art results both in ranking correlation and final accuracy. Codes can be found at https://github.com/lilujunai/Auto-Prox-AAAI24. Zimian Wei, Peijie Dong, Zheng Hui, Anggeng Li, Lujun Li 0001, Menglong Lu, Hengyue Pan, Dongsheng Li 0001 |
AAAI | 7 |
| 2024 | TVT: Training-Free Vision Transformer Search on Tiny Datasets
Zimian Wei, Hengyue Pan, Lujun Li 0001, Peijie Dong, Dongsheng Li 0001 |
ICPR (5) | 2 |
| 2023 | RD-NAS: Enhancing One-Shot Supernet Ranking Ability Via Ranking Distillation From Zero-Cost ProxiesabstractNeural architecture search (NAS) has made tremendous progress in the automatic design of effective neural network structures but suffers from a heavy computational burden. One-shot NAS significantly alleviates the burden through weight sharing and improves computational efficiency. Zero-shot NAS further reduces the cost by predicting the performance of the network from its initial state, which conducts no training. Both methods aim to distinguish between "good" and "bad" architectures, i.e., ranking consistency of predicted and true performance. In this paper, we propose Ranking Distillation one-shot NAS (RD-NAS) to enhance ranking consistency, which utilizes zero-cost proxies as the cheap teacher and adopts the margin ranking loss to distill the ranking knowledge. Specifically, we propose a margin subnet sampler to distill the ranking knowledge from zero-shot NAS to one-shot NAS by introducing Group distance as margin. Our evaluation of the NAS-Bench-201 and ResNet-based search space demonstrates that RD-NAS achieve 10.7% and 9.65% improvements in ranking ability, respectively. Our codes are available at https://github.com/pprp/CVPR2022-NAS-competition-Track1-3th-solution Peijie Dong, Xin Niu 0002, Lujun Li 0001, Zhiliang Tian, Xiaodong Wang 0002, Zimian Wei, Hengyue Pan, Dongsheng Li 0001 |
ICASSP | 7 |
| 2023 | Progressive Meta-Pooling Learning for Lightweight Image Classification ModelabstractPractical networks for edge devices adopt shallow depth and small convolutional kernels to save memory and computational cost, which leads to a restricted receptive field. Conventional efficient learning methods focus on lightweight convolution designs, ignoring the role of the receptive field in neural network design. In this paper, we propose the Meta-Pooling framework to make the receptive field learnable for a lightweight network, which consists of parameterized pooling-based operations. Specifically, we introduce a parameterized spatial enhancer, which is composed of pooling operations to provide versatile receptive fields for each layer of a lightweight model. Then, we present a Progressive Meta-Pooling Learning (PMPL) strategy for the parameterized spatial enhancer to acquire a suitable receptive field size. The results on the ImageNet dataset demonstrate that MobileNetV2 using Meta-Pooling achieves top1 accuracy of 74.6%, which outperforms MobileNetV2 by 2.3%. Peijie Dong, Xin Niu 0002, Zhiliang Tian, Lujun Li 0001, Xiaodong Wang 0002, Zimian Wei, Hengyue Pan, Dongsheng Li 0001 |
ICASSP | 7 |
| 2023 | DMFormer: Closing the gap Between CNN and Vision TransformersabstractVision transformers have shown excellent performance in computer vision tasks. As the computation cost of their self-attention mechanism is expensive, recent works tried to replace the self-attention mechanism in vision transformers with convolutional operations, which is more efficient with built-in inductive bias. However, these efforts either ignore multi-level features or lack dynamic prosperity, leading to sub-optimal performance. In this paper, we propose a Dynamic Multi-level Attention mechanism (DMA), which captures different patterns of input images by multiple kernel sizes and enables input-adaptive weights with a gating mechanism. Based on DMA, we present an efficient backbone network named DMFormer. DMFormer adopts the overall architecture of vision transformers, while replacing the self-attention mechanism with our proposed DMA. Extensive experimental results on ImageNet-1K and ADE20K datasets demonstrated that DMFormer achieves state-of-the-art performance, which outperforms similar-sized vision transformers(ViTs) and convolutional neural networks (CNNs). Zimian Wei, Hengyue Pan, Lujun Li 0001, Menglong Lu, Xin Niu 0002, Peijie Dong, Dongsheng Li 0001 |
ICASSP | 2 |
| 2023 | EMQ: Evolving Training-free Proxies for Automated Mixed Precision QuantizationabstractMixed-Precision Quantization (MQ) can achieve a competitive accuracy-complexity trade-off for models. Conventional training-based search methods require time-consuming candidate training to search optimized per-layer bit-width configurations in MQ. Recently, some training-free approaches have presented various MQ proxies and significantly improve search efficiency. However, the correlation between these proxies and quantization accuracy is poorly understood. To address the gap, we first build the MQ-Bench-101, which involves different bit configurations and quantization results. Then, we observe that the existing training-free proxies perform weak correlations on the MQ-Bench-101. To efficiently seek superior proxies, we develop an automatic search of proxies framework for MQ via evolving algorithms. In particular, we devise an elaborate search space involving the existing proxies and perform an evolution search to discover the best correlated MQ proxy. We proposed a diversity-prompting selection strategy and compatibility screening protocol to avoid premature convergence and improve search efficiency. In this way, our Evolving proxies for Mixed-precision Quantization (EMQ) framework allows the auto-generation of proxies without heavy tuning and expert knowledge. Extensive experiments on ImageNet with various ResNet and MobileNet families demonstrate that our EMQ obtains superior performance than state-of-the-art mixed-precision methods at a significantly reduced cost. The code will be released. Peijie Dong, Lujun Li 0001, Zimian Wei, Xin Niu 0002, Zhiliang Tian, Hengyue Pan |
ICCV | 6 |
| 2023 | MENAS: Multi-trial Evolutionary Neural Architecture Search with Lottery TicketsabstractNeural architecture search (NAS) has brought significant progress in recent image recognition tasks. Most existing NAS methods apply restricted search spaces, which limits the upper-bound performance of searched models. To address this issue, we propose a new search space named MobileNet3-MT. By reducing human-prior knowledge in omni dimensions of networks, MobileNet3-MT accommodates more potential candidates. For searching in this challenging search space, we present an efficient Multi-trial Evolution-based NAS method termed MENAS. Specifically, we accelerate the evolutionary search process by gradually pruning models in the population. Each model is trained with an early stop and replaced by its Lottery Tickets (the explored optimal pruned network). In this way, the full training pipeline of cumbersome networks is prevented and more efficient networks are automatically generated. Extensive experimental results on ImageNet-1K, CIFAR-10, and CIFAR-100 demonstrate that MENAS achieves state-of-the-art performance. Zimian Wei, Hengyue Pan, Lujun Li 0001, Peijie Dong, Xin Niu 0002, Dongsheng Li 0001 |
ICIP | 2 |
| 2023 | AutoRF: Auto Learning Receptive Fields with Spatial Pooling
Peijie Dong, Xin Niu 0002, Zimian Wei, Hengyue Pan, Dongsheng Li 0001, Zhen Huang 0006 |
MMM (2) | 4 |
| 2023 | Recent Trends in Deep Learning Based Textual Emotion Cause ExtractionabstractEmotion Cause Extraction Field (ECEF) focuses on the cause that triggers an emotion in a document and mainly includes Emotion Cause Extraction (ECE) and Emotion Cause Pair Extraction (ECPE). Traditional ECE aims to extract the cause based on a given emotion while ECPE aims to extract both the emotion and its corresponding cause. Recently, ECEF has attracted a lot of attention and most of the advances have benefited from significant developments in deep learning techniques, especially machine reading comprehension and neural-network-based information retrieval. The large pre-trained language model of BERT has also shown effectiveness in this field. Following the proposal of ECPE, the development of ECEF has accelerated. However, a comprehensive review of existing approaches and recent trends in the field is lacking. To address this issue, this survey presents a thorough review to summarise existing methods and recent key advances, illustrate the general technical architecture of traditional ECE, introduce several important variants, in particular ECPE, and provide a detailed comparison of several public datasets. Finally, the limitations of existing work and the prospects for further technological advances in ECEF are discussed. Xinxin Su, Zhen Huang 0006, Yong Dou, Hengyue Pan |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2022 | Cross-Modal Knowledge Distillation in Multi-Modal Fake News DetectionabstractSince the rapid dissemination of fake news brings a lot of negative effects on real society, automatic fake news detection has attracted increasing attention in recent years. In most circumstances, the fake news detection task is a multimodal problem that consists of textual and visual contents. Many existing methods simply integrate the textual and visual features as a shared representation but overlook their correlations, which may lead to sub-optimal results. To address this problem, we propose CMC, a two-stage fake news detection method with a novel knowledge distillation that captures Cross-Modal feature Correlations while training. In the first stage of CMC, the textual and visual networks are trained mutually in an ensemble learning paradigm. The proposed cross-modal knowledge distillation function is presented as a soft target to guide the training of a single-modal network with the correlations from the other peer. In the second stage of CMC, the two well-trained networks are fixed, and their extracted features are fed to a fusion mechanism. The fusion model is then trained to further improve the performance of multi-modal fake news detection. Extensive experiments on Weibo, PolitiFact, and GossipCop databases show that CMC outperforms the existing state-of-the-art methods by a large margin. Zimian Wei, Hengyue Pan, Linbo Qiao, Xin Niu 0002, Peijie Dong, Dongsheng Li 0001 |
ICASSP | 2 |
| 2022 | Fixed-Size Objects Encoding for Visual Relationship Detection
Hengyue Pan, Xin Niu 0002, Yixin Chen 0004, Peng Qiao, Zhen Huang 0006, Dongsheng Li 0001 |
Neural Process. Lett. | 1 |
| 2021 | Inertial Proximal Deep Learning Alternating Minimization for Efficient Neutral Network TrainingabstractIn recent years, the Deep Learning Alternating Minimization (DLAM), which is actually the alternating minimization applied to the penalty form of the deep neutral networks training, has been developed as an alternative algorithm to overcome several drawbacks of Stochastic Gradient Descent (SGD) algorithms. This work develops an improved DLAM by the well-known inertial technique, namely iPDLAM, which predicts a point by linearization of current and last iterates. To obtain further training speed, we apply a warm-up technique to the penalty parameter, that is, starting with a small initial one and increasing it in the iterations. Numerical results on real-world datasets are reported to demonstrate the efficiency of our proposed algorithm. Linbo Qiao, Tao Sun 0005, Hengyue Pan, Dongsheng Li 0001 |
ICASSP | 3 |
| 2021 | Graphcomm: A Graph Neural Network Based Method for Multi-Agent Reinforcement LearningabstractThe communication among agents is important for Multi-Agent Reinforcement Learning (MARL). In this work, we propose GraphComm, a method makes use of the relation-ships among agents for MARL communication. GraphComm takes the explicit relations (e.g., agent types), which can be provided through some knowledge background, into account to better model the relationships among agents. Besides explicit relations, GraphComm considers implicit relations, which are formed by agent interactions. GraphComm use Graph Neural Networks (GNNs) to model the relational information, and use GNNs to assist the learning of agent communication. We show that GraphComm can obtain better results than state-of-the-art methods on the challenging StarCraft II unit micromanagement tasks through extensive experimental evaluation. Yongquan Fu, Huayou Su, Hengyue Pan, Peng Qiao, Yong Dou, Cheng Wang 0003 |
ICASSP | 4 |
| 2020 | A High-Throughput LDPC Decoder Based on GPUs for 5G New RadioabstractIn this paper, we propose a GPU-based QC-LDPC decoder for 5G New Radio(NR). Different from existing LDPC decoders based on GPUs, our decoder achieves high throughput when decoding LDPC codes with high code rates. Moreover, we implement the shortening and puncturing techniques which are exploited by 5G NR. The decoding algorithm Min-Sum approximation algorithm(MSA) is optimized to implement efficient parallel decoding on the GPU. In order to save the on-chip and the off-chip bandwidth, we propose the two-level quantization scheme and implement data packing on the GPU. We also analyse the optimum thread assignment for different code rates based on our implementation. By using the optimum settings on the GPU, the decoding throughput achieves 1.38 Gbps in the case of (2080, 1760), r=5/6 on Nvidia RTX 2080Ti. Rongchun Li, Hengyue Pan, Huayou Su, Yong Dou |
ISCC | 3 |
| 2020 | Annealed gradient descent for deep learning
Hengyue Pan, Xin Niu 0002, Rongchun Li, Yong Dou |
Neurocomputing | 1 |
| 2020 | DropFilterR: A Novel Regularization Method for Learning Convolutional Neural Networks
Hengyue Pan, Xin Niu 0002, Rongchun Li, Yong Dou |
Neural Process. Lett. | 1 |
| 2017 | Learning Convolutional Neural Networks using Hybrid Orthogonal Projection and EstimationabstractConvolutional neural networks (CNNs) have yielded the excellent performance in a variety of computer vision tasks, where CNNs typically adopt a similar structure consisting of convolution layers, pooling layers and fully connected layers. In this paper, we propose to apply a novel method, namely Hybrid Orthogonal Projection and Estimation (HOPE), to CNNs in order to introduce orthogonality into the CNN structure. The HOPE model can be viewed as a hybrid model to combine feature extraction using orthogonal linear projection with mixture models. It is an effective model to extract useful information from the original high-dimension feature vectors and meanwhile filter out irrelevant noises. In this work, we present three different ways to apply the HOPE models to CNNs, i.e., \em HOPE-Input, \em single-HOPE-Block and \em multi-HOPE-Blocks. For \em HOPE-Input CNNs, a HOPE layer is directly used right after the input to de-correlate high-dimension input feature vectors. Alternatively, in \em single-HOPE-Block and \em multi-HOPE-Blocks CNNs, we consider to use HOPE layers to replace one or more blocks in the CNNs, where one block may include several convolutional layers and one pooling layer. The experimental results on CIFAR-10, CIFAR-100 and ImageNet databases have shown that the orthogonal constraints imposed by the HOPE layers can significantly improve the performance of CNNs in these image classification tasks (we have achieved one of the best performance when image augmentation has not been applied, and top 5 performance with image augmentation). Hengyue Pan |
ACML | 1 |
| 2017 | A fast method for saliency detection by back-propagating a convolutional neural network and clamping its partial outputsabstractIn this paper, we propose a fast deep learning method for object saliency detection using convolutional neural networks. In our approach, we use a gradient descent method to iteratively modify the input images based on the pixel-wise gradients to reduce a pre-defined cost function, which is defined to measure the class-specific objectness and clamp the class-irrelevant outputs to maintain image background. The pixel-wise gradients can be efficiently computed using the back-propagation algorithm. We further apply SLIC superpixels and LAB color based low level saliency features to smooth and refine the gradients. Our methods are quite computationally efficient, much faster than other state-of-the-art deep learning based saliency methods. Experimental results on two benchmark tasks, namely Pascal VOC 2012 and MSRA10k, have shown that our proposed methods can generate high-quality salience maps, at least comparable with many slow and complicated deep learning methods. Comparing with the pure low-level methods, our approach excels in handling many difficult images, which contain complex background, highly-variable salient objects, multiple objects, and/or very small salient objects. Hengyue Pan |
IJCNN | 1 |
| 2015 | Annealed Gradient Descent for Deep Learning
Hengyue Pan |
UAI | 1 |