EDBT 2026 Demo / reviewers in the wild / expert
Liqiang He
dblp:47/1699
· DBLP profile ↗
28ranked-venue papers
9as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpikeMixer: Dynamic axial mixing for spiking neural networks
Jiemin Ji, Liqiang He, Jun Li 0138 |
Neurocomputing | 2 |
| 2025 | IGDiT: Illumination-Guided Low-light Image Enhancement with Diffusion Transformer ModelsabstractIn recent years, diffusion models have demonstrated outstanding performance in low-light image enhancement (LLIE). However, their effectiveness is limited due to the insufficient consideration of the physical characteristics of illuminance information. Some methods have attempted to incorporate illuminance information, but the results remain unsatisfactory. To address this, in this paper, we propose an illuminance-guided low-light image enhancement method based on diffusion transformer model (IGDiT). This method enhances image quality by guiding the diffusion process using illuminance information, consisting of two main components: the Retinexbased Illuminance Estimator (ReIE) and the Illuminance-Guided Degradation Restorer (IGDR). ReIE is responsible for estimating the illuminance information, while IGDR uses this information to guide the diffusion model in restoring the image. IGDR as a diffusion Transformer model in this paper, is composed of Multi-Head Cross-Attention (MHCA), Illuminance-Guided Self-Attention (IGSA), and Multi-Scale Cross-Gated Feed-forward Network (MCFN). These modules work together to leverage illuminance information to learn the posterior distribution, thereby simulate the complex distribution of real lighting scenarios, significantly improving the image enhancement capability under low-light conditions. Extensive experiments are conducted, of which the results show that our proposed method outperforms the state-of-the-art methods. Liqiang He |
ICME | 3 |
| 2025 | A Spatial-Frequency Domain Joint Mechanism Network for Cross-modal Semantic SegmentationabstractRGB-Thermal (RGB-T) semantic segmentation has emerged as a promising approach. However, the existing methods ignore the problem of frequency inconsistencies formed by the differences between different modalities. Additionally, traditional global feature construction methods incur significant computational costs to ensure effectiveness. To address these challenges, we propose a spatial-frequency domain joint mechanism network. The Frequency Dilated Hybrid Attention module analyzes the frequency band differences in different modalities, while capturing complex details and high-level semantic information. The Cross-frequency Domain Attention module leverages the frequency domain to estimate scaled dot-product attention, enabling efficient construction of global features. We conducted extensive experiments on the MFNET and PST900 datasets, demonstrating state-of-the-art performance. Yiheng Qu, Zhibing Zhang, Liqiang He |
ICME | 3 |
| 2025 | Expediting the discovery of promising photothermal cyanine molecules through a transfer learning approachabstractCyanine-based molecules have gained significant attention in photothermal therapy due to their unique fluorescence brightness and tunable spectral properties. However, the development of new photothermal agents is often constrained by the complexity of the chemical landscape and the need for biocompatibility. To address these challenges, we present an innovative transfer learning approach for rapidly identifying promising photothermal agent candidates with excellent photothermal properties, high synthetic feasibility, and superior biocompatibility. Using natural language processing, our pretrained model generated a molecular library based on cyanine scaffolds. The most promising candidates were screened rigorously through a weighted analysis of chemical indicators, such as photothermal performance and synthetic accessibility and biological indicators, including bio-toxicity. From these, three molecules were selected for retrosynthetic analysis. This artificial intelligence-driven approach provides a robust solution to the traditional challenges in photothermal agent design, significantly enhancing their potential applications in cancer bioimaging, mitochondrial phototherapy, and image-guided surgery. Siwei Wu, Liqiang He, Guining Cao, Jiacheng Tang, Zhenxing Pan, Zihui Huang, Andi Li, Shuting Cai, Xujie Liu |
Briefings Bioinform. | 3 |
| 2024 | Generative Pre-trained Speech Language Model with Efficient Hierarchical TransformerabstractWhile recent advancements in speech language models have achieved significant progress, they face remarkable challenges in modeling the long acoustic sequences of neural audio codecs.In this paper, we introduce Generative Pretrained Speech Transformer (GPST), a hierarchical transformer designed for efficient speech language modeling.GPST quantizes audio waveforms into two distinct types of discrete speech representations and integrates them within a hierarchical transformer architecture, allowing for a unified one-stage generation process and enhancing Hi-Res audio generation capabilities.By training on large corpora of speeches in an end-to-end unsupervised manner, GPST can generate syntactically consistent speech with diverse speaker identities.Given a brief 3-second prompt, GPST can produce natural and coherent personalized speech, demonstrating in-context learning abilities.Moreover, our approach can be easily extended to spoken cross-lingual speech generation by incorporating multi-lingual semantic tokens and universal acoustic tokens.Experimental results indicate that GPST significantly outperforms the existing speech language models in terms of word error rate, speech quality, and speaker similarity.See https://youngsheen.github. io/GPST/demo for demo samples. Yongxin Zhu 0003, Dan Su 0002, Liqiang He, Linli Xu 0002, Dong Yu 0001 |
ACL (1) | 3 |
| 2024 | Attention Decomposition for Cross-Domain Semantic Segmentation
Liqiang He, Sinisa Todorovic |
ECCV (14) | 1 |
| 2024 | MECNet: Multi-Scale Exposure-Consistency Learning via Fourier Transform for Exposure CorrectionabstractIn the real world, due to various challenging lighting conditions such as low light, underexposure, and overexposure, captured images often exhibit undesirable appearances. Given that images with different exposure levels require different correction processes, a single neural network struggles to produce satisfactory results. We propose a coarse-to-fine exposure correction model for learning exposure consistency representation to address underexposure and overexposure issues. Building upon the bilateral activation mechanism, we introduce the Fourier transform to capture global information and fuse it with locally extracted information through convolution to achieve superior feature representation. Additionally, we employ Laplacian pyramids to decompose the source image into different spatial frequency bands, then the image details are enhanced by denoising high-frequency layers. Experimental results on the MSEC and SICE datasets demonstrate the superiority of our proposed method over current state-of-the-art approaches. Our code will be made available on GitHub. Gaofei Qiao, Liqiang He |
SMC | 3 |
| 2024 | EFD-MVSNet: Enhanced Feature Distinctiveness for Multi-View Stereo
Chengkun Wang, Liqiang He |
SMC | 3 |
| 2023 | Bidirectional Alignment for Domain Adaptive Detection with TransformersabstractWe propose a Bidirectional Alignment for domain adaptive Detection with Transformers (BiADT) to improve cross domain object detection performance. Existing adversarial learning based methods use gradient reverse layer (GRL) to reduce the domain gap between the source and target domains in feature representations. Since different image parts and objects may exhibit various degrees of domain-specific characteristics, directly applying GRL on a global image or object representation may not be suitable. Our proposed BiADT explicitly estimates token-wise domain-invariant and domain-specific features in the image and object token sequences. BiADT has a novel deformable attention and self-attention, aimed at bi-directional domain alignment and mutual information minimization. These two objectives reduce the domain gap in domain-invariant representations, and simultaneously increase the distinctiveness of domain-specific features. Our experiments show that BiADT achieves very competitive performance to SOTA consistently on Cityscapes-to-FoggyCityscapes, Sim10K-to-Citiscapes and Cityscapes-to-BDD100K, outperforming the strong baseline, AQT, by 2.0, 2.1, and 2.4 in mAP50, respectively. The implementation is available at https://github.com/helq2612/biADT Liqiang He, Albert Chen 0001, Min Sun 0001, Cheng-Hao Kuo, Sinisa Todorovic |
ICCV | 1 |
| 2023 | Efficient Rate Control in Versatile Video Coding With Adaptive Spatial-Temporal Bit Allocation and Parameter UpdatingabstractDespite the fact that Versatile Video Coding (VVC) has achieved superior coding performance, two major problems remain for the rate control (RC) model in VVC. First, the regions concerned by human eyes are not clear enough in the coded video due to the deviation between the target bit allocation strategy of the coding tree unit (CTU) in RC and the human visual attention mechanism (HVAM). Second, there are significant quality fluctuations in the coded video frames due to the inappropriate updating speed. To address the above problems, we propose an efficient rate control (ERC) model. Specifically, in order to make the coded video more consistent with the attention of human eyes, we extract texture and motion-based spatial-temporal information to guide the bit allocation at the CTU level. Furthermore, based on the quasi-Newton algorithm and bit error, we propose an adaptive parameter updating (APU) method with the proper updating speed to precisely control the bits per frame. The proposed ERC outperforms the default RC model of VVC Test Model (VTM) 9.1 by saving the average Bjøntegaard Delta Rate (BD-Rate) on full-frame video sequences by 3.60% and 4.94% under low delay P (LDP) and random access (RA) configurations respectively, with higher bitrate accuracy. Moreover, the Peak Signal-to-Noise Ratio (PSNR) and actual coded bits per frame in the video coded by the proposed ERC are more stable. Liqiang He, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | DESTR: Object Detection with Split TransformerabstractSelf- and cross-attention in Transformers provide for high model capacity, making them viable models for object detection. However, Transformers still lag in performance behind CNN-based detectors. This is, we believe, because: (a) Cross-attention is used for both classification and bounding-box regression tasks; (b) Transformer's decoder poorly initializes content queries; and (c) Self-attention poorly accounts for certain prior knowledge which could help improve inductive bias. These limitations are addressed with the corresponding three contributions. First, we propose a new Detection Split Transformer (DESTR) that separates estimation of cross-attention into two independent branches — one tailored for classification and the other for box regression. Second, we use a mini-detector to initialize the content queries in the decoder with classification and regression embeddings of the respective heads in the mini-detector. Third, we augment self-attention in the decoder to additionally account for pairs of adjacent object queries. Our experiments on the MS-COCO dataset show that DESTR outperforms DETR and its successors. Liqiang He, Sinisa Todorovic |
CVPR | 1 |
| 2022 | DP-DWA: Dual-Path Dynamic Weight Attention Network With Streaming Dfsmn-San For Automatic Speech RecognitionabstractIn multi-channel far-field automatic speech recognition (ASR) scenarios, distortion is introduced when the speech signal is processed by the front end, which damages the recognition performance for the ASR tasks. In this paper, we propose a dual-path network for the far-field acoustic model, which uses voice processing (VP) signal and acoustic echo cancellation (AEC) signal as input. Specifically, we design a dynamic weight attention (DWA) module for combining two signals. Besides, we streamline our best deep feed-forward sequential memory network with self-attention (DFSMN-SAN) acoustic model for real-time requirements. Joint-training strategy is adopted to optimize the proposed approach. We find that with dual-path network, we can achieve a 54.5% relative improvement in character error rate (CER) on a 10,000-hour online conference task. In addition, our proposed method is not affected by the arrangement of different microphone arrays. We achieve a 23.56% relative improvement on a vehicle task, which has an array with two microphones. Dongpeng Ma, Liqiang He, Mingjie Jin, Dan Su 0002, Dong Yu 0001 |
ICASSP | 3 |
| 2022 | Efficient Text Analysis with Pre-Trained Neural Network ModelsabstractThis paper investigates the application of pre-trained BERT model in three classic text analysis tasks: Chinese grapheme-to-phoneme(G2P), text normalization(TN) and sentence punctuation annotation. Even though the full-sized BERT has prominent modeling power, there are two challenges for it in real applications: the requirement for annotated training data and the considerable computational cost. In this paper, we propose BERT-based low-latency solutions. To collect sufficient training corpus for G2P, we transfer knowledge from existing rule-based system to BERT through a large amount of unlabeled corpus. The new model could convert all characters directly from raw texts with higher accuracy. We also propose a hybrid two-stage text normalization pipeline which reduces the sentence error rate by 25% compared to the rule-based system. We offer both supervised and weakly supervised versions and find that the latter has only 1% accuracy drop from the former. Jia Cui, Heng Lu 0004, Shiyin Kang, Liqiang He, Guangzhi Li, Dong Yu 0001 |
SLT | 5 |
| 2022 | A polar-edge context-aware (PECA) network for mirror segmentation
Liqiang He, Jiajia Luo, Ke Zhang 0028, Yuyin Sun, Nan Qiao 0009, Cheng-Hao Kuo, Sinisa Todorovic |
Image Vis. Comput. | 1 |
| 2022 | An Optimized Rate Control Algorithm in Versatile Video Coding for 360$^\circ$ VideosabstractToday, 360$°$video has become an integral part of people's lives. Despite the fact that the latest generation standard Versatile Video Coding (VVC) demonstrates a significant gain in encoding capacity over High Efficiency Video Coding (HEVC), it still has room for 360$°$video encoding improvements. To further enhance the applicability of 360$°$video coding, an optimized rate control (RC) algorithm in VVC for 360$°$video is proposed in this paper. We present an efficient extraction algorithm for obtaining the video's saliency feature. Furthermore, for the characteristics of 360$°$video, a partitioning algorithm is also proposed to divide a frame into demand and non-demand regions. Additionally, to achieve precise and rational RC, a Coding Tree Unit (CTU)-level bit allocation strategy is proposed based on the saliency feature for the above-mentioned regions. The experimental results show that the proposed RC algorithm can achieve 11.77$\%$bitrate savings and more accurate allocation compared with the default algorithm of VVC. Also, performance enhancement has been observed in comparison to the most advanced algorithm. Zeming Zhao, Xiaohai He, Shuhua Xiong, Liqiang He, Ray E. Sheriff |
IEEE Signal Process. Lett. | 4 |
| 2021 | Latency-Controlled Neural Architecture Search for Streaming Speech RecognitionabstractNeural architecture search (NAS) has attracted much attention and has been explored for automatic speech recognition (ASR). In this work, we focus on streaming ASR scenarios and propose the latency-controlled NAS for acoustic modeling. First, based on the vanilla neural architecture, normal cells are altered to causal cells to control the total latency of the architecture. Second, a revised operation space with a smaller receptive field is proposed to generate the final architecture with low latency. Extensive experiments show that: 1) Based on the proposed neural architecture, the neural networks with a medium latency of 550ms (millisecond) and a low latency of 190ms can be learned in the vanilla and revised operation space respectively. 2) For the low latency setting, the evaluation network can achieve more than 19% (average on the four test sets) relative improvements compared with the hybrid CLDNN baseline, on a 10k-hour large-scale dataset. Liqiang He, Shulin Feng, Dan Su 0002, Dong Yu 0001 |
ASRU | 1 |
| 2021 | Learned Transferable Architectures Can Surpass Hand-Designed Architectures for Large Scale Speech RecognitionabstractIn this paper, we explore the neural architecture search (NAS) for automatic speech recognition (ASR) systems. We conduct the architecture search on the small proxy dataset, and then evaluate the network, constructed from the searched architecture, on the large dataset. Specially, we propose a revised search space that theoretically facilitates the search algorithm to explore the architectures with low complexity. Extensive experiments show that: (i) the architecture learned in the revised search space can greatly reduce the computational overhead and GPU memory usage with mild performance degradation. (ii) the searched architecture can achieve more than 15% (average on the four test sets) relative improvements on the large dataset, compared with our best hand-designed DFSMN-SAN architecture. To the best of our knowledge, this is the first report of NAS results with a large scale dataset (up to 10K hours), indicating the promising application of NAS to industrial ASR systems. Liqiang He, Dan Su 0002, Dong Yu 0001 |
ICASSP | 1 |
| 2021 | Energy Balance and Cache Optimization Routing Algorithm Based on Communication WillingnessabstractExisting opportunistic network routing algorithms usually have two main problems: excessive calculation of key nodes leads to the uneven energy consumption of nodes, and limited remaining cache of nodes leads to the loss of important messages. To solve the above problems, this paper proposed a new opportunistic network routing algorithm-EC-CW, which forwards messages according to the multi-copy mechanism and the communication willingness between nodes. The simulation results show that EC-CW reduces the average latency and the overhead rate in the nodes-sparse opportunistic network scenarios composed of high-cache nodes; EC-CW improves the delivery rate and reduces the overhead rate in the nodes-intensive opportunistic network scenarios composed of low-cache nodes. JingJian Chen, Xiaorui Wu, Fengqi Wei, Liqiang He |
WCNC | 5 |
| 2021 | Ferry Node Identification Model for the Security of Mobile Ad Hoc NetworkabstractAn opportunistic network is a special type of wireless mobile ad hoc network that does not require any infrastructure, does not have stable links between nodes, and relies on node encounters to complete data forwarding. The unbalanced energy consumption of ferry nodes in an opportunistic network leads to a sharp decline in network performance. Therefore, identifying the ferry node group plays an important role in improving the performance of the opportunistic network and extending its life. Existing research studies have been unable to accurately identify ferry node clusters in opportunistic networks. In order to solve this problem, the concepts of k-core and structural holes have been combined, and a new evaluation indicator, namely, ferry importance rank, has been proposed in this study for analyzing the dynamic importance of nodes in a network. Based on this, a ferry cluster identification model has been designed for accurately identifying the ferry node clusters. The results of the simulations conducted for verifying the performance of the proposed model show that the accuracy of the model to identify the ferry node clusters is 100%. Zhihan Qi, Fengqi Wei, Liqiang He |
Secur. Commun. Networks | 6 |
| 2021 | A Routing Algorithm for the Sparse Opportunistic Networks Based on Node IntimacyabstractOpportunistic networks are becoming more and more important in the Internet of Things. The opportunistic network routing algorithm is a very important algorithm, especially based on the historical encounters of the nodes. Such an algorithm can improve message delivery quality in scenarios where nodes meet regularly. At present, many kinds of opportunistic network routing algorithms based on historical message have been provided. According to the encounter information of the nodes in the last time slice, the routing algorithms predict probability that nodes will meet in the subsequent time slice. However, if opportunistic network is constructed in remote rural and pastoral areas with few nodes, there are few encounters in the network. Then, due to the inability to obtain sufficient encounter information, the existing routing algorithms cannot accurately predict whether there are encounters between nodes in subsequent time slices. For the purpose of improving the accuracy in the environment of sparse opportunistic networks, a prediction model based on nodes intimacy is proposed. And opportunistic network routing algorithm is designed. The experimental results show that the ONBTM model effectively improves the delivery quality of messages in sparse opportunistic networks and reduces network resources consumed during message delivery. Liqiang He |
Wirel. Commun. Mob. Comput. | 6 |
| 2017 | Species Distribution Modeling of Citizen Science Data as a Classification Problem with Class-Conditional NoiseabstractSpecies distribution models relate the geographic occurrence pattern of a species to environmental features and are used for a variety of scientific and management purposes. One source of data for building species distribution models is citizen science, in which volunteers report locations where they observed (or did not observe) sets of species. Since volunteers have variable levels of expertise, citizen science data may contain both false positives and false negatives in the location labels (present vs. absent) they provide, but many common modeling approaches for this task do not address these sources of noise explicitly. In this paper, we propose to formulate the species distribution modeling task as a classification problem with class-conditional noise. Our approach builds on other applications of class-conditional noise models to crowdsourced data, but we focus on leveraging features of the noise processes that are distinct from the class features. We describe the conditions under which the parameters of our proposed model are identifiable and apply it to simulated data and data from the eBird citizen science project. Rebecca A. Hutchinson, Liqiang He, Sarah C. Emerson |
AAAI | 2 |
| 2017 | Distribution of primary additional errors in fractal encoding method
Shuai Liu 0002, Weina Fu, Liqiang He, Jiantao Zhou 0002, Ming Ma 0006 |
Multim. Tools Appl. | 3 |
| 2015 | Neuromorphic accelerators: a comparison between neuroscience and machine-learning approachesabstractA vast array of devices, ranging from industrial robots to self-driven cars or smartphones, require increasingly sophisticated processing of real-world input data (image, voice, radio, ...). Interestingly, hardware neural network accelerators are emerging again as attractive candidate architectures for such tasks. The neural network algorithms considered come from two, largely separate, domains: machine-learning and neuroscience. These neural networks have very different characteristics, so it is unclear which approach should be favored for hardware implementation. Yet, few studies compare them from a hardware perspective. We implement both types of networks down to the layout, and we compare the relative merit of each approach in terms of energy, speed, area cost, accuracy and functionality. Zidong Du, Daniel Ben Dayan Rubin, Yunji Chen, Liqiang He, Tianshi Chen 0002, Lei Zhang 0008, Chengyong Wu, Olivier Temam |
MICRO | 4 |
| 2014 | The improbable but highly appropriate marriage of 3D stacking and neuromorphic acceleratorsabstract3D stacking is a promising technology (low latency/power/area, high bandwidth); its main shortcoming is increased power density. Simultaneously, motivated by energy constraints, architectures are evolving towards greater customization, with tasks delegated to accelerators. Due to the widespread use of machine-learning algorithms and the re-emergence of neural networks (NNs) as the preferred such algorithms, NN accelerators are receiving increased attention. They turn out to be well matched to 3D stacking: inherently 3D structures with a low power density and high across-layer bandwidth requirements. We present what is, to the best of our knowledge, the first 3D stacked NN accelerator Bilel Belhadj, Alexandre Valentian, Pascal Vivet, Marc Duranton, Liqiang He, Olivier Temam |
CASES | 5 |
| 2014 | DaDianNao: A Machine-Learning SupercomputerabstractMany companies are deploying services, either for consumers or industry, which are largely based on machine-learning algorithms for sophisticated processing of large amounts of data. The state-of-the-art and most popular such machine-learning algorithms are Convolutional and Deep Neural Networks (CNNs and DNNs), which are known to be both computationally and memory intensive. A number of neural network accelerators have been recently proposed which can offer high computational capacity/area ratio, but which remain hampered by memory accesses. However, unlike the memory wall faced by processors on general-purpose workloads, the CNNs and DNNs memory footprint, while large, is not beyond the capability of the on chip storage of a multi-chip system. This property, combined with the CNN/DNN algorithmic characteristics, can lead to high internal bandwidth and low external communications, which can in turn enable high-degree parallelism at a reasonable area cost. In this article, we introduce a custom multi-chip machine-learning architecture along those lines. We show that, on a subset of the largest known neural network layers, it is possible to achieve a speedup of 450.65x over a GPU, and reduce the energy by 150.31x on average for a 64-chip system. We implement the node down to the place and route at 28nm, containing a combination of custom storage and computational units, with industry-grade interconnects. Yunji Chen, Shaoli Liu, Shijin Zhang, Liqiang He, Ling Li 0001, Tianshi Chen 0002, Zhiwei Xu 0002, Ninghui Sun, Olivier Temam |
MICRO | 5 |
| 2012 | Improve the Implementation of Pitch Features for Mandarin Digit String Recognition Task
Pei Ding, Liqiang He |
INTERSPEECH | 2 |
| 2009 | A Fast Scheme to Investigate Thermal-Aware Scheduling Policy for Multicore Processors
Liqiang He, Cha Narisu |
APPT | 1 |
| 2005 | Improving Accuracy of Perceptron Predictor Through Correlating Data Values in SMT Processors
Liqiang He |
ISNN (3) | 1 |