VLDB 2026 Research / reviewers in the wild / expert
Changhong Liu
dblp:37/3900
· DBLP profile ↗
31ranked-venue papers
4as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MemoPlan: Long-Horizon Web Agents with Differential Memory Compression and Lookahead Planning
Rongde Zhu, Changhong Liu |
ICIC (14) | 3 |
| 2026 | AgentCache: Self-correcting Dynamic Memory for Training-Free Test-Time Adaptation of Vision-Language Models
Rongde Zhu, Changhong Liu |
ICIC (14) | 3 |
| 2026 | Second-Order Asymptotics for Covert Communication over MIMO AWGN ChannelsabstractAs claimed in International Telecommunications Union (ITU) Recommendation ITU-R M.2160, 6G communication systems require extremely high-security and low-latency communication, necessitating the study of covert communication operating at finite blocklengths. In the point-to-point (P2P) setting, covert communication enables a transmitter to send messages reliably over a noisy channel to a legitimate receiver without being detected by any third party. Under the covertness metric of Kullback–Leibler divergence (KLD), we derive exact second-order asymptotics for P2P covert communication over a multiple-input multiple-output (MIMO) additive white Gaussian noise (AWGN) channel. In particular, our theoretical benchmarks refine the first-order asymptotics of Wang and Bloch (TIFS 2021), which is known as the square root law by showing that the non-asymptotic maximal number of transmitted messages has a back off that scales in the order \(\Theta(n^\frac{1}{4})\) beyond the first-order term scaling in the order \(\Theta(n^\frac{1}{2})\) when the blocklength is \(n\). Thus, our second-order asymptotic bound provides a better approximation to the finite blocklength performance of optimal codes. Furthermore, compared with the single antenna result of Yu et al. (arXiv:2305.17924v3), we demonstrate the impact of the number of antennas \(m\) and reveal spatial diversity gains of MIMO, advocating the use of MIMO for covert communication to achieve a high transmission rate. To prove our results, we extend the quasi-\(\eta\)-neighborhood framework from single-antenna real value channels to multi-antenna complex value MIMO channels. To ensure covertness, the transmission power vanishes as the blocklength \(n\) increases. Thus, we judiciously analyze the finite blocklength performance of MIMO communication by modifying critical steps concerning the Berry–Esseen Theorem to deal with vanishing second and third absolute moments of information densities that rely on blocklength, which is in stark contrast with the non-covert case. Changhong Liu, Jingjing Wang 0001, Lin Zhou 0002 |
ISIT | 1 |
| 2026 | EFD-YOLO: An Improved YOLOv8 Network for River Floating Debris Object DetectionabstractWith the rapid development of unmanned aerial vehicle (UAV) technology, UAVs have provided an innovative solution for floating debris monitoring. However, object detection in UAV images remains challenging due to high miss rates for small objects, insufficient low-level feature extraction and computational redundancy. This letter proposes an Efficient Floating Debris detection model based on YOLOv8n, named EFD-YOLO, to address these issues. First, the Edge Fusion Stem (EFStem) module is proposed to enhance low-level feature extraction through an integrated gate-attention mechanism. Second, the Multi-Branch Efficient Reparameterization Block (MBERB) is designed to achieve efficient cross-layer feature fusion. Experimental results demonstrate that compared to YOLOv8n, our model achieves a 6.3% improvement in mean Average Precision (mAP) on the UAV Floating Debris Dataset, while simultaneously reducing parameters by 26.7% and improving small object recall by 21.9%. The inference time of EFD-YOLO on the RK3588 edge device is as low as 30.5 ms, demonstrating real-time capability. Yier Yan, Zhibin Liang, Changhong Liu, Tao Zou 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | DEEPSERVE: Serverless Large Language Model Serving at Scale
Zhixia Liu, Yuetao Chen, Baoquan Zhang, Shining Wan, Gengyuan Dan, Zhiyu Dong, Zhihao Ren, Changhong Liu, Tao Xie 0001, Dayun Lin, Xusheng Chen, Yizhou Shan |
USENIX ATC | 14 |
| 2025 | Bidirectional Feature Aggregation and Adaptive Fusion Network for ALS Point Cloud Semantic SegmentationabstractSignificant progress has been achieved using deep learning technology for the semantic segmentation of airborne laser scanning (ALS) point clouds. However, there are still challenges in effectively capturing contextual information and fusing network multilevel features to improve ALS point cloud segmentation performance. In this letter, we propose a lightweight bidirectional feature aggregation and adaptive fusion network. First, we propose a novel adaptive bidirectional feature aggregation module (ABA) to adaptively aggregate multiscale local features, effectively expanding the receptive field. After that, we introduce a feature adaptive fusion module (FAF) between upsampling and downsampling in the same layer to explore more comprehensive multilevel features and finer contextual information. Experiments on the ISPRS 3-D, GML, and LASDU datasets verify the superior segmentation performance of the proposed method with mF1 of 72.83% on ISPRS 3-D, 76.12% on GML, and 78.95% on LASDU. Changhong Liu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | ConMSDMamba: Multi-Scale Dilated Mamba Based on Conformer for Speech Emotion RecognitionabstractAlthough the Conformer model excels in speech processing, its core self-attention mechanism is limited in capturing multi-scale temporal dynamics and lacks explicit modeling of frequency-domain features, both crucial for Speech Emotion Recognition (SER). To address this, we propose ConMSDMamba, a novel Conformer-based architecture for SER. Specifically, to overcome the single-scale limitation of the original self-attention, we introduce a multi-scale dilated structure with parallel dilated convolutions to capture diverse temporal contexts. We further find that combining this structure with bidirectional Mamba models long-range temporal dependencies more efficiently than multi-head self-attention. Furthermore, to complement the Conformer's time-domain focus, we design a time-frequency convolution module that incorporates a wavelet-based branch for joint time-frequency perception. Experimental results on the widely used IEMOCAP and MELD datasets demonstrate that ConMSDMamba outperforms state-of-the-art methods. Guangyuan Qian, Zhenchun Lei, Sihong Liu, Changhong Liu, Aiwen Jiang |
IEEE Signal Process. Lett. | 4 |
| 2025 | Open-ended Autoregressive Visual Storytelling via Parameter Efficient Instruction TuningabstractVisual storytelling (VIST) involves generating coherent, creative, and vivid narrative for a collection of images. It remains an immense challenge within cross-modal domain. Traditional mainstream storytelling work were less proficient in handling long sequential relationships. Though large-scale visual-language pre-training (VLP) models demonstrated promising prospect on cross-modal tasks. So far they still were not particularly adept at handling tasks involving image sequences. Moreover, the reference descriptions in the available VIST benchmark dataset are short and simplistic, which constrains model’s potential capabilities. Current models struggle to produce truly rich and vivid narratives. Therefore, in this article, we will address these deficiencies, and contribute from both dataset and innovative model aspects. Firstly, by leveraging large language model (LLM), we have constructed a new dataset VIST++ which can enrich vivid narratives for open-ended image sequences. The dataset has potential on providing beneficial support for future model learning. Secondly, we have proposed an innovative auto-regressive story generation model named ReStoryGen. It can be applied to image sequences of varying lengths in an open-ended way. We have performed extensive experiments and evaluations in terms of visual grounding, coherence and non-redundancy. The experiment results have convincingly demonstrated ReStoryGen achieves impressive outcomes through utilizing parameter-efficient instruction-tuning. Related source codes and models are distributed on Github https://github.com/lixinliu1995/story_gen . Aiwen Jiang, Changhong Liu, Mingwen Wang 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2024 | Gmm-Resnext: Combining Generative and Discriminative Models for Speaker VerificationabstractWith the development of deep learning, many different network architectures have been explored in speaker verification. However, most network architectures rely on a single deep learning architecture, and hybrid networks combining different architectures have been little studied in ASV task. In this paper, we propose the GMM-ResNext model for speaker verification. Conventional GMM does not consider the score distribution of each frame feature over all Gaussian components and ignores the relationship between neighboring speech frames. So, we extract the log Gaussian probability features based on the raw acoustic features and use ResNext-based network as the backbone to extract the speaker embedding. GMM-ResNext combines generative and discriminative models to improve the generalization ability of deep learning models and allows one to more easily specify meaningful priors on model parameters. A two-path GMM-ResNext model based on two gender-related GMMs has also been proposed. The experimental results show that the proposed GMM-ResNext achieves relative improvements of 48.1% and 11.3% in EER compared with ResNet34 and ECAPA-TDNN on VoxCeleb1-O. Zhenchun Lei, Changhong Liu |
ICASSP | 3 |
| 2024 | GMM-ResNet2: Ensemble of Group Resnet Networks for Synthetic Speech DetectionabstractDeep learning models are widely used for speaker recognition and spoofing speech detection. We propose the GMM-ResNet2 for synthesis speech detection. Compared with the previous GMM-ResNet model, GMM-ResNet2 has four improvements. Firstly, the different order GMMs have different capabilities to form smooth approximations to the feature distribution, and multiple GMMs are used to extract multi-scale Log Gaussian Probability features. Secondly, the grouping technique is used to improve the classification accuracy by exposing the group cardinality while reducing both the number of parameters and the training time. The final score is obtained by ensemble of all group classifier outputs using the averaging method. Thirdly, the residual block is improved by including one activation function and one batch normalization layer. Finally, an ensemble-aware loss function is proposed to integrate the independent loss functions of all ensemble members. On the ASVspoof 2019 LA task, the GMM-ResNet2 achieves a minimum t-DCF of 0.0227 and an EER of 0.79%. On the ASVspoof 2021 LA task, the GMM-ResNet2 achieves a minimum t-DCF of 0.2362 and an EER of 2.19%, and represents a relative reductions of 31.4% and 76.3% compared with the LFCC-LCNN baseline. Zhenchun Lei, Changhong Liu, Minglei Ma |
ICASSP | 3 |
| 2024 | Music-driven Character Dance Video Generation based on Pre-trained Diffusion ModelabstractLarge-scale pre-trained models have shown significant progress in cross-modal generation tasks, especially in the text-to-image generation task. However, the pre-trained models for audio-guided video are rare. ControlNet [1] provides a new architecture to enhance the pre-trained diffusion models with task-specific conditions. Following the ControlNet [1] architecture, we propose a music-driven character dance video generation model based on the pre-trained diffusion model by taking the text prompt, music, and character image as the additional guidance conditions to generate dance videos. In this model, multimodal semantic correspondence between text, music, and video is exploited to generate character dance videos better by incorporating the pre-trained CLIP [2] and Wav2CLIP [3] models. Additionally, we design a text prompt to improve the appearance quality of the generated character images. Extensive experiments on the AIST++ dataset show the effectiveness of our method and its ability to generate character dance videos effectively. Changhong Liu, Juan Cai, Ji Ye, Zhenchun Lei, Aiwen Jiang |
IJCNN | 2 |
| 2024 | Low-Light Image Enhancement via FourierTMamba: A Hybrid Frequency-Spatial Approach
Shuwei Peng, Xu Zhang 0079, Aiwen Jiang, Changhong Liu, Jihua Ye |
MMAsia | 4 |
| 2024 | MMIDM: Generating 3D Gesture from Multimodal Inputs with Diffusion Models
Ji Ye, Changhong Liu, Haocong Wan, Aiwen Jiang, Zhenchun Lei |
PRCV (6) | 2 |
| 2024 | Automatic statistical chart analysis based on deep learning method
Changhong Liu, Shengchun Li |
Multim. Tools Appl. | 1 |
| 2024 | DAEA-Net: Dual Attention and Elevation-Aware Networks for Airborne LiDAR Point Cloud Semantic SegmentationabstractSemantic segmentation of airborne laser scanning (ALS) point clouds remains a challenging task due to the complexity and diversity of 3-D scenes in the real world. Currently, most deep learning-based airborne LiDAR point cloud segmentation methods prioritize designing local feature extraction operators while overlooking the long-range dependencies among neighborhoods and the inherently diverse properties of point cloud data. To address these issues, this article introduces a dual-attention and elevation-aware airborne LiDAR point cloud semantic segmentation network (DAEA-Net) built upon an encoding-decoding architecture. First, we develop a cross multiple anti-affine attention (CMAAA) module that effectively captures global contextual information across different neighborhoods through interactive learning of multiple features. Second, we introduce an elevation awareness (EA) module that uses normal vectors to establish a geometric similarity discriminant for each neighboring point. It incorporates an autoencoder architecture to fuse elevation information, enhancing the horizontal structural dissimilarity between objects of similar height while enriching the representation of elevation data. Additionally, to compensate for the potential information loss in the encoding-decoding hierarchical structure, we design a lightweight U-global attention (UGA) module to link decoding and encoding hierarchical levels. It merges features of different resolutions and levels during downsampling and upsampling through pooling while utilizing the self-attention mechanism to enhance the network’s global expression capability. The proposed DAEA-Net enhances ALS semantic segmentation performance by enabling interactive learning of multiple features and effectively representing elevation information. Extensive experiments conducted on two datasets demonstrate that our method delivers superior semantic segmentation performance compared to several existing advanced techniques. Yurong Zhu, Changhong Liu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Snippet-level Supervised Contrastive Learning-based Transformer for Temporal Action DetectionabstractAnchor-free temporal action detection methods have recently achieved many good results in solving the problem of flexible boundaries and different duration of actions. But the anchor-free methods use local features to predict the action boundaries so that it is sensitive to noises and prone to generate incomplete action proposals. Moreover, there exist long-term temporal dependencies between actions and temporal semantic consistency between action primitives in the same classes of actions. Therefore, we propose a snippet-level supervised contrastive learning-based transformer (SSCL-T) model for temporal action detection, which can learn semantically local and global temporal relationships in actions. This model learns the local temporal dynamic features of actions through local temporal coding and uses the transformer to model the global semantic dependencies between long-term actions. In addition, we utilize the action class information to learn the high-level semantic features of actions by designing a snippet-level supervised contrastive learning, forcing the temporal dynamic features of the same class of actions to be as close as possible and the features of different classes of actions to be as far away as possible, thus effectively realizing accurate prediction of action boundaries. Our model has been verified on two benchmark datasets ActivityNet-v1.3 and THUMOS14. The experimental results demonstrate that the proposed model has significantly improved on both datasets. Compared with the benchmark method BMN, the average mAP value has increased by 2.91% and 8.4% on ActivityNet-v1.3 and THUMOS14, respectively. Ronghai Xu, Changhong Liu, Zhenchun Lei |
IJCNN | 2 |
| 2023 | Group GMM-ResNet for Detection of Synthetic Speech Attacks
Zhenchun Lei, Yingen Yang, Changhong Liu, Minglei Ma |
INTERSPEECH | 4 |
| 2022 | Two-Path GMM-ResNet and GMM-SENet for ASV Spoofing DetectionabstractThe automatic speaker verification system is sometimes vulnerable to various spoofing attacks. The 2-class Gaussian Mixture Model classifier for genuine and spoofed speech is usually used as the baseline for spoofing detection. However, the GMM classifier does not separately consider the scores of feature frames on each Gaussian component. In addition, the GMM accumulates the scores on all frames independently, and does not consider their correlations. We propose the two-path GMM-ResNet and GMM-SENet models for spoofing detection, whose input is the Gaussian probability features based on two GMMs trained on genuine and spoofed speech respectively. The models consider not only the score distribution on GMM components, but also the relationship between adjacent frames. A two-step training scheme is applied to improve the system robustness. Experiments on the ASVspoof 2019 show that the LFCC+GMM-ResNet system can relatively reduce min-tDCF and EER by 76.1% and 76.3% on logical access scenario compared with the GMM, and the LFCC+GMM-SENet system by 94.4% and 95.4% on physical access scenario. After score fusion, the systems give the second-best results on both scenarios. Zhenchun Lei, Hui Yang 0007, Changhong Liu, Minglei Ma, Yingen Yang |
ICASSP | 3 |
| 2022 | Multi-Scale Cascaded Generator for Music-driven Dance SynthesisabstractDance is a creative performance art and must keep coherent with the rhythm and style of music. To address these issues, most of the existing music-driven dance synthesis methods utilize deep generative models and capture the dynamic characteristics of dance motions. However, we observe that dance motions contain big-scale body part movements and small-scale joint movements that are mutually coordinated and related, and the generated dance motions, particularly affected by$L_{1}$loss, are too restrictive and conservative. In this paper, we propose a multi-scale cascaded music-driven dance synthesis network (MC-MDSN) that first generates big-scale body motions conditioned on music and then further refines local small-scale joint motions. Furthermore, we design a multi-scale feature loss to capture the dynamic characteristics of each scale motions and the relations between different scale motion joints. Experimental results show that our method generates better dance motions than the baselines. Changhong Liu, Aiwen Jiang, Zhenchun Lei, Mingwen Wang 0001 |
IJCNN | 2 |
| 2022 | Multi-Path GMM-MobileNet Based on Attack Algorithms and Codecs for Synthetic Speech and Deepfake Detection
Zhenchun Lei, Yingen Yang, Changhong Liu, Minglei Ma |
INTERSPEECH | 4 |
| 2022 | Learning Hierarchical Semantic Correspondences for Cross-Modal Image-Text RetrievalabstractCross-modal image-text retrieval is a fundamental task in information retrieval. The key to this task is to address both heterogeneity and cross-modal semantic correlation between data of different modalities. Fine-grained matching methods can nicely model local semantic correlations between image and text but face two challenges. First, images may contain redundant information while text sentences often contain words without semantic meaning. Such redundancy interferes with the local matching between textual words and image regions. Furthermore, the retrieval shall consider not only low-level semantic correspondence between image regions and textual words but also a higher semantic correlation between different intra-modal relationships. We propose a multi-layer graph convolutional network with object-level, object-relational-level, and higher-level learning sub-networks. Our method learns hierarchical semantic correspondences by both local and global alignment. We further introduce a self-attention mechanism after the word embedding to weaken insignificant words in the sentence and a cross-attention mechanism to guide the learning of image features. Extensive experiments on Flickr30K and MS-COCO datasets demonstrate the effectiveness and superiority of our proposed method. Sheng Zeng, Changhong Liu, Jun Zhou 0001, Aiwen Jiang |
ICMR | 2 |
| 2022 | Music-to-Dance Generation with Multiple ConformerabstractIt is necessary for the music-to-dance generation to consider both the kinematics in dance that is highly complex and non-linear and the connection between music and dance movement that is far from deterministic. Existing approaches attempt to address the limited creativity problem, but it is still a very challenging task. First, it is a long-term sequence-to-sequence task. Second, it is noisy in the extracted motion keypoints. Last, there exist local and global dependencies in the music sequence and the dance motion sequence. To address these issues, we propose a novel autoregressive generative framework that predicts future motions based on past motions and music. This framework contains a music conformer, a motion conformer, and a cross-modal conformer, which utilizes the conformer to encode music and motion sequences, and further adapt the cross-modal conformer to the noisy dance motion data that enable it to not only capture local and global dependencies among the sequences but also reduce the effect of noisy data. Quantitative and qualitative experimental results on the publicly available music-to-dance dataset demonstrate our method improves greatly upon the baselines and can generate long-term coherent dance motions well-coordinated with the music. Mingao Zhang, Changhong Liu, Zhenchun Lei, Mingwen Wang 0001 |
ICMR | 2 |
| 2022 | Semantic-aware automatic image colorization via unpaired cycle-consistent self-supervised networkabstractAutomatic image colorization without manual interventions is an ill-conditioned and inherently ambiguous problem. Most of existing methods focus on formulating colorization as a regression problem and learn parametric mappings from grayscale to color through deep neural networks. Due to the multimodalities of color-grayscale space, in many applications, it is not required to recover exact ground-truth color. Pair-wise pixel-to-pixel learning-based algorithms lack rationality. Techniques such as color space conversion techniques are then proposed to avoid such direct pixel learning. However, the coloring results after color space conversion are blunt and unnatural. In this paper, we hold viewpoints that a reasonable solution is to generate some colorized result that looks natural. No matter what color a region is to be assigned, the colorized region should be semantically and spatially consistent. In this paper, we propose an effective semantic-aware automatic colorization model via unpaired cycle-consistent self-supervised network. Low-level monochrome loss, perceptual identity loss and high-level semantic-consistence loss, together with adversarial loss, are introduced to guide network self-training. We train and test our model on randomly selected subsets from PASCAL VOC 2012. The experimental results including human subjective studies demonstrate that, compared with state-of-the-art methods, our proposed model can achieve more convincing and superior results. Relevant source code is available at https://github.com/YuSuen/ACCycleGAN. Aiwen Jiang, Changhong Liu, Mingwen Wang 0001 |
Int. J. Intell. Syst. | 3 |
| 2022 | Beyond Triplet Loss: Person Re-Identification With Fine-Grained Difference-Aware Pairwise LossabstractPerson Re-IDentification (ReID) aims at re-identifying persons from different viewpoints across multiple cameras. Capturing the fine-grained appearance differences is often the key to accurate person ReID, because many identities can be differentiated only when looking into these fine-grained differences. However, most state-of-the-art person ReID approaches, typically driven by a triplet loss, fail to effectively learn the fine-grained features as they are focused more on differentiating large appearance differences. To address this issue, we introduce a novel pairwise loss function that enables ReID models to learn the fine-grained features by adaptively enforcing an exponential penalization on the images of small differences and a bounded penalization on the images of large differences. The proposed loss is generic and can be used as a plugin to replace the triplet loss to significantly enhance different types of state-of-the-art approaches. Experimental results on four benchmark datasets show that the proposed loss substantially outperforms a number of popular loss functions by large margins; and it also enables significantly improved data efficiency. Guansong Pang, Xiao Bai 0001, Changhong Liu, Xin Ning 0001, Lin Gu 0003, Jun Zhou 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | Taking Heuristic Based Graph Edge Partitioning One Step Ahead via OffStream Partitioning ApproachabstractIn the modern era of big data, large-scale graph computing has become challenging because of the dramatic rise in graph data size. Graph edge partitioning (GEP) is a crucial preprocessing step to distributed graph platforms, yet it is challenging to partition the large-scale graphs. GEP has shown better partition quality than the graph vertex partitioning for the graph's skewed degree distribution. Existing GEP approaches are classified into two as stream and offline. The former category assigns edges to the partitions based on the previously received edge information. It has less partitioning quality and is affected by stream order compared to the latter while supporting big graph partitioning. The latter uses complete knowledge of a graph during partitioning and hence has a better partitioning quality than the former; however, it does not support large-scale graphs. In this study, we propose a novel OffStream partitioning approach (OSPA) and hybrid graph edge partitioner OffStreamNH. OSPA leverages both the offline and stream graph partitioning approaches through stateful partitioning by introducing a state layer. This stateful partition state is recorded while offline is partitioning its input graph. It contains partial knowledge of previously partitioned data and is used by the stream partitioner. The OffStreamNH uses Neighborhood Expansion (NE) and Higher Degree Replicated First (HDRF) algorithms for the offline and online; respectively, with minor modifications of both algorithms. Experimental results show that OffStreamNH outperforms the state of the art stream partitioners in terms of replication factor, load balance and tolerates the effect of stream orders. Hancong Duan, Changhong Liu, Fantahun Gereme, Mesay Deleli |
ICDE | 3 |
| 2021 | MSNet: A novel end-to-end single image dehazing network with multiple inter-scale dense skip-connectionsabstractAbstract Dehazing is a challenging ill‐posed image restoration task. Various prior‐based and learning‐based methods have been proposed. Among them, end‐to‐end deep models achieve great success on performance improvement. However, most of them are concentrated on feature learning within the same block scale in isolation, and cannot perform associated analysis well on feature characteristics of different scales. Inter‐scale information reuse which is especially beneficial to image restoration is often neglected. Therefore, in this paper, a novel end‐to‐end network with multiple inter‐scale dense skip‐connections for image dehazing is proposed. Sufficient complementary information combination is considered through dense inter‐scale skip‐connections among encoder and decoder block layers. Besides avoiding gradient vanishing, a kind of bottleneck residual block is proposed to control the importance of local gradients at different scales over global learning process. Extensive comparisons and ablation studies on public dehazing datasets and real‐world images have been conducted. The experiment results demonstrate that the proposed novel elements can ensure more stable training process and superior testing performance with great improvements on PSNR and SSIM. Authors' haze‐removal results consistently comply satisfactorily with real situations, having much higher definition and contrast without colour distortion than those from the state‐of‐the‐art methods compared in this paper. Qiaosi Yi, Aiwen Jiang, Xiaolin Deng, Changhong Liu |
IET Image Process. | 4 |
| 2020 | Siamese Convolutional Neural Network Using Gaussian Probability Feature for Spoofing Speech Detection
Zhenchun Lei, Yingen Yang, Changhong Liu, Jihua Ye |
INTERSPEECH | 3 |
| 2019 | Single Image Colorization Via Modified CycleganabstractIn this paper, we focus on automatically colorizing single grayscale image without manual interventions. Most of existing methods tried to accurately restore unknown ground-truth colors and require paired training data for model optimization. However, the ideal restoration objective and strict training constraints limited their performance. Inspired by CycleGAN, we formulate the process of colorization as image-to-image translation and propose an effective color-CycleGAN solution. High-level semantic identity loss and low-level color loss are additionally suggested for model optimization. Our method allows using unpaired images for training and direct prediction in rgb color space, which makes training data collection much easier and more general. We train our model on randomly selected PASCAL VOC 2007 images. All ablation study on loss function and comparisons with state-of-the-art methods are performed on grayscale SUN data. The experiment results show that our improvements on training loss could achieve better content consistence and generate better reasonable colors with less artifacts. Moreover, due to the bidirectional nature of our model, our proposed method provides a by-product that gives an excellent alternative way on color image graying. Aiwen Jiang, Changhong Liu, Mingwen Wang 0001 |
ICIP | 3 |
| 2017 | Online visual tracking based on subspace representation with continuous occlusion modeling
Chunjuan Bo, Junxing Zhang, Changhong Liu |
Multim. Syst. | 3 |
| 2009 | Comparison of human face matching behavior and computational image similarity measure
Changhong Liu, Karen Lander, Xiaolan Fu |
Sci. China Ser. F Inf. Sci. | 2 |
| 2008 | Low Circle Fatigue Life Model Based on ANFIS
Changhong Liu, Xintian Liu, Hu Huang 0002, Lihui Zhao |
ICIC (3) | 1 |