EDBT 2026 Demo / reviewers in the wild / expert
Zhihong Zhang 0001
dblp:07/5980-1
· DBLP profile ↗
64ranked-venue papers
21as first author
24since 2021 · last 2026
0000-0002-0542-0640ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 15 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 9 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the Topology-Semantic Gap: A Benchmark and Framework for Power Grid Work Ticket GenerationabstractRetrieval-Augmented Generation (RAG) is a promising paradigm for domain-specific knowledge integration. However, applying generic RAG frameworks to power grid operation and maintenance (O&M) faces critical challenges due to the strict physical constraints of electrical networks. We identify a fundamental Topological Association Modeling Deficiency in existing methods, manifesting as Retrieval Linkage Degradation (where semantic ambiguity severs links between abstract intents and specific technical entities) and Topological Reasoning Deficiency (failing to adhere to rigid physical interlocking rules, causing fatal safety violations). To address these bottlenecks, we introduce the Power Grid Work Ticket (PGWT) dataset to systematically evaluate topological reasoning under real-world noise and sparsity. Furthermore, we propose Topology-Semantic Aligned Retrieval (TSA-Retrieval). This framework harmonizes semantic spaces with topological structures by fusing symbolic pattern matching with subgraph structural encoding to ensure physical validity, while explicitly injecting topology-aware reasoning chains to guide constraint-compliant generation. Extensive experiments demonstrate that TSA-Retrieval significantly outperforms state-of-the-art LLMs and generic RAG baselines, offering a robust and scalable solution for safety-critical applications. Jiangbing Mao, Tianke Xiang, Yantong Zhu, Qinggang Zhang, Zhihong Zhang 0001 |
SIGIR | 6 |
| 2026 | Facilitating Generative Retrieval with Logical Denoising for Interpretable Conversational Search
Qichuan Liu, Chentao Zhang, Chenfeng Zheng, Qinggang Zhang, Zhihong Zhang 0001 |
WWW | 6 |
| 2025 | Optimize Battery Control: A Multi-Objective Evolutionary Ensemble Reinforcement Learning ApproachabstractThe Dynamically Reconfigurable Battery (DRB) systems, which use high-speed power electronic switches to dynamically adjust battery interconnections in real-time, are critical to the performance of the battery pack. Traditional battery management strategies often fail to address multi-objective optimization, leading to imbalanced performance and inadequate energy utilization. To enhance decision-making across multiple objectives, an Evolutionary Ensemble Reinforcement Learning (EERL) framework is proposed in this paper. This framework incorporates evolutionary algorithms to associate ensemble learning, thus improving reinforcement learning (RL) performance. It decomposes a complex objective into multiple sub-objectives, each optimized independently, while incorporating diverse performance metrics into the correlation stage to derive the Pareto optimal solution. The EERL can efficiently mitigate potential adverse effects such as short circuits, disconnections, and reverse charging, thereby effectively reducing capacity differences among various batteries. Simulations and real-world testing demonstrate that the proposed approach overcomes the issue of local optima entrapment in multi-objective optimization scenarios. In a real-world system, an 11.08 % increase in energy efficiency is observed compared to existing approaches. Junchi Yan, Zhihong Zhang 0001 |
IJCAI | 6 |
| 2025 | Multi-view short-term photovoltaic power prediction combining satellite images feature learning and graph mutual information feature representationabstractAbstract With the introduction of national policies, photovoltaic (PV) power forecasting requirements for PV power plants are becoming increasingly stringent. It is particularly critical that PV power predictions are accurate while new energy is being consumed. It is also important to consider the satellite imagery of the location of the PV power plant and the meteorological information of the plant itself. The authors aim to explore the impact of these two elements on PV power prediction to better support PV power prediction. Therefore, this paper explores the cloud information elements of the satellite images from a multi‐view perspective and performs feature extraction and processing of the meteorological information to learn the impact of cloud cover on PV power prediction. Meanwhile, this paper introduces the mutual information mechanism for the influence of meteorological factors on PV power generation. It constructs the mutual information matrix and adopts the graph neural network for representation learning. A time‐series prediction model for short‐term PV power prediction is constructed and more accurate prediction results are obtained. The experimental results demonstrate that the proposed method is effective, has generalisation ability, and improved performance compared with the traditional model. The proposed method can also provide a novel approach and solution for short‐term PV power prediction. Yuxing Dai, Jing Lai, Xuexin Xu, Jianbing Xiahou, Jie Lian 0005, Zhihong Zhang 0001 |
IET Comput. Vis. | 6 |
| 2024 | Two-Level Graph Neural NetworkabstractGraph neural networks (GNNs) are recently proposed neural network structures for the processing of graph-structured data. Due to their employed neighbor aggregation strategy, existing GNNs focus on capturing node-level information and neglect high-level information. Existing GNNs, therefore, suffer from representational limitations caused by the local permutation invariance (LPI) problem. To overcome these limitations and enrich the features captured by GNNs, we propose a novel GNN framework, referred to as the two-level GNN (TL-GNN). This merges subgraph-level information with node-level information. Moreover, we provide a mathematical analysis of the LPI problem, which demonstrates that subgraph-level information is beneficial to overcoming the problems associated with LPI. A subgraph counting method based on the dynamic programming algorithm is also proposed, and this has the time complexity of O(n³), where n is the number of nodes of a graph. Experiments show that TL-GNN outperforms existing GNNs and achieves state-of-the-art performance. Xing Ai, Zhihong Zhang 0001, Edwin R. Hancock |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | ESSEN: Improving Evolution State Estimation for Temporal Networks using Von Neumann EntropyabstractTemporal networks are widely used as abstract graph representations for real-world dynamic systems. Indeed, recognizing the network evolution states is crucial in understanding and analyzing temporal networks. For instance, social networks will generate the clustering and formation of tightly-knit groups or communities over time, relying on the triadic closure theory. However, the existing methods often struggle to account for the time-varying nature of these network structures, hindering their performance when applied to networks with complex evolution states. To mitigate this problem, we propose a novel framework called ESSEN, an Evolution StateS awarE Network, to measure temporal network evolution using von Neumann entropy and thermodynamic temperature. The developed framework utilizes a von Neumann entropy aware attention mechanism and network evolution state contrastive learning in the graph encoding. In addition, it employs a unique decoder the so-called Mixture of Thermodynamic Experts (MoTE) for decoding. ESSEN extracts local and global network evolution information using thermodynamic features and adaptively recognizes the network evolution states. Moreover, the proposed method is evaluated on link prediction tasks under both transductive and inductive settings, with the corresponding results demonstrating its effectiveness compared to various state-of-the-art baselines. Qiyao Huang, Zhihong Zhang 0001, Edwin R. Hancock |
NeurIPS | 3 |
| 2023 | Image feature learning combined with attention-based spectral representation for spatio-temporal photovoltaic power predictionabstractAbstract Clean energy is a major trend. The importance of photovoltaic power generation is also growing. Photovoltaic power generation is mainly affected by the weather. It is full of uncertainties. Previous work has relied chiefly on historical photovoltaics data for time series forecasts. However, unforeseen weather conditions can sometimes skew. Consequently, a spatial‐temporal‐meteorological‐long short‐term memory prediction model (STM‐LSTM) is proposed to compensate for the shortage of photovoltaic prediction models for uncertainties. This model can simultaneously process satellite image data, historical meteorological data, and historical power generation data. In this way, historical patterns and meteorological change information are extracted to improve the accuracy of photovoltaic prediction. STM‐LSTM processes raw satellite data to obtain cloud image data. It can extract cloud motion information using the dense optical flow method. First, the cloud images are processed to extract cloud position information. By adaptive attentive learning of images in different bands, a better representation for subsequent tasks can be obtained. Second, it is important to process historical meteorological data to learn meteorological change patterns. Last but not least, the historical photovoltaic power generation sequences are combined to obtain the final photovoltaic prediction results. After a series of experimental validation, the performance of the proposed STM‐LSTM model has a good improvement compared with the baseline model. Xingchen Guo, Jing Lai, Chenxiang Lin, Yuxing Dai, Xuexin Xu, Haisheng San, Rong Jia, Zhihong Zhang 0001 |
IET Comput. Vis. | 9 |
| 2023 | Position-aware and structure embedding networks for deep graph matching
Dongdong Chen 0003, Yuxing Dai, Lichi Zhang, Zhihong Zhang 0001, Edwin R. Hancock |
Pattern Recognit. | 4 |
| 2023 | Any-to-Any Voice Conversion With Multi-Layer Speaker Adaptation and Content SupervisionabstractAny-to-any voice conversion can be performed among arbitrary speakers, even with a single reference utterance. Many related studies have demonstrated that it can be effectively implemented by speech representation disentanglement. However, most existing solutions fuse the speaker representations into the content features globally without considering their distribution difference. Additionally, in the any-to-any scenario, there is no effective method ensuring the consistency of linguistic content without text transcription or additional information extracted from additional modules (e.g., automatic speech recognition). Hence, to alleviate the above problems, this paper proposes SACS-VC, a novel any-to-any voice conversion method that combines two principal modules: Speaker Adaptation and Content Supervision. Specifically, we rearrange the timbre representations according to the content distribution using a temporal attention mechanism to obtain finer-grained speaker timbre information for each content feature. Meanwhile, we associate the converted outputs and source utterances directly to supervise the consistency of the semantic content in an unsupervised manner. This is achieved using contrastive learning based on the corresponding and non-corresponding locations of content features. It should be noted that SACS-VC can be implemented using a non-parallel speech corpus without any pertaining. The experimental results demonstrate that the proposed method outperforms current state-of-the-art any-to-any voice conversion systems in objective and subjective evaluation settings. Xuexin Xu, Xunquan Chen, Pingyuan Lin, Jie Lian 0005, Zhihong Zhang 0001, Edwin R. Hancock |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2023 | Spatial Adaptive Fusion Consistency Contrastive Constraint: Weakly Supervised Building Facade Point Cloud Semantic SegmentationabstractSemantic segmentation of building facade point clouds has diverse applications. The development of semantic segmentation methods is inextricably linked to datasets. The available building facade datasets suffer from a lack of abundant semantic categories and data completeness. To compensate for these shortcomings, we propose a new building facade dataset characterized by various categories and relatively complete 3-D building facades. In addition, most existing methods focus on fully supervised learning, which relies on manually labeling large-scale point cloud data and results in high time and labor costs. In this article, we propose an effective weakly supervised building facade segmentation approach, called spatial adaptive fusion consistency contrastive constraint (SAF-C3), to solve the above problem. We first design a multirandom point cloud augmentor as an auxiliary supervision branch to enhance the learning ability of the original network branch. Then, we present a spatial adaptive fusion (SAF) module to extract discriminative features for building facade point clouds. Finally, we propose a spatial consistency contrastive constraint to explore the contrastive property in feature space and to ensure the predictive consistency among the augmentation and original branches. The proposed method achieves a significant performance improvement against the state-of-the-art methods on two building facade point cloud datasets through extensive experiments. In particular, the performance of SAF-C3 with 1% labels significantly surpasses the baseline network with 100% labels. Yanfei Su, Ming Cheng 0002, Zhimin Yuan, Weiquan Liu, Wankang Zeng, Zhihong Zhang 0001, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Speaker-Independent Emotional Voice Conversion via Disentangled RepresentationsabstractEmotional Voice Conversion (EVC) technology aims to transfer emotional state in speech while keeping the linguistic information and speaker identity unchanged. Prior studies on EVC have been limited to perform the conversion for a specific speaker or a predefined set of multiple speakers seen in the training stage. When encountering arbitrary speakers that may be unseen during training (outside the set of speakers used in training), existing EVC methods have limited conversion capabilities. However, converting the emotion of arbitrary speakers, even those unseen during the training procedure, in one model is much more challenging and much more attractive in real-world scenarios. To address this problem, in this study, we propose SIEVC, a novel speaker-independent emotional voice conversion framework for arbitrary speakers via disentangled representation learning. The proposed method employs the autoencoder framework to disentangle the emotion information and emotion-independent information of each input speech into separated representation spaces. To achieve better disentanglement, we incorporate mutual information minimization into the training process. In addition, adversarial training is applied to enhance the quality of the generated audio signals. Finally, speaker-independent EVC for arbitrary speakers could be achieved by only replacing the emotion representations of source speech with the target ones. The experimental results demonstrate that the proposed EVC model outperforms the baseline models in terms of objective and subjective evaluation for both seen and unseen speakers. Xunquan Chen, Xuexin Xu, Zhihong Zhang 0001, Tetsuya Takiguchi, Edwin R. Hancock |
IEEE Trans. Multim. | 4 |
| 2023 | Entropic Dynamic Time Warping Kernels for Co-Evolving Financial Time Series AnalysisabstractNetwork representations are powerful tools to modeling the dynamic time-varying financial complex systems consisting of multiple co-evolving financial time series, e.g., stock prices. In this work, we develop a novel framework to compute the kernel-based similarity measure between dynamic time-varying financial networks. Specifically, we explore whether the proposed kernel can be employed to understand the structural evolution of the financial networks with time associated with standard kernel machines. For a set of time-varying financial networks with each vertex representing the individual time series of a different stock and each edge between a pair of time series representing the absolute value of their Pearson correlation, our start point is to compute the commute time (CT) matrix associated with the weighted adjacency matrix of the network structures, where each element of the matrix can be seen as the enhanced correlation value between pairwise stocks. For each network, we show how the CT matrix allows us to identify a reliable set of dominant correlated time series as well as an associated dominant probability distribution of the stock belonging to this set. Furthermore, we represent each original network as a discrete dominant Shannon entropy time series computed from the dominant probability distribution. With the dominant entropy time series for each pair of financial networks to hand, we develop an entropic dynamic time warping kernels through the classical dynamic time warping framework, for analyzing the financial time-varying networks. We show that the proposed kernel bridges the gap between graph kernels and the classical dynamic time warping framework for multiple financial time series analysis. Experiments on time-varying networks extracted through New York Stock Exchange (NYSE) database demonstrate that the effectiveness of the proposed method. Lu Bai 0001, Lixin Cui, Zhihong Zhang 0001, Lixiang Xu, Yue Wang 0014, Edwin R. Hancock |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Graph Motif Entropy for Understanding Time-Evolving NetworksabstractThe structure of networks can be efficiently represented using motifs, which are those subgraphs that recur most frequently. One route to understanding the motif structure of a network is to study the distribution of subgraphs using statistical mechanics. In this article, we address the use of motifs as network primitives using the cluster expansion from statistical physics. By mapping the network motifs to clusters in the gas model, we derive the partition function for a network, and this allows us to calculate global thermodynamic quantities, such as energy and entropy. We present analytical expressions for the number of certain types of motifs, and compute their associated entropy. We conduct numerical experiments for synthetic and real-world data sets and evaluate the qualitative and quantitative characterizations of the motif entropy derived from the partition function. We find that the motif entropy for real-world networks, such as financial stock market networks, is sensitive to the variance in network structure. This is in line with recent evidence that network motifs can be regarded as basic elements with well-defined information-processing functions. Zhihong Zhang 0001, Dongdong Chen 0003, Lu Bai 0001, Jianjia Wang, Edwin R. Hancock |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Direction of arrival estimation for indoor environments based on acoustic composition model with a single microphone
Xingchen Guo, Xuexin Xu, Xunquan Chen, Rong Jia, Zhihong Zhang 0001, Tetsuya Takiguchi, Edwin R. Hancock |
Pattern Recognit. | 6 |
| 2022 | DLA-Net: Learning dual local attention features for semantic segmentation of large-scale building facade point clouds
Yanfei Su, Weiquan Liu, Zhimin Yuan, Ming Cheng 0002, Zhihong Zhang 0001, Xuelun Shen, Cheng Wang 0003 |
Pattern Recognit. | 5 |
| 2022 | LiDAR-based localization using universal encoding and memory-aware regression
Shangshu Yu, Cheng Wang 0003, Chenglu Wen, Ming Cheng 0002, Minghao Liu 0007, Zhihong Zhang 0001, Xin Li 0003 |
Pattern Recognit. | 6 |
| 2022 | Corrigendum to "LiDAR-based localization using universal encoding and memory-aware regression" Pattern Recognition Volume 128 (2022) 108685
Shangshu Yu, Cheng Wang 0003, Chenglu Wen, Ming Cheng 0002, Minghao Liu 0007, Zhihong Zhang 0001, Xin Li 0003 |
Pattern Recognit. | 6 |
| 2022 | Feature refinement: An expression-specific feature learning and fusion method for micro-expression recognition
Ling Zhou 0005, Qirong Mao, Xiaohua Huang 0002, Feifei Zhang 0001, Zhihong Zhang 0001 |
Pattern Recognit. | 5 |
| 2021 | Blockchain Based Trusted Identity Authentication in Ubiquitous Power Internet of Things
Jie Lian 0005, Dianwei Jin, Aleksei Balabontsev, Zhihong Zhang 0001 |
ICIC (2) | 9 |
| 2021 | Two-Pathway Style Embedding for Arbitrary Voice ConversionabstractArbitrary voice conversion, also referred to as zero-shot voice conversion, has recently attracted increased attention in the literature.Although disentangling the linguistic and style representations for acoustic features is an effective way to achieve zero-shot voice conversion, the problem of how to convert to a natural speaker style is challenging because of the intrinsic variabilities of speech and the difficulties of completely decoupling them.For this reason, in this paper, we propose a Two-Pathway Style Embedding Voice Conversion framework (TPSE-VC) for realistic and natural speech conversion.The novel feature of this method is to simultaneously embed sentence-level and phoneme-level style information.A novel attention mechanism is proposed to implement the implicit alignment for timbre style and phoneme content, further embedding a phoneme-level style representation.In addition, we consider embedding the complete set of time steps of audio style into a fixed-length vector to obtain the sentence-level style representation.Moreover, TPSE-VC does not require any pre-trained models, and is only trained with non-parallel speech data.Experimental results demonstrate that the proposed TPSE-VC outperforms the state-of-theart results on zero-shot voice conversion. Xuexin Xu, Xunquan Chen, Jie Lian 0005, Pingyuan Lin, Zhihong Zhang 0001, Edwin R. Hancock |
Interspeech | 7 |
| 2021 | Thermodynamic motif analysis for directed stock market networks
Dongdong Chen 0003, Xingchen Guo, Jianjia Wang, Zhihong Zhang 0001, Edwin R. Hancock |
Pattern Recognit. | 5 |
| 2021 | Multimodal fusion for indoor sound source localization
Ryoichi Takashima, Xingchen Guo, Zhihong Zhang 0001, Xuexin Xu, Tetsuya Takiguchi, Edwin R. Hancock |
Pattern Recognit. | 4 |
| 2021 | Statistical mechanical analysis for unweighted and weighted stock market networks
Jianjia Wang, Xingchen Guo, Weimin Li 0001, Xing Wu 0001, Zhihong Zhang 0001, Edwin R. Hancock |
Pattern Recognit. | 5 |
| 2021 | Semi-Supervised Face Frontalization in the WildabstractSynthesizing a frontal view face from a single nonfrontal image, i.e. face frontalization, is a task of practical importance in a wide range of facial image analysis applications. However, to train the frontalization model in a supervised manner, most existing face frontalization methods rely on the availability of nonfrontal-frontal face pairs (typically from the Multi-PIE dataset) captured in a constrained environment. Such approaches, in return, limit the generalizability of their application to unconstrained scenarios. Unfortunately, although a large amount of in-the-wild face datasets are available, they cannot easily be utilized for face frontalization training since the nonfrontal and frontal facial images are not paired. To train a frontalization network which generalizes well to both constrained and unconstrained environments, we propose a semi-supervised learning framework which effectively uses both (labeled) indoor and (unlabeled) outdoor faces. Specifically, to achieve this goal, this article presents a Cycle-Consistent Face Frontalization Generative Adversarial Network (CCFF-GAN) which consists of both (1) the supervised and (2) the unsupervised components. For (1), we use the indoor paired (labeled) data to learn a roughly accurate frontalization network which may not generalize well to outdoor (in-the-wild) scenarios. For (2), to cope with the generalization issue, the unsupervised part uses the unpaired (unlabeled) images under the perceptual cycle consistency constraint in the semantic feature space to generalize the network from controlled (indoor) to uncontrolled (outdoor) environment. Extensive experiments demonstrate the effectiveness of the proposed method in comparison with the state-of-the-art face frontalization methods, especially under the in-the-wild scenarios. Zhihong Zhang 0001, Ruiyang Liang, Xu Chen 0020, Xuexin Xu, Guosheng Hu, Wangmeng Zuo, Edwin R. Hancock |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Learning EEG topographical representation for classification via convolutional neural network
Meiyan Xu, Junfeng Yao, Zhihong Zhang 0001, Baorong Yang, Chunyan Li 0002, Junsong Zhang |
Pattern Recognit. | 3 |
| 2020 | Spectral bounding: Strictly satisfying the 1-Lipschitz property for generative adversarial networks
Zhihong Zhang 0001, Yangbin Zeng, Lu Bai 0001, Yiqun Hu, Meihong Wu, Shuai Wang 0049, Edwin R. Hancock |
Pattern Recognit. | 1 |
| 2019 | A graph-based approach to automated EUS image layer segmentation and abnormal region detection
Xu Chen 0020, Yiqun Hu, Zhihong Zhang 0001, Beizhan Wang, Lichi Zhang, Xinjian Chen 0001, Xiaoyi Jiang 0001 |
Neurocomputing | 3 |
| 2019 | Identifying the most informative features using a structurally interacting elastic net
Lixin Cui, Lu Bai 0001, Zhihong Zhang 0001, Yue Wang 0014, Edwin R. Hancock |
Neurocomputing | 3 |
| 2019 | Quantum-based subgraph convolutional neural networks
Zhihong Zhang 0001, Dongdong Chen 0003, Jianjia Wang, Lu Bai 0001, Edwin R. Hancock |
Pattern Recognit. | 1 |
| 2019 | Depth-based subgraph convolutional auto-encoder for network representation learning
Zhihong Zhang 0001, Dongdong Chen 0003, Zeli Wang, Heng Li 0001, Lu Bai 0001, Edwin R. Hancock |
Pattern Recognit. | 1 |
| 2019 | Single image super resolution via neighbor reconstruction
Zhihong Zhang 0001, Yide Cai, Zeli Wang, Heng Li 0001, Edwin R. Hancock |
Pattern Recognit. Lett. | 1 |
| 2019 | Structural network inference from time-series data using a generative model and transfer entropy
Zhihong Zhang 0001, Genzhou Zhang, Yangbin Zeng, Beizhan Wang, Edwin R. Hancock |
Pattern Recognit. Lett. | 1 |
| 2019 | Face Frontalization Using an Appearance-Flow-Based Convolutional Neural NetworkabstractFacial pose variation is one of the major factors making face recognition (FR) a challenging task. One popular solution is to convert non-frontal faces to frontal ones on which FR is performed. Rotating faces causes facial pixel value changes. Therefore, existing CNN-based methods learn to synthesize frontal faces in color space. However, this learning problem in a color space is highly non-linear, causing the synthetic frontal faces to lose fine facial textures. In this paper, we take the view that the nonfrontal-frontal pixel changes are essentially caused by geometric transformations (rotation, translation, and so on) in space. Therefore, we aim to learn the nonfrontal-frontal facial conversion in the spatial domain rather than the color domain to ease the learning task. To this end, we propose an appearance-flow-based face frontalization convolutional neural network (A3F-CNN). Specifically, A3F-CNN learns to establish the dense correspondence between the non-frontal and frontal faces. Once the correspondence is built, frontal faces are synthesized by explicitly "moving" pixels from the non-frontal one. In this way, the synthetic frontal faces can preserve fine facial textures. To improve the convergence of training, an appearance-flow-guided learning strategy is proposed. In addition, generative adversarial network loss is applied to achieve a more photorealistic face, and a face mirroring method is introduced to handle the self-occlusion problem. Extensive experiments are conducted on face synthesis and pose invariant FR. Results show that our method can synthesize more photorealistic faces than the existing methods in both the controlled and uncontrolled lighting environments. Moreover, we achieve a very competitive FR performance on the Multi-PIE, LFW and IJB-A databases. Zhihong Zhang 0001, Xu Chen 0020, Beizhan Wang, Guosheng Hu, Wangmeng Zuo, Edwin R. Hancock |
IEEE Trans. Image Process. | 1 |
| 2019 | Polar Transformation on Image Features for Orientation-Invariant RepresentationsabstractThe choice of image feature representation plays a crucial role in the analysis of visual information. Although vast numbers of alternative robust feature representation models have been proposed to improve the performance of different visual tasks, most existing feature representations [e.g., handcrafted features or convolutional neural networks (CNNs)] have a relatively limited capacity to capture the highly orientation-invariant (rotation/reversal) features. The net consequence is suboptimal visual performance. To address these problems, this study adopts a novel transformational approach, which investigates the potential of using polar feature representations. Our low level consists of a histogram of oriented gradient, which is then binned using annular spatial bin-type cells applied to the polar gradient. This gives gradient binning invariance for feature extraction. In this way, the descriptors have significantly enhanced orientation-invariant capabilities. The proposed feature representation, calledorientation-invariant histograms of oriented gradients, is capable of accurately processing visual tasks (e.g., facial expression recognition). In the context of the CNN architecture, we propose two polar convolution operations, referred to as full polar convolution and local polar convolution, and use these to develop polar architectures for the CNN orientation-invariant representation. Experimental results show that the proposed orientation-invariant image representation, based on polar models for both handcrafted features and deep learning features, is both competitive with state-of-the-art methods and maintains compact representation on a set of challenging benchmark image datasets. Zhaojie Luo, Zhihong Zhang 0001, Faliang Huang, Zhiling Ye, Tetsuya Takiguchi, Edwin R. Hancock |
IEEE Trans. Multim. | 3 |
| 2018 | Deep Multi-task Learning to Recognise Subtle Facial Expressions of Mental States
Guosheng Hu, Li Liu 0004, Yang Hua 0001, Zhihong Zhang 0001, Fumin Shen, Ling Shao 0001, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang |
ECCV (12) | 6 |
| 2018 | Deep Stock Representation Learning: From Candlestick Charts to Investment DecisionsabstractWe propose a novel investment decision strategy (IDS) based on deep learning. The performance of many IDSs is affected by stock similarity. Most existing stock similarity measurements have the problems: (a) The linear nature of many measurements cannot capture nonlinear stock dynamics; (b) The estimation of many similarity metrics (e.g. covariance) needs very long period historic data (e.g. 3K days) which cannot represent current market effectively; (c) They cannot capture translation-invariance. To solve these problems, we apply Convolutional AutoEncoder to learn a stock representation, based on which we propose a novel portfolio construction strategy by: (i) using the deeply learned representation and modularity optimisation to cluster stocks and identify diverse sectors, (ii) picking stocks within each cluster according to their Sharpe ratio (Sharpe 1994). Overall this strategy provides low-risk high-return portfolios. We use the Financial Times Stock Exchange 100 Index (FTSE 100) data for evaluation. Results show our portfolio outperforms FTSE 100 index and many well known funds in terms of total return in 2000 trading days. Guosheng Hu, Kai Yang 0031, Flood Sung, Zhihong Zhang 0001, Neil Robertson 0002, Timothy M. Hospedales, Qiangwei Miemie |
ICASSP | 6 |
| 2018 | Depth-based Subgraph Convolutional Neural NetworksabstractThis paper proposes a new graph convolutional neural architecture based on a depth-based representation of graph structure, called the depth-based subgraph convolutional neural networks (DS-CNNs), which integrates both the global topological and local connectivity structures within a graph. Our idea is to decompose a graph into a family of$K$-layer expansion subgraphs rooted at each vertex, and then a set of convolution filters are designed over these subgraphs to capture local connectivity structural information. Specifically, we commence by establishing a family of$K$-layer expansion subgraphs for each vertex of graph by mapping graph to tree procedures, which can provide global topological arrangement information contained within a graph. We then design a set of fixed-size convolution filters and integrate them with these subgraphs (depicted in Figure 1). The idea is to apply convolution filters sliding over the entire subgraphs of a vertex to extract the local features analogous to the standard convolution operation on grid data. In particular, the convolution operation captures the local structural information within the graph, and has the weight sharing property among different positions of subgraph; the pooling operation acts directly on the output of the preceding layer without any preprocessing scheme (e.g., clustering or other techniques). Experiments on three graph-structured datasets demonstrate that our model DS-CNNs are able to outperform six state-of-the-art methods at the task of node classification. Chuanyu Xu, Zhihong Zhang 0001, Beizhan Wang, Da Zhou, Guijun Ren, Lu Bai 0001, Lixin Cui, Edwin R. Hancock |
ICPR | 3 |
| 2018 | A Unified Neighbor Reconstruction Method for EmbeddingsabstractIn this work we propose a novel and compact Neighbor Reconstruction Method (NRM) which is a unified pre-processing method for graph-based sparse spectral algorithms. This method is conducted by vector operations on a central point and its corresponding neighbor points. NRM generates new neighbor points which can capture the local space structure of the central point more appropriately than original neighbor points. With NRM, a large number of sparse spectral based nonlinear feature extraction and selection algorithms gain significant improvement. Specifically, we embedded NRM to several classical algorithms, Local Linear Embedding (LLE) [1], Laplacian Eigenmaps (LE) [2] and Unsupervised Feature Selection for Multi-cluster Data (MCFS) [3], with accuracy improvement of up to 7%, 2.6%, 2.4% on ORL, CIFAR 10, and MINST data sets respectively. We also apply NRM to a Super Resolution algorithm, A+ [5], and obtain 0.12dB improvement than original method. Zhiling Ye, Zhihong Zhang 0001, Lu Bai 0001, Guosheng Hu, Zheng-Jian Bai, Yiqun Hu, Edwin R. Hancock |
ICPR | 2 |
| 2018 | H-Net: Neural Network for Cross-domain Image Patch MatchingabstractDescribing the same scene with different imaging style or rendering image from its 3D model gives us different domain images. Different domain images tend to have a gap and different local appearances, which raise the main challenge on the cross-domain image patch matching. In this paper, we propose to incorporate AutoEncoder into the Siamese network, named as H-Net, of which the structural shape resembles the letter H. The H-Net achieves state-of-the-art performance on the cross-domain image patch matching. Furthermore, we improved H-Net to H-Net++. The H-Net++ extracts invariant feature descriptors in cross-domain image patches and achieves state-of-the-art performance by feature retrieval in Euclidean space. As there is no benchmark dataset including cross-domain images, we made a cross-domain image dataset which consists of camera images, rendering images from UAV 3D model, and images generated by CycleGAN algorithm. Experiments show that the proposed H-Net and H-Net++ outperform the existing algorithms. Our code and cross-domain image dataset are available at https://github.com/Xylon-Sean/H-Net. Weiquan Liu, Xuelun Shen, Cheng Wang 0003, Zhihong Zhang 0001, Chenglu Wen, Jonathan Li 0001 |
IJCAI | 4 |
| 2018 | Accurate geometry modeling of vasculatures using implicit fitting with 2D radial basis functions
Qingqi Hong, Qingde Li, Beizhan Wang, Kunhong Liu 0001, Fan Lin, Juncong Lin, Zhihong Zhang 0001, Ming Zeng 0008 |
Comput. Aided Geom. Des. | 8 |
| 2018 | Recovering variations in facial albedo from low resolution images
Xu Chen 0020, Zhihong Zhang 0001, Beizhan Wang, Guosheng Hu, Edwin R. Hancock |
Pattern Recognit. | 2 |
| 2017 | Attribute-Enhanced Face Recognition with Neural Tensor Fusion NetworksabstractDeep learning has achieved great success in face recognition, however deep-learned features still have limited invariance to strong intra-personal variations such as large pose changes. It is observed that some facial attributes (e.g. eyebrow thickness, gender) are robust to such variations. We present the first work to systematically explore how the fusion of face recognition features (FRF) and facial attribute features (FAF) can enhance face recognition performance in various challenging scenarios. Despite the promise of FAF, we find that in practice existing fusion methods fail to leverage FAF to boost face recognition performance in some challenging scenarios. Thus, we develop a powerful tensor-based framework which formulates feature fusion as a tensor optimisation problem. It is nontrivial to directly optimise this tensor due to the large number of parameters to optimise. To solve this problem, we establish a theoretical equivalence between low-rank tensor optimisation and a two-stream gated neural network. This equivalence allows tractable learning using standard neural network optimisation tools, leading to accurate and stable optimisation. Experimental results show the fused feature works better than individual features, thus proving for the first time that facial attributes aid face recognition. We achieve state-of-the-art performance on three popular databases: MultiPIE (cross pose, lighting and expression), CASIA NIR-VIS2.0 (cross-modality environment) and LFW (uncontrolled environment). Guosheng Hu, Yang Hua 0001, Zhihong Zhang 0001, Sankha S. Mukherjee, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang |
ICCV | 4 |
| 2017 | Joint hypergraph learning and sparse regression for feature selection
Zhihong Zhang 0001, Lu Bai 0001, Yuanheng Liang, Edwin R. Hancock |
Pattern Recognit. | 1 |
| 2017 | Quantum kernels for unattributed graphs using discrete-time quantum walks
Lu Bai 0001, Luca Rossi 0004, Lixin Cui, Zhihong Zhang 0001, Peng Ren 0001, Xiao Bai 0001, Edwin R. Hancock |
Pattern Recognit. Lett. | 4 |
| 2017 | High-order covariate interacted Lasso for feature selection
Zhihong Zhang 0001, Yiyang Tian, Lu Bai 0001, Jianbing Xiahou, Edwin R. Hancock |
Pattern Recognit. Lett. | 1 |
| 2016 | Face image super-resolution via weighted patches regressionabstractRecently sparse representation has gained great success in face image super-resolution. The conventional sparsity-based methods enforce sparse coding on face image patches and the representation fidelity is measured by ℓ2-norm. Such a sparse coding model regularizes all facial patches equally, which however ignores the natures of facial patches, where the facial patches in the different regions (patch positions) of human face may have distinct contributions to face image reconstruction. In this paper, we propose to weight facial patches based on their discriminative abilities in regression for robust face hallucination reconstruction. Specifically, we learn the weights for facial patches according to the information entropy in each face region, so as to highlight higher frequency details in face images and the facial discriminability can be well retrieved. Furthermore, the weighted sparse coding can reasonable represent the less sparse nature of noisy images and thus remarkably boosts noise robust performance in face image super-resolution. Various experimental results on standard face databased show that our proposed method outperforms state-of-the-art methods in terms of both objective metrics and visual quality. Zhihong Zhang 0001, Guosheng Hu, Edwin R. Hancock |
ICPR | 2 |
| 2016 | High-order graph matching kernel for early carcinoma EUS image classification
Zhihong Zhang 0001, Lu Bai 0001, Peng Ren 0001, Edwin R. Hancock |
Multim. Tools Appl. | 1 |
| 2016 | Discriminative sparse representation for face recognition
Zhihong Zhang 0001, Yuanheng Liang, Lu Bai 0001, Edwin R. Hancock |
Multim. Tools Appl. | 1 |
| 2015 | A High-Order Depth-Based Graph Matching Method
Lu Bai 0001, Zhihong Zhang 0001, Peng Ren 0001, Edwin R. Hancock |
CAIP (1) | 2 |
| 2015 | An Edge-Based Matching Kernel for Graphs Through the Directed Line Graphs
Lu Bai 0001, Zhihong Zhang 0001, Chaoyan Wang, Edwin R. Hancock |
CAIP (2) | 2 |
| 2015 | Adaptive Graph Learning for Unsupervised Feature Selection
Zhihong Zhang 0001, Lu Bai 0001, Yuanheng Liang, Edwin R. Hancock |
CAIP (1) | 1 |
| 2015 | An Aligned Subtree Kernel for Weighted GraphsabstractIn this paper, we develop a new entropic matching kernel for weighted graphs by aligning depth-based representations. We demonstrate that this kernel can be seen as an \textbfaligned subtree kernel that incorporates explicit subtree correspondences, and thus addresses the drawback of neglecting the relative locations between substructures that arises in the R-convolution kernels. Experiments on standard datasets demonstrate that our kernel can easily outperform state-of-the-art graph kernels in terms of classification accuracy. Lu Bai 0001, Luca Rossi 0004, Zhihong Zhang 0001, Edwin R. Hancock |
ICML | 3 |
| 2015 | A Graph Kernel Based on the Jensen-Shannon Representation Alignment
Lu Bai 0001, Zhihong Zhang 0001, Chaoyan Wang, Xiao Bai 0001, Edwin R. Hancock |
IJCAI | 2 |
| 2014 | Adaptive Object Retrieval with Kernel Reconstructive HashingabstractHashing is very useful for fast approximate similarity search on large database. In the unsupervised settings, most hashing methods aim at preserving the similarity defined by Euclidean distance. Hash codes generated by these approaches only keep their Hamming distance corresponding to the pairwise Euclidean distance, ignoring the local distribution of each data point. This objective does not hold for k-nearest neighbors search. In this paper, we firstly propose a new adaptive similarity measure which is consistent with k-NN search, and prove that it leads to a valid kernel. Then we propose a hashing scheme which uses binary codes to preserve the kernel function. Using low-rank approximation, our hashing framework is more effective than existing methods that preserve similarity over arbitrary kernel. The proposed kernel function, hashing framework, and their combination have demonstrated significant advantages compared with several state-of-the-art methods. Haichuan Yang, Xiao Bai 0001, Jun Zhou 0001, Peng Ren 0001, Zhihong Zhang 0001, Jian Cheng 0001 |
CVPR | 5 |
| 2013 | A hypergraph based semi-supervised band selection method for hyperspectral image classificationabstractBand selection is a fundamental problem in hyperspectral data processing. In this paper, we present a semi-supervised learning approach and a hypergraph model to select useful bands based on few labeled object information. The contributions of this paper are two-fold. Firstly, the hypergraph model captures multiple relationships between hyperspectral image samples. Secondly, the semi-supervised learning method not only utilizes unlabeled samples in the learning process to improve model performance, but also requires little labeled samples which can significantly reduce large amount of human labor and costs. The proposed approach is evaluated on AVIRIS and APHI datasets, which demonstrate its advantages over several other band selection methods. Zhouxiao Guo, Xiao Bai 0001, Zhihong Zhang 0001, Jun Zhou 0001 |
ICIP | 3 |
| 2013 | Semi-supervised hyperspectral band selection via sparse linear regression and hypergraph modelsabstractBand selection is an important step towards effective and efficient object classification in hyperspectral imagery. In this paper, we propose a semi-supervised learning method for band selection based on a sparse linear regression model. This model uses a least absolute shrinkage and selection operator to compute the regression coefficients from both labeled and unlabeled samples. These coefficients are then used to compute a contribution score for each band, which allows bands with high scores being selected for the testing step. During this process, unlabeled samples also contribute to the coefficients calculation. In order to propagate the labels to these samples, a hypergraph is first built to describe the relationship between labeled and unlabeled samples. This leads to an adjacency matrix whose entries are the sum of corresponding weights of hyperedges. Then matrix subspace learning method is used to estimate the labels of unlabeled samples. The proposed method is evaluated on the APHI dataset. Comparison with several baseline methods has shown the advantages of the proposed method on the pixel-level classification. Zhouxiao Guo, Haichuan Yang, Xiao Bai 0001, Zhihong Zhang 0001, Jun Zhou 0001 |
IGARSS | 4 |
| 2012 | Unsupervised Feature Selection Via Hypergraph EmbeddingabstractMost existing feature selection methods focus on ranking individual features based on a utility criterion, and select the optimal feature set in a greedy manner.However, the feature combinations found in this way do not give optimal classification performance, since they tend to neglect the correlations among features.In an attempt to overcome this problem, we develop a novel unsupervised feature selection technique by using hypergraph spectral embedding, where the projection matrix is constrained to be a selection matrix designed to select the optimal feature subset.Specifically, by using multidimensional interaction information (MII) as a higher order similarity measure, we establish a novel hypergraph framework which is used for characterizing the multiple relationships within a set of samples.Thus, the structural information latent in the data can be more effectively modeled.We then derive a hypergraph embedding view of feature selection which casts the feature discriminant analysis into a regression framework that considers the correlations among features.Within our framework, features are evaluated in combinations rather than considered individually, and feature redundancies can thus be addressed accordingly.Experimental results demonstrate the effectiveness of our feature selection method on a number of standard datasets. Zhihong Zhang 0001, Peng Ren 0001, Edwin R. Hancock |
BMVC | 1 |
| 2012 | Face recognition using semi-supervised spectral feature selection
Zhihong Zhang 0001, Edwin R. Hancock |
ICPR | 1 |
| 2012 | Unsupervised spectral feature selection for face recognition
Zhihong Zhang 0001, Edwin R. Hancock |
ICPR | 1 |
| 2012 | Hypergraph based semi-supervised learning for gender classification
Zhihong Zhang 0001, Edwin R. Hancock, Peng Ren 0001 |
ICPR | 1 |
| 2012 | Hypergraph Spectra for Semi-supervised Feature Selection
Zhihong Zhang 0001, Edwin R. Hancock, Xiao Bai 0001 |
ECML/PKDD (1) | 1 |
| 2012 | Kernel Entropy-Based Unsupervised spectral Feature SelectionabstractMost existing feature selection methods focus on ranking individual features based on a utility criterion, and select the optimal feature set in a greedy manner. However, the feature combinations found in this way do not give optimal classification performance, since they neglect the correlations among features. In an attempt to overcome this problem, we develop a novel feature selection technique using the spectral data transformation and by using ℓ1-norm regularized models for subset selection. Specifically, we propose a new two-step spectral regression technique for unsupervised feature selection. In the first step, we use kernel entropy component analysis (kECA) to transform the data into a lower-dimensional space so as to improve class separation. Second, we use ℓ1-norm regularization to select the features that best align with the data embedding resulting from kECA. The advantage of kECA is that dimensionality reducing data transformation maximally preserves entropy estimates for the input data whilst also best preserving the cluster structure of the data. Using ℓ1-norm regularization, we cast feature discriminant analysis into a regression framework which accommodates the correlations among features. As a result, we can evaluate joint feature combinations, rather than being confined to consider them individually. Experimental results demonstrate the effectiveness of our feature selection method on a number of standard face datasets. Zhihong Zhang 0001, Edwin R. Hancock |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2012 | Hypergraph based information-theoretic feature selection
Zhihong Zhang 0001, Edwin R. Hancock |
Pattern Recognit. Lett. | 1 |
| 2011 | A Hypergraph-Based Approach to Feature Selection
Zhihong Zhang 0001, Edwin R. Hancock |
CAIP (1) | 1 |