VLDB 2026 Research / reviewers in the wild / expert
Guoqiang Zhang 0003
dblp:24/5631-3
· DBLP profile ↗
46ranked-venue papers
26as first author
12since 2021 · last 2025
0000-0003-4521-542XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 19 first-author · 3 since 2021Artificial intelligence and machine learning · 22 · 9 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On Exact Bit-level Reversible Transformers Without Changing ArchitectureabstractIn this work we present the BDIA-transformer, which is an exact bit-level reversible transformer that uses an unchanged standard architecture for inference. The basic idea is to first treat each transformer block as the Euler integration approximation for solving an ordinary differential equation (ODE) and then incorporate the technique of bidirectional integration approximation (BDIA) (originally designed for diffusion inversion) into the neural architecture, together with activation quantization to make it exactly bit-level reversible. In the training process, we let a hyper-parameter $\gamma$ in BDIA-transformer randomly take one of the two values $\{0.5, -0.5\}$ per training sample per transformer block for averaging every two consecutive integration approximations. As a result, BDIA-transformer can be viewed as training an ensemble of ODE solvers parameterized by a set of binary random variables, which regularizes the model and results in improved validation accuracy. Lightweight side information is required to be stored in the forward process to account for binary quantization loss to enable exact bit-level reversibility. In the inference procedure, the expectation $\mathbb{E}(\gamma)=0$ is taken to make the resulting architecture identical to transformer up to activation quantization. Our experiments in natural language generation, image classification, and language translation show that BDIA-transformers outperform their conventional counterparts significantly in terms of validation performance while also requiring considerably less training memory. Thanks to the regularizing effect of the ensemble, the BDIA-transformer is particularly suitable for fine-tuning with limited data. Source-code can be found via https://github.com/guoqiang-zhang-x/BDIA-Transformer. Guoqiang Zhang 0003, John P. Lewis, W. Bastiaan Kleijn |
ICML | 1 |
| 2025 | Revisiting 1-peer exponential graph for enhancing decentralized learning efficiencyabstractFor communication-efficient decentralized learning, it is essential to employ dynamic graphs designed to improve the expected spectral gap by reducing deviations from global averaging. The $1$-peer exponential graph demonstrates its finite-time convergence property--achieved by maximizing the expected spectral gap--but only when the number of nodes $n$ is a power of two. However, its efficiency across any $n$ and the commutativity of mixing matrices remain unexplored. We delve into the principles underlying the $1$-peer exponential graph to explain its efficiency across any $n$ and leverage them to develop new dynamic graphs. We propose two new dynamic graphs: the $k$-peer exponential graph and the null-cascade graph. Notably, the null-cascade graph achieves finite-time convergence for any $n$ while ensuring commutativity. Our experiments confirm the effectiveness of these new graphs, particularly the null-cascade graph, in most test settings. Kenta Niwa, Yuki Takezawa, Guoqiang Zhang 0003, W. Bastiaan Kleijn |
NeurIPS | 3 |
| 2024 | Exact Diffusion Inversion via Bidirectional Integration Approximation
Guoqiang Zhang 0003, John P. Lewis, W. Bastiaan Kleijn |
ECCV (57) | 1 |
| 2024 | On Accelerating Diffusion-Based Sampling Processes via Improved Integration ApproximationabstractA popular approach to sample a diffusion-based generative model is to solve an ordinary differential equation (ODE). In existing samplers, the coefficients of the ODE solvers are pre-determined by the ODE formulation, the reverse discrete timesteps, and the employed ODE methods. In this paper, we consider accelerating several popular ODE-based sampling processes (including EDM, DDIM, and DPM-Solver) by optimizing certain coefficients via improved integration approximation (IIA). We propose to minimize, for each time step, a mean squared error (MSE) function with respect to the selected coefficients. The MSE is constructed by applying the original ODE solver for a set of fine-grained timesteps, which in principle provides a more accurate integration approximation in predicting the next diffusion state. The proposed IIA technique does not require any change of a pre-trained model, and only introduces a very small computational overhead for solving a number of quadratic optimization problems. Extensive experiments show that considerably better FID scores can be achieved by using IIA-EDM, IIA-DDIM, and IIA-DPM-Solver than the original counterparts when the neural function evaluation (NFE) is small (i.e., less than 25). Guoqiang Zhang 0003, Kenta Niwa, W. Bastiaan Kleijn |
ICLR | 1 |
| 2024 | SSFG: Stochastically Scaling Features and Gradients for Regularizing Graph Convolutional NetworksabstractGraph convolutional networks (GCNs) have been successfully applied in various graph-based tasks. In a typical graph convolutional layer, node features are updated by aggregating neighborhood information. Repeatedly applying graph convolutions can cause the oversmoothing issue, i.e., node features at deep layers converge to similar values. Previous studies have suggested that oversmoothing is one of the major issues that restrict the performance of GCNs. In this article, we propose a stochastic regularization method to tackle the oversmoothing problem. In the proposed method, we stochastically scale features and gradients (SSFG) by a factor sampled from a probability distribution in the training procedure. By explicitly applying a scaling factor to break feature convergence, the oversmoothing issue is alleviated. We show that applying stochastic scaling at the gradient level is complementary to that applied at the feature level to improve the overall performance. Our method does not increase the number of trainable parameters. When used together with ReLU, our SSFG can be seen as a stochastic ReLU activation function. We experimentally validate our SSFG regularization method on three commonly used types of graph networks. Extensive experimental results on seven benchmark datasets for four graph-based tasks demonstrate that our SSFG regularization is effective in improving the overall performance of the baseline graph networks. The code is available at https://github.com/vailatuts/SSFG-regularization. Haimin Zhang 0001, Min Xu 0001, Guoqiang Zhang 0003, Kenta Niwa |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Learning Graph Representations Through Learning and Propagating Edge FeaturesabstractGraph convolutional networks have achieved considerable success in various graph domain tasks. Recently, numerous types of graph convolutional networks have been developed. A typical rule for learning a node's feature in these graph convolutional networks is to aggregate node features from the node's local neighborhood. However, in these models, the interrelation information between adjacent nodes is not well-considered. This information could be helpful to learn improved node embeddings. In this article, we present a graph representation learning framework that generates node embeddings through learning and propagating edge features. Instead of aggregating node features from a local neighborhood, we learn a feature for each edge and update a node's representation by aggregating local edge features. The edge feature is learned from the concatenation of the edge's starting node feature, the input edge feature, and the edge's end node feature. Unlike node feature propagation-based graph networks, our model propagates different features from a node to its neighbors. In addition, we learn an attention vector for each edge in aggregation, enabling the model to focus on important information in each feature dimension. By learning and aggregating edge features, the interrelation between a node and its neighboring nodes is integrated in the aggregated feature, which helps learn improved node embeddings in graph representation learning. Our model is evaluated on graph classification, node classification, graph regression, and multitask binary graph classification on eight popular datasets. The experimental results demonstrate that our model achieves improved performance compared with a wide variety of baseline models. Haimin Zhang 0001, Jiahao Xia 0001, Guoqiang Zhang 0003, Min Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Lookahead Diffusion Probabilistic Models for Refining Mean EstimationabstractWe propose lookahead diffusion probabilistic models (LA-DPMs) to exploit the correlation in the outputs of the deep neural networks (DNNs) over subsequent timesteps in diffusion probabilistic models (DPMs) to refine the mean estimation of the conditional Gaussian distributions in the backward process. A typical DPM first obtains an estimate of the original data sample$x$by feeding the most recent state$z_{i}$and index$i$into the DNN model and then computes the mean vector of the conditional Gaussian distribution for$z_{i-1}$. We propose to calculate a more accurate estimate for$x$by performing extrapolation on the two estimates of$x$that are obtained by feeding ($z_{i+1},i+1$) and ($z_{i},i$) into the DNN model. The extrapolation can be easily integrated into the backward process of existing DPMs by introducing an additional connection over two consecutive timesteps, and fine-tuning is not required. Extensive experiments showed that plugging in the additional connection into DDPM, DDIM, DEIS, S-PNDM, and high-order DPM-Solvers leads to a significant performance gain in terms of Fréchet inception distance (FID) score. Our implementation is available at https://github.com/guoqiang-zhang-x/LA-DPM. Guoqiang Zhang 0003, Kenta Niwa, W. Bastiaan Kleijn |
CVPR | 1 |
| 2022 | GPCA: A Probabilistic Framework for Gaussian Process Embedded Channel AttentionabstractChannel attention mechanisms have been commonly applied in many visual tasks for effective performance improvement. It is able to reinforce the informative channels as well as to suppress the useless channels. Recently, different channel attention modules have been proposed and implemented in various ways. Generally speaking, they are mainly based on convolution and pooling operations. In this paper, we propose Gaussian process embedded channel attention (GPCA) module and further interpret the channel attention schemes in a probabilistic way. The GPCA module intends to model the correlations among the channels, which are assumed to be captured by beta distributed variables. As the beta distribution cannot be integrated into the end-to-end training of convolutional neural networks (CNNs) with a mathematically tractable solution, we utilize an approximation of the beta distribution to solve this problem. To specify, we adapt a Sigmoid-Gaussian approximation, in which the Gaussian distributed variables are transferred into the interval [0,1]. The Gaussian process is then utilized to model the correlations among different channels. In this case, a mathematically tractable solution is derived. The GPCA module can be efficiently implemented and integrated into the end-to-end training of the CNNs. Experimental results demonstrate the promising performance of the proposed GPCA module. Codes are available at https://github.com/PRIS-CV/GPCA. Jiyang Xie 0001, Zhanyu Ma, Dongliang Chang, Guoqiang Zhang 0003, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Advanced Dropout: A Model-Free Methodology for Bayesian Dropout OptimizationabstractDue to lack of data, overfitting ubiquitously exists in real-world applications of deep neural networks (DNNs). We propose advanced dropout, a model-free methodology, to mitigate overfitting and improve the performance of DNNs. The advanced dropout technique applies a model-free and easily implemented distribution with parametric prior, and adaptively adjusts dropout rate. Specifically, the distribution parameters are optimized by stochastic gradient variational Bayes in order to carry out an end-to-end training. We evaluate the effectiveness of the advanced dropout against nine dropout techniques on seven computer vision datasets (five small-scale datasets and two large-scale datasets) with various base models. The advanced dropout outperforms all the referred techniques on all the datasets. We further compare the effectiveness ratios and find that advanced dropout achieves the highest one on most cases. Next, we conduct a set of analysis of dropout rate characteristics, including convergence of the adaptive dropout rate, the learned distributions of dropout masks, and a comparison with dropout rate generation without an explicit distribution. In addition, the ability of overfitting prevention is evaluated and confirmed. Finally, we extend the application of the advanced dropout to uncertainty inference, network pruning, text classification, and regression. The proposed advanced dropout is also superior to the corresponding referred methods. Codes are available at https://github.com/PRIS-CV/AdvancedDropout. Jiyang Xie 0001, Zhanyu Ma, Jianjun Lei 0001, Guoqiang Zhang 0003, Jing-Hao Xue, Zheng-Hua Tan, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Asynchronous Decentralized Optimization With Implicit Stochastic Variance ReductionabstractA novel asynchronous decentralized optimization method that follows Stochastic Variance Reduction (SVR) is proposed. Average consensus algorithms, such as Decentralized Stochastic Gradient Descent (DSGD), facilitate distributed training of machine learning models. However, the gradient will drift within the local nodes due to statistical heterogeneity of the subsets of data residing on the nodes and long communication intervals. To overcome the drift problem, (i) Gradient Tracking-SVR (GT-SVR) integrates SVR into DSGD and (ii) Edge-Consensus Learning (ECL) solves a model constrained minimization problem using a primal-dual formalism. In this paper, we reformulate the update procedure of ECL such that it implicitly includes the gradient modification of SVR by optimally selecting a constraint-strength control parameter. Through convergence analysis and experiments, we confirmed that the proposed ECL with Implicit SVR (ECL-ISVR) is stable and approximately reaches the reference performance obtained with computation on a single-node using full data set. Kenta Niwa, Guoqiang Zhang 0003, W. Bastiaan Kleijn, Noboru Harada, Hiroshi Sawada, Akinori Fujino |
ICML | 2 |
| 2021 | Sentence Semantic Matching with Hierarchical CNN Based on Dimension-augmented RepresentationabstractAs a fundamental task in natural language processing, sentence semantic matching (SSM) is critical yet challenging due to difficulties in learning expressive sentence representation while capturing complex interactions between sentences. Recent work has shown the great potential of deep neural models in improving the performance of SSM task. However, existing work usually employs recurrent neural networks (RNNs) or 1D (one-dimensional) convolutional neural networks (CNNs) to learn sentence representation, leading to limited performance improvement. Benefiting from the multi-dimensional structure, 2D convolutional neural networks are expected to be more powerful to learn expressive sentence representation by capturing the implicit inter-sentence interactions and thus can further improve the performance of SSM. To this end, in this paper, we propose a novel sentence semantic matching model named Hierarchical CNN based on Dimension-augmented Representation (HiDR). In HiDR, first, bidirectional long short-term memory networks (LSTMs) are utilized to generate dimension-augmented representation for each of the input sentences; then, a hierarchical 2D CNN is devised to learn sentence representation while capturing the inter-sentence interactions, followed by a prediction layer based on sigmoid function to output the matching degree between sentences. To evaluate the performance of our proposed model, we conducted extensive experiments on two public real-world data sets. The empirical results show that HiDR has achieved remarkable results, which demonstrates either better or comparable performance w.r.t. BERT-based models. Rui Yu 0005, Wenpeng Lu, Yifeng Li 0001, Jiguo Yu, Guoqiang Zhang 0003, Xu Zhang 0053 |
IJCNN | 5 |
| 2021 | DS-UI: Dual-Supervised Mixture of Gaussian Mixture Models for Uncertainty Inference in Image RecognitionabstractThis paper proposes a dual-supervised uncertainty inference (DS-UI) framework for improving Bayesian estimation-based UI in DNN-based image recognition. In the DS-UI, we combine the classifier of a DNN, i.e., the last fully-connected (FC) layer, with a mixture of Gaussian mixture models (MoGMM) to obtain an MoGMM-FC layer. Unlike existing UI methods for DNNs, which only calculate the means or modes of the DNN outputs' distributions, the proposed MoGMM-FC layer acts as a probabilistic interpreter for the features that are inputs of the classifier to directly calculate the probabilities of them for the DS-UI. In addition, we propose a dual-supervised stochastic gradient-based variational Bayes (DS-SGVB) algorithm for the MoGMM-FC layer optimization. Unlike conventional SGVB and optimization algorithms in other UI methods, the DS-SGVB not only models the samples in the specific class for each Gaussian mixture model (GMM) in the MoGMM, but also considers the negative samples from other classes for the GMM to reduce the intra-class distances and enlarge the inter-class margins simultaneously for enhancing the learning ability of the MoGMM-FC layer in the DS-UI. Experimental results show the DS-UI outperforms the state-of-the-art UI methods in misclassification detection. We further evaluate the DS-UI in open-set out-of-domain/-distribution detection and find statistically significant improvements. Visualizations of the feature spaces demonstrate the superiority of the DS-UI. Codes are available at https://github.com/PRIS-CV/DS-UI. Jiyang Xie 0001, Zhanyu Ma, Jing-Hao Xue, Guoqiang Zhang 0003, Yinhe Zheng, Jun Guo 0002 |
IEEE Trans. Image Process. | 4 |
| 2020 | Intra-Correlation Encoding for Chinese Sentence Intention MatchingabstractSentence intention matching is vital for natural language understanding.Especially for Chinese sentence intention matching task, due to the ambiguity of Chinese words, semantic missing or semantic confusion are more likely to occur in the encoding process.Although the existing methods have enriched text representation through pre-trained word embedding to solve this problem, due to the particularity of Chinese text, different granularities of pre-trained word embedding will affect the semantic description of a piece of text.In this paper, we propose an effective approach that combines charactergranularity and word-granularity features to perform sentence intention matching, and we utilize soft alignment attention to enhance the local information of sentences on the corresponding levels.The proposed method can capture sentence feature information from multiple perspectives and correlation information between different levels of sentences.By evaluating on BQ and LCQMC datasets, our model has achieved remarkable results, and demonstrates better or comparable performance with BERT-based models. Xu Zhang 0053, Yifeng Li 0001, Wenpeng Lu, Ping Jian, Guoqiang Zhang 0003 |
COLING | 5 |
| 2020 | Projected Weight Regularization to Improve Neural Network GeneralizationabstractGeneralization of a deep neural network (DNN) is one major concern when employing the deep learning approach for solving practical problems. In this paper we propose a new technique, named projected weight regularization (PWR), to improve the generalization capacity of a DNN model. Consider a weight matrix W from a particular neural layer in the model. Our objective is to make the eigenvalues of the matrix product WWThave comparable or roughly the same magnitudes while allowing the DNN model to fit the training data sufficiently accurate. Intuitively speaking, by doing so, it would prevent the W matrix from matching the training data too well. Specifically, at each iteration, we first project the W matrix to a number of vectors along randomly generated directions. After that, we build an objective function of the projected vectors to regularize their behaviours towards comparable eigenvalue magnitudes of WWT. Experimental results on training VGG16 for CIFAR10 show that PWR combined with centered weight normalization (CWN) yields promising validation performance compared to orthonormal regularisation combined with CWN. Guoqiang Zhang 0003, Kenta Niwa, W. Bastiaan Kleijn |
ICASSP | 1 |
| 2020 | Object Detection and 3d Estimation Via an FMCW Radar Using a Fully Convolutional NetworkabstractThis paper considers object detection and 3D estimation using an FMCW radar. The state-of-the-art deep learning framework is employed instead of using traditional signal processing. In preparing the radar training data, the ground truth of an object orientation in 3D space is provided by conducting image analysis, of which the images are obtained through a coupled camera to the radar device. To ensure successful training of a fully convolutional network (FCN), we propose a normalization method, which is found to be essential to be applied to the radar signal before feeding into the neural network. The system after proper training is able to first detect the presence of an object in an environment. If it does, the system then further produces an estimation of its 3D position. Experimental results show that the proposed system can be successfully trained and employed for detecting a car and further estimating its 3D position in a noisy environment. Guoqiang Zhang 0003, Fabian Wenger |
ICASSP | 1 |
| 2020 | Edge-consensus Learning: Deep Learning on P2P Networks with Nonhomogeneous DataabstractAn effective Deep Neural Network (DNN) optimization algorithm that can use decentralized data sets over a peer-to-peer (P2P) network is proposed. In applications such as medical data analysis, the aggregation of data in one location may not be possible due to privacy issues. Hence, we formulate an algorithm to reach a global DNN model that does not require transmission of data among nodes. An existing solution for this issue is gossip stochastic gradient descend (SGD), which updates by averaging node models over a P2P network. However, in practical situations where the data are statistically heterogeneous across the nodes and/or where communication is asynchronous, gossip SGD often gets trapped in local minimum since the model gradients are noticeably different. To overcome this issue, we solve a linearly constrained DNN cost minimization problem, which results in variable update rules that restrict differences among all node models. Our approach can be based on the Primal-Dual Method of Multipliers (PDMM) or the Alternating Direction Method of Multiplier (ADMM), but the cost function is linearized to be suitable for deep learning. It facilitates asynchronous communication. The results of our numerical experiments using CIFAR-10 indicate that the proposed algorithms converge to a global recognition model even though statistically heterogeneous data sets are placed on the nodes. Kenta Niwa, Noboru Harada, Guoqiang Zhang 0003, W. Bastiaan Kleijn |
KDD | 3 |
| 2020 | Chinese Sentence Semantic Matching Based on Multi-Granularity Fusion Model
Xu Zhang 0053, Wenpeng Lu, Guoqiang Zhang 0003, Shoujin Wang |
PAKDD (2) | 3 |
| 2020 | Chinese medical question answer selection via hybrid models based on CNN and GRU
Wenpeng Lu, Weihua Ou, Guoqiang Zhang 0003, Xu Zhang 0053, Jinyong Cheng, Weiyu Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2020 | Microphone Array Wiener Post Filtering Using Monotone Operator SplittingabstractFor array-based acoustic source enhancement, variants of multi-channel Wiener filters are commonly used. The approach includes a Wiener post-filter that requires the simultaneous estimation of the power spectral density (PSD) of the target source and of noise sources for each time-frame. Conventional methods generally do not exploit prior knowledge, such as sparsity of the source, in solving this simultaneous estimation problem. We show that, for common scenarios, the simultaneous PSD estimation with consideration of prior knowledge can be formulated as a convex optimization problem with linear constraints. We use monotone operator splitting (MOS) to solve the constrained optimization problem. Our experiments confirm that the proposed method improves the accuracy of the noise PSD estimation, and that the resulting enhanced target signal is of higher quality. Kenta Niwa, Hironobu Chiba, Noboru Harada, Guoqiang Zhang 0003, W. Bastiaan Kleijn |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Period adding bifurcations in dynamic pricing processesabstractPrice information enables consumers to anticipate a price and to make purchasing decisions based on their price expectations, which are critical for agents with pricing decisions or price regulations. A company with pricing decisions can aim to optimise the short-term or the long-term revenue, each of which leads to different pricing strategies thereby different price expectations. Two key ingredients play important roles in the choosing of the short-term or the long-term optimisation objectives: the maximal revenue and the robustness of the chosen pricing strategy against market volatility. However the robustness is rarely identified in a volatile market. Here, we investigate the robustness of optimal pricing strategies with the short-term or long-term optimisation objectives through the analysis of nonlinear dynamics of price expectations. Bifurcation diagrams and period diagrams are introduced to compare the change in dynamics of the optomal pricing strategies. Our results highlight that period adding bifurcations occur during the dynamic pricing processes studied. These bifurcations would challenge the robustness of an optimal pricing strategy. The consideration of the long-term revenue allows a company to charge a higher price, which in turn increases the revenue. However, the consideration of the short-term revenue can reduce the occurrence of period adding bifurcations, contributing to a robust pricing strategy. For a company, this strategy is a robust guarantee of optimal revenue in a volatile market; for consumers, this strategy avoids rapid changes in price and reduce their dissatisfaction of price variations. Shuixiu Lu, Sebastian Oberst, Guoqiang Zhang 0003, Zongwei Luo |
CIFEr | 3 |
| 2019 | Fast Edge-consensus Computing Based on Bregman Monotone Operator SplittingabstractEdge-consensus computing is a framework to optimize a global cost function when distributed nodes observe distinct data sets. The distributed primal-dual method of multipliers (PDMM) and distributed alternating direction method of multipliers (ADMM) find network-global optima for edge-consensus algorithms by exchanging variables rather than data sets among the nodes. Since the distributed PDMM follows traditional Peaceman-Rachford splitting, it has a faster convergence rate than the distributed ADMM. To further speed up the convergence rate, we propose a new edge-consensus computing algorithm based on Bregman Peaceman-Rachford splitting. In traditional Peaceman-Rachford splitting, the variable update is defined based on a Euclidean metric and the convergence rate and is a form of first-order gradient descent. By generalizing the metric to a Bregman divergence and designing the divergence adaptively, our fast edge-consensus computing algorithm corresponds to the Newton or an accelerated gradient descent method. The results of our experiments confirm that the proposed algorithm can significantly improve the convergence rate of edge-consensus computing over state-of-the-art algorithms. Kenta Niwa, Guoqiang Zhang 0003, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2019 | Decentralized Two-Channel Active Noise Control for Single Frequency by Shaping Matrix EigenvaluesabstractIn an active noise control (ANC) system, computational complexity is one major concern when designing practical control algorithms. For an ANC system with multiple secondary sources and error microphones, one approach to reducing computational complexity is to apply a decentralized control scheme rather than centralized approaches. A decentralized scheme attempts to control a number of small-size ANC subsystems independently. In this paper, we consider the decentralized control of a two-channel ANC system tackling a noise disturbance in the frequency domain, where each channel consists of one secondary source and one error microphone. We propose a decentralized control method that is able to achieve the same noise reduction performance as the centralized controller with guaranteed convergence. The key step in designing the control method is to properly shape the eigenvalues of a matrix that models the two-channel secondary paths for each frequency index. Guoqiang Zhang 0003, Jiancheng Tao, Xiaojun Qiu, Ian S. Burnett |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Training Deep Neural Networks via Optimization Over GraphsabstractIn this work, we propose to train a deep neural network by distributed optimization over a graph. Two nonlinear functions are considered: the rectified linear unit (ReLU) and a linear unit with both lower and upper cutoffs (DCutLU). The problem reformulation over a graph is realized by explicitly representing ReLU or DCutLU using a set of slack variables. We then apply the alternating direction method of multipliers (ADMM) to update the weights of the network layer-wise by solving subproblems of the reformulated problem. Empirical results suggest that the ADMM-based method is less sensitive to overfitting than the stochastic gradient descent (SGD) and Adam methods. Guoqiang Zhang 0003, W. Bastiaan Kleijn |
ICASSP | 1 |
| 2016 | On simplifying the primal-dual method of multipliersabstractRecently, the primal-dual method of multipliers (PDMM) has been proposed to solve a convex optimization problem defined over a general graph. In this paper, we consider simplifying PDMM for a subclass of the convex optimization problems. This subclass includes the consensus problem as a special form. By using algebra, we show that the update expressions of PDMM can be simplified significantly. We then evaluate PDMM for training a support vector machine (SVM). The experimental results indicate that PDMM converges considerably faster than the alternating direction method of multipliers (ADMM). Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 1 |
| 2015 | Bi-alternating direction method of multipliers over graphsabstractIn this paper, we extend the bi-alternating direction method of multipliers (BiADMM) designed on a graph of two nodes to a graph of multiple nodes. In particular, we optimize a sum of convex functions defined over a general graph, where every edge carries a linear equality constraint. In designing the new algorithm, an augmented primal-dual Lagrangian function is carefully constructed which naturally captures the associated graph topology. We show that under both the synchronous and asynchronous updating schemes, the extended BiADMM has the convergence rate of O(1/K) (where K denotes the iteration index) for general closed, proper and convex functions. As an example, we apply the new algorithm for distributed averaging. Experimental results show that the new algorithm remarkably outperforms the state-of-the-art methods. Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 1 |
| 2014 | On the convergence rate of the bi-alternating direction method of multipliersabstractIn this paper, we analyze the convergence rate of the bi-alternating direction method of multipliers (BiADMM). Differently from ADMM that optimizes an augmented Lagrangian function, Bi-ADMM optimizes an augmented primal-dual Lagrangian function. The new function involves both the objective functions and their conjugates, thus incorporating more information of the objective functions than the augmented Lagrangian used in ADMM. We show that BiADMM has a convergence rate of O(K-1) (K denotes the number of iterations) for general convex functions. We consider the lasso problem as an example application. Our experimental results show that BiADMM outperforms not only ADMM, but fast-ADMM as well. Guoqiang Zhang 0003, Richard Heusdens, W. Bastiaan Kleijn |
ICASSP | 1 |
| 2014 | Convergence of Min-Sum-Min Message-Passing for Quadratic Optimization
Guoqiang Zhang 0003, Richard Heusdens |
ECML/PKDD (3) | 1 |
| 2013 | Bi-alternating direction method of multipliersabstractThe alternating-direction method of multipliers (ADMM) has been widely applied in the field of distributed optimization and statistic learning. ADMM iteratively approaches the saddle point of an augmented Lagrangian function by performing three updates per-iteration. In this paper, we propose a bi-alternating direction method of multipliers (BiADMM) that iteratively minimizes an augmented bi-conjugate function. As a result, the convergence of BiADMM is naturally established. Unlike ADMM that always involves three updates per iteration, BiADMM opens up an avenue to perform either two or three updates per iteration, depending on the functional construction. As an application, we consider applying BiADMM for the lasso problem. Experimental results demonstrate the effectiveness of our new method. Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 1 |
| 2013 | Simplified alternating-direction message passing for dual MAP LP-relaxationabstractThe approximate MAP inference over (factor) graphic models is of great importance in many applications. Due to its simplicity, linear-programming (LP) relaxation has become one of the most popular approaches to approximate MAP. In this paper, we propose a new message passing algorithm for the MAP LP-relaxation problem by using the alternating-direction method of multipliers (ADMM). At each iteration, the new algorithm performs two layers of optimization sequentially, that is node-oriented optimization and factor-oriented optimization. On the other hand, the recently proposed augmented dual LP (ADLP) algorithm, also based on the ADMM, has to perform three layers of optimization. We refer to our new algorithm as the simplified ADLP (SiADLP) algorithm. The design of the SiADLP algorithm stems from a new formulation for the dual LP problem. Experimental results show that the SiADLP algorithm outperforms the ADLP method. Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 1 |
| 2013 | Proximal alternating-direction message-passing for MAP LP relaxationabstractLinear programming (LP) relaxation for MAP inference over (factor) graphic models is one of the fundamental problems in machine learning. In this paper, we propose a new message-passing algorithm for the MAP LP-relaxation by using the proximal alternating-direction method of multipliers (PADMM). At each iteration, the new algorithm performs two layers of optimization, that is node-oriented optimization and factor-oriented optimization. On the other hand, the recently proposed augmented primal LP (APLP) algorithm, based on the ADMM, has to perform three layers of optimization. Our algorithm simplifies the APLP algorithm by removing one layer of optimization, thus reducing the computational complexities and further accelerating the convergence rate. We refer to our new algorithm as the proximal alternating-direction (PAD) algorithm. Experimental results confirm that the PAD algorithm indeed converges faster than the APLP method. Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 1 |
| 2012 | Linear coordinate-descent message-passing for quadratic optimizationabstractIn this paper we propose a new message-passing algorithm for quadratic optimization. The design of the new algorithm is based on linear coordinate-descent between neighboring nodes. The updating messages are in a form of linear functions as compared to the min-sum algorithm of which the messages are in a form of quadratic functions. Therefore, the linear coordinate-descent (LiCD) algorithm has simpler updating rules than the min-sum algorithm. It is shown that when the quadratic matrix is walk-summable, the LiCD algorithm converges. As an application, the LiCD algorithm is utilized in solving general linear systems. The performance of the LiCD algorithm is found empirically to be comparable to that of the min-sum algorithm, but at lower complexity in terms of computation and storage. Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 1 |
| 2012 | Generalized linear coordinate-descent message-passing for convex optimizationabstractIn this paper we propose a generalized linear coordinate-descent (GLiCD) algorithm for a class of unconstrained convex optimization problems. The considered objective function can be decomposed into edge-functions and node-functions of a graphical model. The messages of the GLiCD algorithm are in a form of linear functions, as compared to the min-sum algorithm of which the form of messages depends on the objective function. Thus, the implementation of the GLiCD algorithm is much simpler than that of the min-sum algorithm. A theorem is stated according to which the algorithm converges to the optimal solution if the objective function satisfies a diagonal-dominant condition. As an application, the GLiCD algorithm is exploited in solving the averaging problem in sensor networks, where the performance is compared to that of the min-sum algorithm. Guoqiang Zhang 0003, Richard Heusdens |
ICASSP | 1 |
| 2012 | Convergence of generalized linear coordinate-descent message-passing for quadratic optimizationabstractWe study the generalized linear coordinate-descent (GLiCD) algorithm for the quadratic optimization problem. As an extension of the linear coordinate-descent (LiCD) algorithm, the GLiCD algorithm incorporates feedback from last iteration in generating new messages. We show that if the amount of feedback signal from last iteration is above a threshold and the GLiCD algorithm converges, it computes the optimal solution. Based on the result, we further show that if the feedback signal is large enough, the GLiCD algorithm is guaranteed to converge. Guoqiang Zhang 0003, Richard Heusdens |
ISIT | 1 |
| 2012 | Linear Coordinate-Descent Message Passing for Quadratic OptimizationabstractIn this letter, we propose a new message-passing algorithm for quadratic optimization. The design of the new algorithm is based on linear coordinate descent between neighboring nodes. The updating messages are in a form of linear functions as compared to the min-sum algorithm of which the messages are in a form of quadratic functions. As a result, the linear coordinate-descent (LiCD) algorithm transmits only one parameter per message as opposed to the min-sum algorithm, which transmits two parameters per message. We show that when the quadratic matrix is walk-summable, the LiCD algorithm converges. By taking the LiCD algorithm as a subroutine, we also fix the convergence issue for a general quadratic matrix. The LiCD algorithm works in either a synchronous or asynchronous message-passing manner. Experimental results show that for a general graph with multiple cycles, the LiCD algorithm has comparable convergence speed to the min-sum algorithm, thereby reducing the number of parameters to be transmitted and the computational complexity. Guoqiang Zhang 0003, Richard Heusdens |
Neural Comput. | 1 |
| 2011 | Bounding the Rate Region of the Two-Terminal Vector Gaussian CEO ProblemabstractThe rate region of the two-terminal vector Gaussian CEO problem is studied. A lower bound on the rate region is derived. It is obtained by lower-bounding a weighted sum rate for each supporting hyper plane of the rate region. The bound is in the form of a closed-form expression rather than the form of an optimization problem. Guoqiang Zhang 0003, W. Bastiaan Kleijn |
DCC | 1 |
| 2011 | High-Rate Analysis of Symmetric L-Channel Multiple Description CodingabstractThis paper studies the tight rate-distortion bound for L-channel symmetric multiple-description coding of a scalar Gaussian source with two levels of receivers. Each of the first-level receivers obtains κ of the L descriptions (κ < L). The second-level receiver obtains all L descriptions. We find that if the central distortion (corresponding to the second-level receiver) is much smaller than the side distortion (corresponding to the first-level receivers), the product of a function of the side distortions and the central distortion is asymptotically independent of the redundancy between the descriptions. Using this property, we analyze the asymptotic behavior of a practical multiple-description lattice vector quantizer (MDLVQ). Our analysis includes the treatment of the MDLVQ system from a new geometric viewpoint, which results in an expression for the side distortions using the normalized second moment of a sphere of higher dimensionality than the quantization space. The expression of the distortion product derived from the lower bound is then applied as a criterion to assess the performance loss of the considered MDLVQ system. In principle, the efficiency of other practical MD systems can also be evaluated using the derived distortion product. Guoqiang Zhang 0003, Jan Østergaard, Janusz Klejsa, W. Bastiaan Kleijn |
IEEE Trans. Commun. | 1 |
| 2011 | A Stratified Approach for Camera Calibration Using SpheresabstractThis paper proposes a stratified approach for camera calibration using spheres. Previous works have exploited epipolar tangents to locate frontier points on spheres for estimating the epipolar geometry. It is shown in this paper that other than the frontier points, two additional point features can be obtained by considering the bitangent envelopes of a pair of spheres. A simple method for locating the images of such point features and the sphere centers is presented. An algorithm for recovering the fundamental matrix in a plane plus parallax representation using these recovered image points and the epipolar tangents from three spheres is developed. A new formulation of the absolute dual quadric as a cone tangent to a dual sphere with the plane at infinity being its vertex is derived. This allows the recovery of the absolute dual quadric, which is used to upgrade the weak calibration to a full calibration. Experimental results on both synthetic and real data are presented, which demonstrate the feasibility and the high precision achieved by our proposed algorithm. Kwan-Yee Kenneth Wong, Guoqiang Zhang 0003, Zhihu Chen |
IEEE Trans. Image Process. | 2 |
| 2010 | Bounding the Rate Region of Vector Gaussian Multiple Descriptions with Individual and Central ReceiversabstractThe problem of the rate region of the vector Gaussian multiple description with individual and central quadratic distortion constraints is studied. We have two main contributions. First, a lower bound on the rate region is derived. The bound is obtained by lower-bounding a weighted sum rate for each supporting hyperplane of the rate region. Second, the rate region for the scenario of the scalar Gaussian source is fully characterized by showing that the lower bound is tight. The optimal weighted sum rate for each supporting hyperplane is obtained by solving a single maximization problem. This is contrary to existing results, which require solving a min-max optimization problem. Guoqiang Zhang 0003, W. Bastiaan Kleijn, Jan Østergaard |
DCC | 1 |
| 2009 | Analysis of K-Channel Multiple Description QuantizationabstractThis paper studies the tight rate-distortion bound for K-channel symmetric multiple-description coding for a memory less Gaussian source. We find that the product of a function of the individual side distortions (for single received descriptions) and the central distortion (for K received descriptions) is asymptotically independent of the redundancy among the descriptions. Using this property, we analyze the asymptotic behaviors of two different practical multiple-description lattice vector quantizers (MDLVQ). Our analysis includes the treatment of a MDLVQ system from a new geometric viewpoint, which results in an expression for the side distortions using the normalized second moment of a sphere of higher dimensionality than the quantization space. The expression of the distortion product derived from the lower bound is then applied as a criterion to assess the performance losses of the considered MDLVQ systems. Guoqiang Zhang 0003, Janusz Klejsa, W. Bastiaan Kleijn |
DCC | 1 |
| 2008 | Autoregressive model-based speech packet-loss concealmentabstractWe study packet-loss concealment for speech based on autoregressive modeling using a rigorous minimum mean square error (MMSE) approach. The effect of the model estimation error on predicting the missing segment is studied and an upper bound on the mean square error is derived. Our experiments show that the upper bound is tight when the estimation error is less than the signal variance. We also consider the usage of perceptual weighting on prediction to improve speech quality. A rigorous argument is presented to show that perceptual weighting is not useful in this context. We create simple and practical MMSE-based systems using two signal models: a basic model capturing the short-term correlation and a more sophisticated model that also captures the long-term correlation. Subjective quality comparison tests show that the proposed MMSE-based system provides state-of-the-art performance. Guoqiang Zhang 0003, W. Bastiaan Kleijn |
ICASSP | 1 |
| 2008 | 1D Camera Geometry and Its Application to the Self-Calibration of Circular Motion SequencesabstractThis paper proposes a novel method for robustly recovering the camera geometry of an uncalibrated image sequence taken under circular motion. Under circular motion, all the camera centers lie on a circle and the mapping from the plane containing this circle to the horizon line observed in the image can be modelled as a 1D projection. A 2 x 2 homography is introduced in this paper to relate the projections of the camera centers in two 1D views. It is shown that the two imaged circular points of the motion plane and the rotation angle between the two views can be derived directly from such a homography. This way of recovering the imaged circular points and rotation angles is intrinsically a multiple view approach, as all the sequence geometry embedded in the epipoles is exploited in the estimation of the homography for each view pair. This results in a more robust method compared to those computing the rotation angles using adjacent views only. The proposed method has been applied to self-calibrate turntable sequences using either point features or silhouettes, and highly accurate results have been achieved. Kwan-Yee Kenneth Wong, Guoqiang Zhang 0003, Chen Liang 0004, Hui Zhang 0062 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Camera Calibration from Images of SpheresabstractThis paper introduces a novel approach for solving the problem of camera calibration from spheres. By exploiting the relationship between the dual images of spheres and the dual image of the absolute conic (IAC), it is shown that the common pole and polar with regard to the conic images of two spheres are also the pole and polar with regard to the IAC. This provides two constraints for estimating the IAC and, hence, allows a camera to be calibrated from an image of at least three spheres. Experimental results show the feasibility of the proposed approach. Hui Zhang 0062, Kwan-Yee Kenneth Wong, Guoqiang Zhang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2006 | 1D Camera Geometry and Its Application to Circular Motion EstimationabstractThis paper describes a new and robust method for estimating circular motion geometry from an uncalibrated image sequence. Under circular motion, all the camera centers lie on a circle, and the mapping of the plane containing this circle to the horizon line in the image can be modelled as a 1D projection. A 2×2 homography is introduced in this paper to relate the projections of the camera centers in two 1D views. It is shown that the two imaged circular points and the rotation angle between the two views can be derived directly from the eigenvectors and eigenvalues of such a homography respectively. The proposed 1D geometry can be nicely applied to circular motion estimation using either point correspondences or silhouettes. The method introduced here is intrinsically a multiple view approach as all the sequence geometry embedded in the epipoles is exploited in the computation of the homography for a view pair. This results in a robust method which gives accurate estimated rotation angles and imaged circular points. Experimental results are presented to demonstrate the simplicity and applicability of the new method. Guoqiang Zhang 0003, Hui Zhang 0062, Kwan-Yee Kenneth Wong |
BMVC | 1 |
| 2006 | Motion Estimation from SpheresabstractThis paper addresses the problem of recovering epipolar geometry from spheres. Previous works have exploited epipolar tangencies induced by frontier points on the spheres for motion recovery. It will be shown in this paper that besides epipolar tangencies, N^2 point features can be extracted from the apparent contours of the N spheres whenN \gt 2. An algorithm for recovering the fundamental matrices from such point features and the epipolar tangencies from 3 or more spheres is developed, with the point features providing a homography over the view pairs and the epipolar tangencies determining the epipoles. In general, there will be two solutions to the locations of the epipoles. One of the solutions corresponds to the true camera configuration, while the other corresponds to a mirrored configuration. Several methods are proposed to select the right solution. Experiments on using 3 and 4 spheres demonstrate that our algorithm can be carried out easily and can achieve a high precision. Guoqiang Zhang 0003, Kwan-Yee Kenneth Wong |
CVPR (1) | 1 |
| 2005 | Auto-Calibration and Motion Recovery from Silhouettes for Turntable SequencesabstractThis paper addresses the problem of structure and motion from silhou-ettes for turntable sequences. Previous works have exploited corresponding points induced by epipolar tangencies to estimate the image invariants un-der turntable motion and recover the epipolar geometry. In these approaches, however, camera intrinsics are needed in order to obtain Euclidean motion and reconstruction. This paper proposes a novel approach to precisely esti-mate the image invariants and the rotation angles in the absence of the camera intrinsics, and to perform auto-calibration. By exploiting a special parame-terization of the epipoles, it is shown that the imaged circular points can be formulated in terms of the image invariants. A fixed scalar κ, introduced to account for the different scales in the homogeneous representations of the image invariants used in the parameterizations, is found crucial in both cali-bration and motion estimation. Given the image invariants, namely the hori-zon, the imaged rotation axis and its orthogonal vanishing point, this scalar can be determined from the epipoles in an image triplet. A robust method for estimating κ is proposed and the rotation angles can be recovered using this estimated value of κ. All the estimated variables are then refined using bundle-adjustment and auto-calibration is performed using the imaged circu-lar points, the imaged rotation axis and the associated vanishing point. This allows the recovery of the full camera positions and orientations, and hence Euclidean reconstruction. Experimental results demonstrate the simplicity of this novel approach and the high precision in the estimated motion and reconstruction. 1 Hui Zhang 0062, Guoqiang Zhang 0003, Kwan-Yee Kenneth Wong |
BMVC | 2 |
| 2005 | Camera calibration with spheres: linear approachesabstractThis paper addresses the problem of camera calibration from spheres. By studying the relationship between the dual images of spheres and that of the absolute conic, a linear solution has been derived from a recently proposed non-linear semi-definite approach. However, experiments show that this approach is quite sensitive to noise. In order to overcome this problem, a second approach has been proposed, where the orthogonal calibration relationship is obtained by regarding any two spheres as a surface of revolution. This allows a camera to be fully calibrated from an image of three spheres. Besides, a conic homography is derived from the imaged spheres, and from its eigenvectors the orthogonal invariants can be computed directly. Experiments on synthetic and real data show the practicality of such an approach. Hui Zhang 0062, Guoqiang Zhang 0003, Kwan-Yee Kenneth Wong |
ICIP (2) | 2 |