VLDB 2026 Research / reviewers in the wild / expert
Hui Qian 0001
dblp:66/5293
· DBLP profile ↗
53ranked-venue papers
0as first author
26since 2021 · last 2026
0000-0003-0293-2656ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 18 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mass Concept Erasure in Diffusion Models with Concept HierarchyabstractThe success of diffusion models has raised concerns about the generation of unsafe or harmful content, prompting concept erasure approaches that fine-tune modules to suppress specific concepts while preserving general generative capabilities. However, as the number of erased concepts grows, these methods often become inefficient and ineffective, since each concept requires a separate set of fine-tuned parameters and may degrade the overall generation quality. In this work, we propose a supertype-subtype concept hierarchy that organizes erased concepts into a parent–child structure. Each erased concept is treated as a child node, and semantically related concepts (e.g., macaw, and bald eagle) are grouped under a shared parent node, referred to as a supertype concept (e.g., bird). Rather than erasing concepts individually, we introduce an effective and efficient group-wise suppression method, where semantically similar concepts are grouped and erased jointly by sharing a single set of learnable parameters. During the erasure phase, standard diffusion regularization is applied to preserve denoising process in unmasked regions. To mitigate the degradation of supertype generation caused by excessive erasure of semantically related subtypes, we propose a novel method called Supertype-Preserving Low-Rank Adaptation (SuPLoRA), which encodes the supertype concept information in the frozen down-projection matrix and updates only the up-projection matrix during erasure. Theoretical analysis demonstrates the effectiveness of SuPLoRA in mitigating generation performance degradation. We construct a more challenging benchmark that requires simultaneous erasure of concepts across diverse domains, including celebrities, objects, and pornographic content. Comprehensive experiments demonstrate that our method achieves a superior balance between effective multi-concept erasure and the preservation of desirable generative performance. Jiahang Tu, Ye Li 0043, Hanbin Zhao, Chao Zhang 0029, Hui Qian 0001 |
AAAI | 6 |
| 2026 | CE-SDWV: Effective and Efficient Concept Erasure for Text-to-Image Diffusion Models via a Semantic-Driven Word Vocabulary
Jiahang Tu, Jiahua Dong 0001, Hanbin Zhao, Chao Zhang 0001, Nicu Sebe, Hui Qian 0001 |
Int. J. Comput. Vis. | 7 |
| 2026 | IAP: Improving Continual Learning of Vision-Language Models via Instance-Aware PromptingabstractRecent pre-trained vision-language models (PT-VLMs) often face a Multi-Domain Task Incremental Learning (MTIL) scenario in practice, where several classes and domains of multi-modal tasks are arrive incrementally. Without access to previously seen tasks and unseen tasks, memory-constrained MTIL suffers from forward and backward forgetting. To alleviate the above challenges, parameter-efficient fine-tuning techniques (PEFT), such as prompt tuning, are employed to adapt the PT-VLM to the diverse incrementally learned tasks. To achieve effective new task adaptation, existing methods only consider the effect of PEFT strategy selection, but neglect the influence of PEFT parameter setting (e.g., prompting). In this paper, we tackle the challenge of optimizing prompt designs for diverse tasks in MTIL and propose an Instance-Aware Prompting (IAP) framework. Specifically, our Instance-Aware Gated Prompting (IA-GP) strategy enhances adaptation to new tasks while mitigating forgetting by adaptively assigning prompts across transformer layers at the instance level. Our Instance-Aware Class-Distribution-Driven Prompting (IA-CDDP) improves the task adaptation process by determining an accurate task-label-related confidence score for each instance. Experimental evaluations across 11 datasets, using three performance metrics, demonstrate the effectiveness of our proposed method. The source codes are available at https://github.com/FerdinandZJU/IAP. Hao Fu 0023, Hanbin Zhao, Jiahua Dong 0001, Henghui Ding, Chao Zhang 0001, Hui Qian 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | TextToucher: Fine-Grained Text-to-Touch GenerationabstractTactile sensation plays a crucial role in the development of multi-modal large models and embodied intelligence. To collect tactile data with minimal cost as possible, a series of studies have attempted to generate tactile images by vision-to-touch image translation. However, compared to text modality, visual modality-driven tactile generation cannot accurately depict human tactile sensation. In this work, we analyze the characteristics of tactile images in detail from two granularities: object-level (tactile texture, tactile shape), and sensor-level (gel status). We model these granularities of information through text descriptions and propose a fine-grained Text-to-Touch generation method (TextToucher) to generate high-quality tactile samples. Specifically, we introduce a multimodal large language model to build the text sentences about object-level tactile information and employ a set of learnable text prompts to represent the sensor-level tactile information. To better guide the tactile generation process with the built text information, we fuse the dual grains of text information and explore various dual-grain text conditioning methods within the diffusion transformer architecture. Furthermore, we propose a Contrastive Text-Touch Pre-training (CTTP) metric to precisely evaluate the quality of text-driven generated tactile data. Extensive experiments demonstrate the superiority of our TextToucher method. Jiahang Tu, Hao Fu 0023, Fengyu Yang 0003, Hanbin Zhao, Chao Zhang 0029, Hui Qian 0001 |
AAAI | 6 |
| 2025 | FG-OrIU: Towards Better Forgetting via Feature-Gradient Orthogonality for Incremental Unlearning
Jiahang Tu, Mintong Kang, Hanbin Zhao, Chao Zhang 0001, Hui Qian 0001 |
ICCV | 6 |
| 2025 | Unleashing High-Quality Image Generation in Diffusion Sampling Using Second-Order Levenberg-Marquardt-LangevinabstractThe diffusion models (DMs) have demonstrated the remarkable capability of generating images via learning the noised score function of data distribution. Current DM sampling techniques typically rely on first-order Langevin dynamics at each noise level, with efforts concentrated on refining inter-level denoising strategies. While leveraging additional second-order Hessian geometry to enhance the sampling quality of Langevin is a common practice in Markov chain Monte Carlo (MCMC), the naive attempts to utilize Hessian geometry in high-dimensional DMs lead to quadratic-complexity computational costs, rendering them non-scalable. In this work, we introduce a novel Levenberg-Marquardt-Langevin (LML) method that approximates the diffusion Hessian geometry in a training-free manner, drawing inspiration from the celebrated Levenberg-Marquardt optimization algorithm. Our approach introduces two key innovations: (1) A low-rank approximation of the diffusion Hessian, leveraging the DMs' inherent structure and circumventing explicit quadratic-complexity computations; (2) A damping mechanism to stabilize the approximated Hessian. This LML approximated Hessian geometry enables the diffusion sampling to execute more accurate steps and improve the image generation quality. We further conduct a theoretical analysis to substantiate the approximation error bound of low-rank approximation and the convergence property of the damping mechanism. Extensive experiments across multiple pretrained DMs validate that the LML method significantly improves image generation quality, with negligible computational overhead. Fangyikang Wang, Hubery Yin, Shaobin Zhuang, Huminhao Zhu, Yanlong Tang, Chao Zhang 0029, Hanbin Zhao, Hui Qian 0001, Chen Li 0031 |
ICCV | 11 |
| 2025 | Efficiently Access Diffusion Fisher: Within the Outer Product Span SpaceabstractRecent Diffusion models (DMs) advancements have explored incorporating the second-order diffusion Fisher information (DF), defined as the negative Hessian of log density, into various downstream tasks and theoretical analysis.
However, current practices typically approximate the diffusion Fisher by applying auto-differentiation to the learned score network. This black-box method, though straightforward, lacks any accuracy guarantee and is time-consuming.
In this paper, we show that the diffusion Fisher actually resides within a space spanned by the outer products of score and initial data.
Based on the outer-product structure, we develop two efficient approximation algorithms to access the trace and matrix-vector multiplication of DF, respectively.
These algorithms bypass the auto-differentiation operations with time-efficient vector-product calculations.
Furthermore, we establish the approximation error bounds for the proposed algorithms.
Experiments in likelihood evaluation and adjoint optimization demonstrate the superior accuracy and reduced computational cost of our proposed algorithms.
Additionally, based on the novel outer-product formulation of DF, we design the first numerical verification experiment for the optimal transport property of the general PF-ODE deduced map. Fangyikang Wang, Hubery Yin, Shaobin Zhuang, Huminhao Zhu, Chao Zhang 0029, Hanbin Zhao, Hui Qian 0001, Chen Li 0031 |
ICML | 9 |
| 2025 | Rebalancing Return Coverage for Conditional Sequence Modeling in Offline Reinforcement LearningabstractRecent advancements in offline reinforcement learning (RL) have underscored the capabilities of conditional sequence modeling (CSM), a paradigm that models the action distribution conditioned on both historical trajectories and target returns associated with each state. However, due to the imbalanced return distribution caused by suboptimal datasets, CSM is grappling with a serious distributional shift problem when conditioning on high returns. While recent approaches attempt to empirically tackle this challenge through return rebalancing techniques such as weighted sampling and value-regularized supervision, the relationship between return rebalancing and the performance of CSM methods is not well understood. In this paper, we reveal that both expert-level and full-spectrum return-coverage critically influence the performance and sample efficiency of CSM policies. Building on this finding, we devise a simple yet effective return-coverage rebalancing mechanism that can be seamlessly integrated into common CSM frameworks, including the most widely used one, Decision Transformer (DT). The resulting CSM algorithm, referred to as Return-rebalanced Value-regularized Decision Transformer (RVDT), integrates both implicit and explicit return-coverage rebalancing mechanisms, and achieves state-of-the-art performance in the D4RL experiments. Wensong Bai, Chufan Chen, Yichao Fu, Qihang Xu, Chao Zhang 0029, Hui Qian 0001 |
NeurIPS | 6 |
| 2025 | Efficient projection-free online convex optimization using stochastic gradients
Jiahao Xie 0001, Chao Zhang 0029, Zebang Shen, Hui Qian 0001 |
Mach. Learn. | 4 |
| 2025 | PECTP: Parameter-Efficient Cross-Task Prompts for Incremental Vision TransformerabstractIncremental Learning (IL) aims to learn deep models on sequential tasks continually, where each new task includes a batch of new classes and deep models have no access to task ID information at the inference time. Recent vast pre-trained models (PTMs) have achieved outstanding performance by prompt technique in practical IL without the old samples (rehearsal-free) and with a memory constraint (memory-constrained): Prompt-extending and Prompt-fixed methods. However, prompt-extending methods need a large memory buffer to maintain an ever-expanding prompt pool and meet an extra challenging prompt selection problem. Prompt-fixed methods only learn a single set of prompts on one of the incremental tasks and can not handle all the incremental tasks effectively. To achieve a good balance between the memory cost and the performance on all the tasks, we propose a Parameter-Efficient Cross-Task Prompt (PECTP) framework with Prompt Retention Module (PRM) and classifier Head Retention Module (HRM). To make the final learned prompts effective on all incremental tasks, PRM constrains the evolution of cross-task prompts’ parameters from Outer Prompt Granularity and Inner Prompt Granularity. Besides, we employ HRM to inherit old knowledge in the previously learned classifier heads to facilitate the cross-task prompts’ generalization ability. Extensive experiments show the effectiveness of our method. The source codes are available at https://github.com/RAIAN08/PECTP. Hanbin Zhao, Chao Zhang 0001, Jiahua Dong 0001, Henghui Ding, Yu-Gang Jiang 0001, Hui Qian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Visuo-Tactile Class-Incremental LearningabstractThe ability to associate sight with touch is essential for human and robot agents to understand material properties and to interact with the physical world. In the real-world scenarios, the robot agents often operate in a dynamically changing environment where new classes of objects are continually collected by visual and tactile sensors. In this article, we define this scenario as V isuo- T actile C lass- I ncremental L earning (VT-CIL). In practical VT-CIL, the robot needs to adapt to a new environment with constrained storage and computing resources, and suffers from the severe forgetting of vision and touch knowledge about old environments. To alleviate this problem, we consider visuo-tactile correlations in VT-CIL and propose a novel framework. It efficiently incorporates the Visuo-Tactile Cross-Modal Pseudo-Label-Consistent (VT-CMPLC) constraint, Dual-Visuo-Tactile Exemplars (DVT-E), and the Dual-Visuo-Tactile-Compatible (DVT-C) constraint. The old visual–tactile classes are preserved by the VT-CMPLC constraint and DVT-E, while the visuo-tactile correlations and the VT-CMPLC and DVT-E capabilities are enhanced by the DVT-C constraint. We built two benchmarks, the Touch-and-Go Class-Incremental (TaG-CI) benchmark and the ObjectFolder-Real Class-Incremental (OFR-CI) benchmark. Experimental results on TaG-CI and OFR-CI benchmarks demonstrate the effectiveness of our method against previous state-of-the-art class-incremental learning methods in VT-CIL. Hao Fu 0023, Fengyu Yang 0003, Boyang Wang 0008, Wei Ji 0008, Hanbin Zhao, Chao Zhang 0029, Roger Zimmermann, Hui Qian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2025 | DriveDiTFit: Fine-tuning Diffusion Transformers for Autonomous Driving Data GenerationabstractIn autonomous driving, deep models have shown remarkable performance across various visual perception tasks with the demand of high-quality and huge-diversity training datasets. Such datasets are expected to cover various driving scenarios with adverse weather, lighting conditions, and diverse moving objects. However, manually collecting these data presents huge challenges and is expensive. With the rapid development of large generative models, we propose DriveDiTFit, a novel method for efficiently generating autonomous Driv ing data by Fi ne- t uning pre-trained Di ffusion T ransformers (DiTs). Specifically, DriveDiTFit utilizes a gap-driven modulation technique to carefully select and efficiently fine-tune a few parameters in DiTs according to the discrepancy between the pre-trained source data and the target driving data. Additionally, DriveDiTFit develops an effective weather and lighting condition embedding module to ensure diversity in the generated data, which is initialized by a nearest-semantic-similarity initialization approach. Through progressive tuning scheme to refine the process of detail generation in early diffusion process and enlarging the weights corresponding to small objects in training loss, DriveDiTFit ensures high-quality generation of small moving objects in the generated data. Extensive experiments conducted on driving datasets confirm that our method could efficiently produce diverse real driving data. Jiahang Tu, Wei Ji 0008, Hanbin Zhao, Chao Zhang 0029, Roger Zimmermann, Hui Qian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | GAD-PVI: A General Accelerated Dynamic-Weight Particle-Based Variational Inference FrameworkabstractParticle-based Variational Inference (ParVI) methods approximate the target distribution by iteratively evolving finite weighted particle systems. Recent advances of ParVI methods reveal the benefits of accelerated position update strategies and dynamic weight adjustment approaches. In this paper, we propose the first ParVI framework that possesses both accelerated position update and dynamical weight adjustment simultaneously, named the General Accelerated Dynamic-Weight Particle-based Variational Inference (GAD-PVI) framework. Generally, GAD-PVI simulates the semi-Hamiltonian gradient flow on a novel Information-Fisher-Rao space, which yields an additional decrease on the local functional dissipation. GAD-PVI is compatible with different dissimilarity functionals and associated smoothing approaches under three information metrics. Experiments on both synthetic and real-world data demonstrate the faster convergence and reduced approximation error of GAD-PVI methods over the state-of-the-art. Fangyikang Wang, Huminhao Zhu, Chao Zhang 0029, Hanbin Zhao, Hui Qian 0001 |
AAAI | 5 |
| 2024 | D-LLM: A Token Adaptive Computing Resource Allocation Strategy for Large Language ModelsabstractLarge language models have shown an impressive societal impact owing to their excellent understanding and logical reasoning skills. However, such strong ability relies on a huge amount of computing resources, which makes it difficult to deploy LLMs on computing resource-constrained platforms. Currently, LLMs process each token equivalently, but we argue that not every word is equally important. Some words should not be allocated excessive computing resources, particularly for dispensable terms in simple questions. In this paper, we propose a novel dynamic inference paradigm for LLMs, namely D-LLMs, which adaptively allocate computing resources in token processing. We design a dynamic decision module for each transformer layer that decides whether a network unit should be executed or skipped. Moreover, we tackle the issue of adapting D-LLMs to real-world applications, specifically concerning the missing KV-cache when layers are skipped. To overcome this, we propose a simple yet effective eviction policy to exclude the skipped layers from subsequent attention calculations. The eviction policy not only enables D-LLMs to be compatible with prevalent applications but also reduces considerable storage resources. Experimentally, D-LLMs show superior performance, in terms of computational cost and KV storage utilization. It can reduce up to 45\% computational cost and KV storage on Q\&A, summarization, and math solving tasks, 50\% on commonsense reasoning tasks. Yikun Jiang, Hanbin Zhao, Zhang Chao, Hui Qian 0001, John C. S. Lui |
NeurIPS | 6 |
| 2024 | BELM: Bidirectional Explicit Linear Multi-step Sampler for Exact Inversion in Diffusion ModelsabstractThe inversion of diffusion model sampling, which aims to find the corresponding initial noise of a sample, plays a critical role in various tasks.
Recently, several heuristic exact inversion samplers have been proposed to address the inexact inversion issue in a training-free manner.
However, the theoretical properties of these heuristic samplers remain unknown and they often exhibit mediocre sampling quality.
In this paper, we introduce a generic formulation, \emph{Bidirectional Explicit Linear Multi-step} (BELM) samplers, of the exact inversion samplers, which includes all previously proposed heuristic exact inversion samplers as special cases.
The BELM formulation is derived from the variable-stepsize-variable-formula linear multi-step method via integrating a bidirectional explicit constraint. We highlight this bidirectional explicit constraint is the key of mathematically exact inversion.
We systematically investigate the Local Truncation Error (LTE) within the BELM framework and show that the existing heuristic designs of exact inversion samplers yield sub-optimal LTE.
Consequently, we propose the Optimal BELM (O-BELM) sampler through the LTE minimization approach.
We conduct additional analysis to substantiate the theoretical stability and global convergence property of the proposed optimal sampler.
Comprehensive experiments demonstrate our O-BELM sampler establishes the exact inversion property while achieving high-quality sampling.
Additional experiments in image editing and image interpolation highlight the extensive potential of applying O-BELM in varying applications. Fangyikang Wang, Hubery Yin, Yuejiang Dong, Huminhao Zhu, Zhang Chao, Hanbin Zhao, Hui Qian 0001, Chen Li 0031 |
NeurIPS | 7 |
| 2024 | Solving Zero-Sum Markov Games with Continuous State via Spectral Dynamic EmbeddingabstractIn this paper, we propose a provably efficient natural policy gradient algorithm called Spectral Dynamic Embedding Policy Optimization (\SDEPO) for two-player zero-sum stochastic Markov games with continuous state space and finite action space.
In the policy evaluation procedure of our algorithm, a novel kernel embedding method is employed to construct a finite-dimensional linear approximations to the state-action value function.
We explicitly analyze the approximation error in policy evaluation, and show that \SDEPO\ achieves an $\tilde{O}(\frac{1}{(1-\gamma)^3\epsilon})$ last-iterate convergence to the $\epsilon-$optimal Nash equilibrium, which is independent of the cardinality of the state space.
The complexity result matches the best-known results for global convergence of policy gradient algorithms for single agent setting.
Moreover, we also propose a practical variant of \SDEPO\ to deal with continuous action space and empirical results demonstrate the practical superiority of the proposed method. Chenhao Zhou 0003, Zebang Shen, Zhang Chao, Hanbin Zhao, Hui Qian 0001 |
NeurIPS | 5 |
| 2023 | Towards Optimal Randomized Strategies in Adversarial Example GameabstractThe vulnerability of deep neural network models to adversarial example attacks is a practical challenge in many artificial intelligence applications. A recent line of work shows that the use of randomization in adversarial training is the key to find optimal strategies against adversarial example attacks. However, in a fully randomized setting where both the defender and the attacker can use randomized strategies, there are no efficient algorithm for finding such an optimal strategy. To fill the gap, we propose the first algorithm of its kind, called FRAT, which models the problem with a new infinite-dimensional continuous-time flow on probability distribution spaces. FRAT maintains a lightweight mixture of models for the defender, with flexibility to efficiently update mixing weights and model parameters at each iteration. Furthermore, FRAT utilizes lightweight sampling subroutines to construct a random strategy for the attacker. We prove that the continuous-time limit of FRAT converges to a mixed Nash equilibria in a zero-sum game formed by a defender and an attacker. Experimental results also demonstrate the efficiency of FRAT on CIFAR-10 and CIFAR-100 datasets. Jiahao Xie 0001, Chao Zhang 0029, Weijie Liu 0006, Wensong Bai, Hui Qian 0001 |
AAAI | 5 |
| 2023 | CDMA: A Practical Cross-Device Federated Learning Algorithm for General Minimax ProblemsabstractMinimax problems arise in a wide range of important applications including robust adversarial learning and Generative Adversarial Network (GAN) training. Recently, algorithms for minimax problems in the Federated Learning (FL) paradigm have received considerable interest. Existing federated algorithms for general minimax problems require the full aggregation (i.e., aggregation of local model information from all clients) in each training round. Thus, they are inapplicable to an important setting of FL known as the cross-device setting, which involves numerous unreliable mobile/IoT devices. In this paper, we develop the first practical algorithm named CDMA for general minimax problems in the cross-device FL setting. CDMA is based on a Start-Immediately-With-Enough-Responses mechanism, in which the server first signals a subset of clients to perform local computation and then starts to aggregate the local results reported by clients once it receives responses from enough clients in each round. With this mechanism, CDMA is resilient to the low client availability. In addition, CDMA is incorporated with a lightweight global correction in the local update steps of clients, which mitigates the impact of slow network connections. We establish theoretical guarantees of CDMA under different choices of hyperparameters and conduct experiments on AUC maximization, robust adversarial network training, and GAN training tasks. Theoretical and experimental results demonstrate the efficiency of CDMA. Jiahao Xie 0001, Chao Zhang 0029, Zebang Shen, Weijie Liu 0006, Hui Qian 0001 |
AAAI | 5 |
| 2023 | Robust Graph Dictionary Learning
Weijie Liu 0006, Jiahao Xie 0001, Chao Zhang 0029, Makoto Yamada, Nenggan Zheng, Hui Qian 0001 |
ICLR | 6 |
| 2023 | Vision Transformers for Single Image DehazingabstractImage dehazing is a representative low-level vision task that estimates latent haze-free images from hazy images. In recent years, convolutional neural network-based methods have dominated image dehazing. However, vision Transformers, which has recently made a breakthrough in high-level vision tasks, has not brought new dimensions to image dehazing. We start with the popular Swin Transformer and find that several of its key designs are unsuitable for image dehazing. To this end, we propose DehazeFormer, which consists of various improvements, such as the modified normalization layer, activation function, and spatial information aggregation scheme. We train multiple variants of DehazeFormer on various datasets to demonstrate its effectiveness. Specifically, on the most frequently used SOTS indoor set, our small model outperforms FFA-Net with only 25% #Param and 5% computational cost. To the best of our knowledge, our large model is the first method with the PSNR over 40 dB on the SOTS indoor set, dramatically outperforming the previous state-of-the-art methods. We also collect a large-scale realistic remote sensing dehazing dataset for evaluating the method's capability to remove highly non-homogeneous haze. We share our code and dataset at https://github.com/IDKiro/DehazeFormer. Yuda Song 0002, Zhuqing He, Hui Qian 0001, Xin Du 0005 |
IEEE Trans. Image Process. | 3 |
| 2022 | From One to All: Learning to Match Heterogeneous and Partially Overlapped GraphsabstractRecent years have witnessed a flurry of research activity in graph matching, which aims at finding the correspondence of nodes across two graphs and lies at the heart of many artificial intelligence applications. However, matching heterogeneous graphs with partial overlap remains a challenging problem in real-world applications. This paper proposes the first practical learning-to-match method to meet this challenge. The proposed unsupervised method adopts a novel partial optimal transport paradigm to learn a transport plan and node embeddings simultaneously. In a from-one-to-all manner, the entire learning procedure is decomposed into a series of easy-to-solve sub-procedures, each of which only handles the alignment of a single type of nodes. A mechanism for searching the transport mass is also proposed. Experimental results demonstrate that the proposed method outperforms state-of-the-art graph matching methods. Weijie Liu 0006, Hui Qian 0001, Chao Zhang 0029, Jiahao Xie 0001, Zebang Shen, Nenggan Zheng |
AAAI | 2 |
| 2022 | Multi-Curve Translator for High-Resolution Photorealistic Image Translation
Yuda Song 0002, Hui Qian 0001, Xin Du 0005 |
ECCV (15) | 2 |
| 2022 | DPVI: A Dynamic-Weight Particle-Based Variational Inference FrameworkabstractThe recently developed Particle-based Variational Inference (ParVI) methods drive the empirical distribution of a set of fixed-weight particles towards a given target distribution by iteratively updating particles' positions. However, the fixed weight restriction greatly confines the empirical distribution's approximation ability, especially when the particle number is limited. In this paper, we propose to dynamically adjust particles' weights according to a Fisher-Rao reaction flow. We develop a general Dynamic-weight Particle-based Variational Inference (DPVI) framework according to a novel continuous composite flow, which evolves the positions and weights of particles simultaneously. We show that the mean-field limit of our composite flow is actually a Wasserstein-Fisher-Rao gradient flow of the associated dissimilarity functional. By using different finite-particle approximations in our general framework, we derive several efficient DPVI algorithms. The empirical results demonstrate the superiority of our derived DPVI algorithms over their fixed-weight counterparts. Chao Zhang 0029, Xin Du 0005, Hui Qian 0001 |
IJCAI | 4 |
| 2021 | A Hybrid Stochastic Gradient Hamiltonian Monte Carlo Method
Chao Zhang 0029, Zebang Shen, Jiahao Xie 0001, Hui Qian 0001 |
AAAI | 5 |
| 2021 | StarEnhancer: Learning Real-Time and Style-Aware Image EnhancementabstractImage enhancement is a subjective process whose targets vary with user preferences. In this paper, we propose a deep learning-based image enhancement method covering multiple tonal styles using only a single model dubbed StarEnhancer. It can transform an image from one tonal style to another, even if that style is unseen. With a simple one-time setting, users can customize the model to make the enhanced images more in line with their aesthetics. To make the method more practical, we propose a well-designed enhancer that can process a 4K-resolution image over 200 FPS but surpasses the contemporaneous single style image enhancement methods in terms of PSNR, SSIM, and LPIPS. Finally, our proposed enhancement method has good inter-actability, which allows the user to fine-tune the enhanced image using intuitive options. Yuda Song 0002, Hui Qian 0001, Xin Du 0005 |
ICCV | 2 |
| 2021 | SHPOS: A Theoretical Guaranteed Accelerated Particle Optimization Sampling MethodabstractRecently, the Stochastic Particle Optimization Sampling (SPOS) method is proposed to solve the particle-collapsing pitfall of deterministic Particle Variational Inference methods by ultilizing the stochastic Overdamped Langevin dynamics to enhance exploration. In this paper, we propose an accelerated particle optimization sampling method called Stochastic Hamiltonian Particle Optimization Sampling (SHPOS). Compared to the first-order dynamics used in SPOS, SHPOS adopts an augmented second-order dynamics, which involves an extra momentum term to achieve acceleration. We establish a non-asymptotic convergence analysis for SHPOS, and show that it enjoys a faster convergence rate than SPOS. Besides, we also propose a variance-reduced stochastic gradient variant of SHPOS for tasks with large-scale datasets and complex models. Experiments on both synthetic and real data validate our theory and demonstrate the superiority of SHPOS over the state-of-the-art. Chao Zhang 0029, Hui Qian 0001, Xin Du 0005, Lingwei Peng |
IJCAI | 3 |
| 2020 | Efficient Projection-Free Online Methods with Stochastic Recursive GradientabstractThis paper focuses on projection-free methods for solving smooth Online Convex Optimization (OCO) problems. Existing projection-free methods either achieve suboptimal regret bounds or have high per-round computational costs. To fill this gap, two efficient projection-free online methods called ORGFW and MORGFW are proposed for solving stochastic and adversarial OCO problems, respectively. By employing a recursive gradient estimator, our methods achieve optimal regret bounds (up to a logarithmic factor) while possessing low per-round computational costs. Experimental results demonstrate the efficiency of the proposed methods compared to state-of-the-arts. Jiahao Xie 0001, Zebang Shen, Chao Zhang 0029, Hui Qian 0001 |
AAAI | 5 |
| 2020 | Aggregated Gradient Langevin Dynamics
Chao Zhang 0029, Jiahao Xie 0001, Zebang Shen, Peilin Zhao, Hui Qian 0001 |
AAAI | 6 |
| 2020 | Accelerating Stratified Sampling SGD by Reconstructing StrataabstractIn this paper, a novel stratified sampling strategy is designed to accelerate the mini-batch SGD. We derive a new iteration-dependent surrogate which bound the stochastic variance from above. To keep the strata minimizing this surrogate with high probability, a stochastic stratifying algorithm is adopted in an adaptive manner, that is, in each iteration, strata are reconstructed only if an easily verifiable condition is met. Based on this novel sampling strategy, we propose an accelerated mini-batch SGD algorithm named SGD-RS. Our theoretical analysis shows that the convergence rate of SGD-RS is superior to the state-of-the-art. Numerical experiments corroborate our theory and demonstrate that SGD-RS achieves at least 3.48-times speed-ups compared to vanilla minibatch SGD. Weijie Liu 0006, Hui Qian 0001, Chao Zhang 0029, Zebang Shen, Jiahao Xie 0001, Nenggan Zheng |
IJCAI | 2 |
| 2019 | Complexities in Projection-Free Stochastic Non-convex MinimizationabstractFor constrained nonconvex minimization problems, we propose a meta stochastic projection-free optimization algorithm, named Normalized Frank Wolfe Updating, that can take any Gradient Estimator (GE) as input. For this algorithm, we prove its convergence rate, regardless of the choice of GE. Using a sophisticated GE, this algorithm can significantly improve the Stochastic First order Oracle (SFO) complexity. Further, a new second order GE strategy is proposed to incorporate curvature information, which enjoys theoretical advantage over the first order ones. Besides, this paper also provides a lower bound of Linear-optimization Oracle (LO) queried to achieve an approximate stationary point. Simulation studies validate our analysis under various parameter settings. Zebang Shen, Cong Fang 0001, Peilin Zhao, Junzhou Huang, Hui Qian 0001 |
AISTATS | 5 |
| 2019 | Decentralized Gradient Tracking for Continuous DR-Submodular MaximizationabstractIn this paper, we focus on the continuous DR-submodular maximization over a network. By using the gradient tracking technique, two decentralized algorithms are proposed for deterministic and stochastic settings, respectively. The proposed methods attain the $\epsilon$-accuracy tight approximation ratio for monotone continuous DR-submodular functions in only $O(1/\epsilon)$ and $\tilde{O}(1/\epsilon)$ rounds of communication, respectively, which are superior to the state-of-the-art. Our numerical results show that the proposed methods outperform existing decentralized methods in terms of both computation and communication complexity. Jiahao Xie 0001, Chao Zhang 0029, Zebang Shen, Chao Mi, Hui Qian 0001 |
AISTATS | 5 |
| 2019 | Hessian Aided Policy GradientabstractReducing the variance of estimators for policy gradient has long been the focus of reinforcement learning research. While classic algorithms like REINFORCE find an $\epsilon$-approximate first-order stationary point in $\OM({1}/{\epsilon^4})$ random trajectory simulations, no provable improvement on the complexity has been made so far. This paper presents a Hessian aided policy gradient method with the first improved sample complexity of $\OM({1}/{\epsilon^3})$. While our method exploits information from the policy Hessian, it can be implemented in linear time with respect to the parameter dimension and is hence applicable to sophisticated DNN parameterization. Simulations on standard tasks validate the efficiency of our method. Zebang Shen, Alejandro Ribeiro, Seyed Hamed Hassani, Hui Qian 0001, Chao Mi |
ICML | 4 |
| 2018 | Towards Memory-Friendly Deterministic Incremental Gradient MethodabstractIncremental Gradient (IG) methods are classical strategies in solving finite sum minimization problems. Deterministic IG methods are particularly favorable in handling massive scale problem due to its memory-friendly data access pattern. In this paper, we propose a new deterministic variant of the IG method SVRG that blends a periodically updated full gradient with a component function gradient selected in a cyclic order. Our method uses only $O(1)$ extra gradient storage without compromising the linear convergence. Empirical results demonstrate that the proposed method is advantageous over existing incremental gradient algorithms, especially on problems that does not fit into physical memory. Jiahao Xie 0001, Hui Qian 0001, Zebang Shen, Chao Zhang 0029 |
AISTATS | 2 |
| 2018 | Towards More Efficient Stochastic Decentralized Learning: Faster Convergence and Sparse CommunicationabstractRecently, the decentralized optimization problem is attracting growing attention. Most existing methods are deterministic with high per-iteration cost and have a convergence rate quadratically depending on the problem condition number. Besides, the dense communication is necessary to ensure the convergence even if the dataset is sparse. In this paper, we generalize the decentralized optimization problem to a monotone operator root finding problem, and propose a stochastic algorithm named DSBA that (1) converges geometrically with a rate linearly depending on the problem condition number, and (2) can be implemented using sparse communication only. Additionally, DSBA handles important learning problems like AUC-maximization which can not be tackled efficiently in the previous problem setting. Experiments on convex minimization and AUC-maximization validate the efficiency of our method. Zebang Shen, Aryan Mokhtari, Peilin Zhao, Hui Qian 0001 |
ICML | 5 |
| 2018 | JUMP: a Jointly Predictor for User Click and Dwell TimeabstractWith the recent proliferation of recommendation system, there have been a lot of interests in session-based prediction methods, particularly those based on Recurrent Neural Network (RNN) and their variants. However, existing methods either ignore the dwell time prediction that plays an important role in measuring user's engagement on the content, or fail to process very short or noisy sessions. In this paper, we propose a joint predictor, JUMP, for both user click and dwell time in session-based settings. To map its input into a feature vector, JUMP adopts a novel three-layered RNN structure which includes a fast-slow layer for very short sessions and an attention layer for noisy sessions. Experiments demonstrate that JUMP outperforms state-of-the-art methods in both user click and dwell time prediction. Hui Qian 0001, Zebang Shen, Chao Zhang 0029, Chengwei Wang, Shichen Liu, Wenwu Ou |
IJCAI | 2 |
| 2017 | Loosecut: Interactive image segmentation with loosely bounded boxesabstractOne popular approach to interactively segment an object of interest from an image is to annotate a bounding box that covers the object, followed by a binary labeling. However, the existing algorithms for such interactive image segmentation prefer a bounding box that tightly encloses the object. This increases the annotation burden, and prevents these algorithms from utilizing automatically detected bounding boxes. In this paper, we develop a new LooseCut algorithm that can handle cases where the bounding box only loosely covers the object. We propose a new Markov Random Fields (MRF) model for segmentation with loosely bounded boxes, including an additional energy term to encourage consistent labeling of similar-appearance pixels and a global similarity constraint to better distinguish the foreground and background. This MRF model is then solved by an iterated max-flow algorithm. We evaluate LooseCut in three public image datasets, and show its better performance against several state-of-the-art methods when increasing the bounding-box size. Hongkai Yu, Youjie Zhou, Hui Qian 0001, Min Xian, Song Wang 0002 |
ICIP | 3 |
| 2017 | Accelerated Doubly Stochastic Gradient Algorithm for Large-scale Empirical Risk MinimizationabstractNowadays, algorithms with fast convergence, small memory footprints, and low per-iteration complexity are particularly favorable for artificial intelligence applications. In this paper, we propose a doubly stochastic algorithm with a novel accelerating multi-momentum technique to solve large scale empirical risk minimization problem for learning tasks. While enjoying a provably superior convergence rate, in each iteration, such algorithm only accesses a mini batch of samples and meanwhile updates a small block of variable coordinates, which substantially reduces the amount of memory reference when both the massive sample size and ultra-high dimensionality are involved. Specifically, to obtain an ε-accurate solution, our algorithm requires only O(log(1/ε)/sqrt(ε)) overall computation for the general convex case and O((n+sqrt{nκ})log(1/ε)) for the strongly convex case. Empirical studies on huge scale datasets are conducted to illustrate the efficiency of our method in practice. Zebang Shen, Hui Qian 0001, Tongzhou Mu, Chao Zhang 0029 |
IJCAI | 2 |
| 2017 | Tensor Completion with Side Information: A Riemannian Manifold ApproachabstractBy restricting the iterate on a nonlinear manifold, the recently proposed Riemannian optimization methods prove to be both efficient and effective in low rank tensor completion problems. However, existing methods fail to exploit the easily accessible side information, due to their format mismatch. Consequently, there is still room for improvement. To fill the gap, in this paper, a novel Riemannian model is proposed to tightly integrate the original model and the side information by overcoming their inconsistency. For this model, an efficient Riemannian conjugate gradient descent solver is devised based on a new metric that captures the curvature of the objective. Numerical experiments suggest that our method is more accurate than the state-of-the-art without compromising the efficiency. Hui Qian 0001, Zebang Shen, Chao Zhang 0029, Congfu Xu |
IJCAI | 2 |
| 2016 | An Alternating Proximal Splitting Method with Global Convergence for Nonconvex Structured Sparsity OptimizationabstractIn many learning tasks with structural properties, structured sparse modeling usually leads to better interpretability and higher generalization performance. While great efforts have focused on the convex regularization, recent studies show that nonconvex regularizers can outperform their convex counterparts in many situations. However, the resulting nonconvex optimization problems are still challenging, especially for the structured sparsity-inducing regularizers. In this paper, we propose a splitting method for solving nonconvex structured sparsity optimization problems. The proposed method alternates between a gradient step and an easily solvable proximal step, and thus enjoys low per-iteration computational complexity. We prove that the whole sequence generated by the proposed method converges to a critical point with at least sublinear convergence rate, relying on the Kurdyka-Łojasiewicz inequality. Experiments on both simulated and real-world data sets demonstrate the efficiency and efficacy of the proposed method. Shubao Zhang, Hui Qian 0001, Xiaojin Gong |
AAAI | 2 |
| 2016 | Fast Hybrid Algorithm for Big Matrix Recovery
Hui Qian 0001, Zebang Shen, Congfu Xu |
AAAI | 2 |
| 2016 | Adaptive Variance Reducing for Stochastic Gradient Descent
Zebang Shen, Hui Qian 0001, Tongzhou Mu |
IJCAI | 2 |
| 2016 | Constrained Preference Embedding for Item Recommendation
Xin Wang 0060, Congfu Xu, Yunhui Guo, Hui Qian 0001 |
IJCAI | 4 |
| 2015 | Co-Interest Person Detection from Multiple Wearable Camera VideosabstractWearable cameras, such as Google Glass and Go Pro, enable video data collection over larger areas and from different views. In this paper, we tackle a new problem of locating the co-interest person (CIP), i.e., the one who draws attention from most camera wearers, from temporally synchronized videos taken by multiple wearable cameras. Our basic idea is to exploit the motion patterns of people and use them to correlate the persons across different videos, instead of performing appearance-based matching as in traditional video co-segmentation/localization. This way, we can identify CIP even if a group of people with similar appearance are present in the view. More specifically, we detect a set of persons on each frame as the candidates of the CIP and then build a Conditional Random Field (CRF) model to select the one with consistent motion patterns in different videos and high spacial-temporal consistency in each video. We collect three sets of wearable-camera videos for testing the proposed algorithm. All the involved people have similar appearances in the collected videos and the experiments demonstrate the effectiveness of the proposed algorithm. Yuewei Lin, Kareem Abdelfatah, Youjie Zhou, Xiaochuan Fan, Hongkai Yu, Hui Qian 0001, Song Wang 0002 |
ICCV | 6 |
| 2015 | Simple Atom Selection Strategy for Greedy Matrix Completion
Zebang Shen, Hui Qian 0001, Song Wang 0002 |
IJCAI | 2 |
| 2014 | Using The Matrix Ridge Approximation to Speedup Determinantal Point Processes Sampling AlgorithmsabstractDeterminantal point process (DPP) is an important probabilistic model that has extensive applications in artificial intelligence. The exact sampling algorithm of DPP requires the full eigenvalue decomposition of the kernel matrix which has high time and space complexities. This prohibits the applications of DPP from large-scale datasets. Previous work has applied the Nystrom method to speedup the sampling algorithm of DPP, and error bounds have been established for the approximation. In this paper we employ the matrix ridge approximation (MRA) to speedup the sampling algorithm of DPP, showing that our approach MRA-DPP has stronger error bound than the Nystrom-DPP. In certain circumstances our MRA-DPP is provably exact, whereas the Nystrom-DPP is far from the ground truth. Finally, experiments on several real-world datasets show that our MRA-DPP is more accurate than the other approximation approaches. Shusen Wang, Chao Zhang 0029, Hui Qian 0001, Zhihua Zhang 0008 |
AAAI | 3 |
| 2014 | Making Fisher Discriminant Analysis ScalableabstractThe Fisher linear discriminant analysis (LDA) is a classical method for classification and dimension reduction jointly. A major limitation of the conventional LDA is a so-called singularity issue. Many LDA variants, especially two-stage methods such as PCA+LDA and LDA/QR, were proposed to solve this issue. In the two-stage methods, an intermediate stage for dimension reduction is developed before the actual LDA method works. These two-stage methods are scalable because they are an approximate alternative of the LDA method. However, there is no theoretical analysis on how well they approximate the conventional LDA problem. In this paper we present theoretical analysis on the approximation error of a two-stage algorithm. Accordingly, we develop a new two-stage algorithm. Furthermore, we resort to a random projection approach, making our algorithm scalable. We also provide an implemention on distributed system to handle large scale problems. Our algorithm takes LDA/QR as its special case, and outperforms PCA+LDA while having a similar scalability. We also generalize our algorithm to kernel discriminant analysis, a nonlinear version of the classical LDA. Extensive experiments show that our algorithms outperform PCA+LDA and have a similar scalability with it. Bojun Tu, Zhihua Zhang 0008, Shusen Wang, Hui Qian 0001 |
ICML | 4 |
| 2014 | Improving the modified nyström method using spectral shiftingabstractThe Nystrom method is an efficient approach to enabling large-scale kernel methods. The Nystrom method generates a fast approximation to any large-scale symmetric positive semidefinete (SPSD) matrix using only a few columns of the SPSD matrix. However, since the Nystrom approximation is low-rank, when the spectrum of the SPSD matrix decays slowly, the Nystrom approximation is of low accuracy. In this paper, we propose a variant of the Nystrom method called the modified Nystrom by spectral shifting (SS-Nystrom). The SS-Nystrom method works well no matter whether the spectrum of SPSD matrix decays fast or slow. We prove that our SS-Nystrom has a much stronger error bound than the standard and modified Nystrom methods, and that SS-Nystrom can be even more accurate than the truncated SVD of the same scale in some cases. We also devise an algorithm such that the SS-Nystrom approximation can be computed nearly as efficient as the modified Nystrom approximation. Finally, our SS-Nystrom method demonstrates significant improvements over the standard and modified Nystrom methods on several real-world datasets. Shusen Wang, Chao Zhang 0029, Hui Qian 0001, Zhihua Zhang 0008 |
KDD | 3 |
| 2013 | Large-Scale Hierarchical Classification via Stochastic Perceptron
Dehua Liu, Bojun Tu, Hui Qian 0001, Zhihua Zhang 0008 |
AAAI | 3 |
| 2013 | A Concave Conjugate Approach for Nonconvex Penalized Regression with the MCP PenaltyabstractThe minimax concave plus penalty (MCP) has been demonstrated to be effective in nonconvex penalization for feature selection. In this paper we propose a novel construction approach for MCP. In particular, we show that MCP can be derived from a concave conjugate of the Euclidean distance function. This construction approach in turn leads us to an augmented Lagrange multiplier method for solving the penalized regression problem with MCP. In our method each tuning parameter corresponds to a feature, and these tuning parameters can be automatically updated. We also develop a d.c. (difference of convex functions) programming approach for the penalized regression problem. We find that the augmented Lagrange multiplier method degenerates into the d.c. programming method under specific conditions. Experimental analysis is conducted on a set of simulated data. The result is encouraging. Shubao Zhang, Hui Qian 0001, Zhihua Zhang 0008 |
AAAI | 2 |
| 2013 | A Nearly Unbiased Matrix Completion Approach
Dehua Liu, Hui Qian 0001, Congfu Xu, Zhihua Zhang 0008 |
ECML/PKDD (2) | 3 |
| 2013 | An iterative SVM approach to feature selection and classification in high-dimensional datasets
Dehua Liu, Hui Qian 0001, Guang Dai, Zhihua Zhang 0008 |
Pattern Recognit. | 2 |
| 2010 | Modified reward function on abstract features in inverse reinforcement learningabstractWe improve inverse reinforcement learning (IRL) by applying dimension reduction methods to automatically extract abstract features from human-demonstrated policies, to deal with the cases where features are either unknown or numerous. The importance rating of each abstract feature is incorporated into the reward function. Simulation is performed on a task of driving in a five-lane highway, where the controlled car has the largest fixed speed among all the cars. Performance is almost 10.6% better on average with than without importance ratings. Shenyi Chen, Hui Qian 0001, Jia Fan, Zhuo-Jun Jin, Miaoliang Zhu |
J. Zhejiang Univ. Sci. C | 2 |
| 2008 | 3D object recognition by fast spherical correlation between combined view EGIs and PFTabstractThis paper proposes a method to recognize 3D object from probed range image under arbitrary pose by fast spherical correlation. First, all view EGIs under different viewpoints are extracted and combined onto a Gaussian sphere to form a feature description for each object. Second, the probed range image at arbitrary pose is represented as a PFT feature by phase-encoded Fourier transform and the PFT feature is mapped onto the Gaussian hemisphere by coordinates transforms and intensity scaling. Third, the spherical correlation algorithm based on spherical harmonic functions is used to do matching and similarity measurement between mapped PFT and combined view EGIs. The spherical correlation peak can output both of the recognition result and pose estimation. The experimental results proved that the proposed method can not only recognize totally different objects but also has enough discriminating capability for scalable dataset of similar objects. Hui Qian 0001 |
ICPR | 2 |