EDBT 2026 Demo / reviewers in the wild / expert
Dingjiang Huang
dblp:120/4493
· DBLP profile ↗
25ranked-venue papers
4as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating OOD overoptimism via in-sample value function in offline reinforcement learning
Kangyang Luo, Zhijian Wu, Shanfeng Hao, Dingjiang Huang |
Neural Networks | 5 |
| 2025 | Federated Prototype Guided Adaption for Vision-Language ModelsabstractFederated Learning (FL) is a new pivotal paradigm for decentralized training on heterogeneous data. Recently fine-tuning of Vision-Language Models (VLMs) has been extended to the federated setting to improve overall performance. Unfortunately, in this case, FL still faces two critical challenges that hinder its actual performance: data distribution heterogeneity and high resource costs brought by large VLMs. In this paper, we introduce FedPGA, a prototype-guided method, for achieving performance improvements in the federated setting for VLMs. Concretely, we design a prototype-based adapter for the vision-language model, CLIP. The lightweight adapter updates the prior knowledge encoded in CLIP to enhance its adaption capability further and avoid the effects of data distribution heterogeneity in the federated setting. Simultaneously, small-scale operations can mitigate the computational and communication burden caused by large VLMs. Our comprehensive empirical evaluations of nine diverse image classification datasets show that our method is superior to existing FL methods under VLMs. Youchao Liu, Dingjiang Huang |
ICASSP | 2 |
| 2025 | GCAT: Gated Convolutional Attention Transformer for Efficient Image Super-ResolutionabstractRecently, Transformer-based methods have achieved impressive performance in many computer vision tasks (e.g., image super-resolution (SR)) due to the advantages of long-range modeling. However, the computational cost requirement renders these methods unsuitable on resource-constrain devices, especially for image SR tasks involving high-resolution images. In this paper, we propose a concise and effective Gated Convolutional Attention Unit (GCAU) that uses cheap convolutional operations. Specifically, GCAU consists of Convolutional Transposed Attention (CTA) and Locally-enhanced Gating (LeG) in parallel. The former allows for efficient modeling of the global relational interactions by calculating cross-covariance across channels dimension, while the latter controls the information flow from the former directing the network to focus on more refined image attributes. Without bells and whistles, we present a simple SR Transformer GCAT by cascading the GCAUs. Extensive experimental results demonstrate that our GCAT achieves state-of-the-art performance among the existing efficient SR methods with significantly less complexity. Especially, GCAT is on average 5× faster than SwinIR-light with comparable performance. Zhijian Wu, Kaiyi Feng, Dingjiang Huang |
ICASSP | 3 |
| 2025 | EvaSR: Rethinking Efficient Visual Attention Design for Image Super-ResolutionabstractDue to the advantages of long-range modeling via the self-attention mechanism, Transformer has taken various vision tasks by storm, including image super-resolution (SR). In this study, we reveal that the convolutional neural network (CNN) with proper visual attention is a more simple and effective paradigm than Transformer in image SR tasks. We reexamine the successful SR models and discover several key characteristics that contribute to accurate image reconstruction. Built on this recipe, we propose a pure CNN-based SR network using efficient visual attention, dubbed EvaSR. Benefiting from the carefully designed visual attention, our EvaSR can favorably capture both local structure and long-range dependencies, and achieve adaptivity in spatial and channel dimensions while retaining the simplicity and efficiency of CNNs. The experimental results demonstrate that our EvaSR achieves state-of-the-art performance among the existing efficient SR methods. Especially, the tiny version of EvaSR needs 21.4% and 15.2% parameters of IMDN and SMSR with better performance. Zhijian Wu, Chenhan Zhang, Dingjiang Huang |
ICASSP | 3 |
| 2025 | Overcoming Feature Contamination by Unidirectional Information Modeling for Vision-Language TrackingabstractBenefiting from the advantages of multi-modal learning, Vision-Language Tracking shows greater potential than Visual Tracking. Existing work utilizes one-stream structures to fuse vision and language features, resulting in noise propagation from the search region into the language features. This contamination weakens the guidance of language information, consequently limiting the robustness of the tracking model. To solve this problem, we propose a Unidirectional Information modeling (UITracker) to explicitly fuse the language and visual features for Vision-Language Tracking. Specifically, we introduce a plug-and-play lightweight modal adapter to unidirectionally inject language guidance into the visual template and search region across all layers. This allows the tracker to make full use of rich semantic information while overcoming language feature contamination in the feature interaction process. Extensive ablation studies demonstrate the superiority and effectiveness of our UITracker. Code and raw results are available at https://github.com/jcwang0602/UITrack. Zhijian Wu, Dingjiang Huang |
ICME | 6 |
| 2025 | A Simple and Better Baseline for Visual GroundingabstractVisual grounding aims to predict the locations of target objects specified by textual descriptions. For this task with linguistic and visual modalities, there is a latest research line that focuses on only selecting the linguistic-relevant visual regions for object localization to reduce the computational overhead. Albeit achieving impressive performance, it is iteratively performed on different image scales, and at every iteration, linguistic features and visual features need to be stored in a cache, incurring extra overhead. To facilitate the implementation, in this paper, we propose a feature selection-based simple yet effective baseline for visual grounding, called FSVG. Specifically, we directly encapsulate the linguistic and visual modalities into an overall network architecture without complicated iterative procedures, and utilize the language in parallel as guidance to facilitate the interaction between linguistic modal and visual modal for extracting effective visual features. Furthermore, to reduce the computational cost, during the visual feature learning, we introduce a similarity-based feature selection mechanism to only exploit language-related visual features for faster prediction. Extensive experiments conducted on several benchmark datasets comprehensively substantiate that the proposed FSVG achieves a better balance between accuracy and efficiency beyond the current state-of-the-art methods. Code is available at https://github.com/jcwang0602/FSVG. Dingjiang Huang, Hong Wang 0021, Yefeng Zheng 0001 |
ICME | 3 |
| 2025 | Imagination-Limited Q-Learning for Offline Reinforcement LearningabstractOffline reinforcement learning seeks to derive improved policies entirely from historical data but often struggles with over-optimistic value estimates for out-of-distribution (OOD) actions. This issue is typically mitigated via policy constraint or conservative value regularization methods. However, these approaches may impose overly constraints or biased value estimates, potentially limiting performance improvements. To balance exploitation and restriction, we propose an Imagination-Limited Q-learning (ILQ) method, which aims to maintain the optimism that OOD actions deserve within appropriate limits. Specifically, we utilize the dynamics model to imagine OOD action-values, and then clip the imagined values with the maximum behavior values. Such design maintains reasonable evaluation of OOD actions to the furthest extent, while avoiding its over-optimism. Theoretically, we prove the convergence of the proposed ILQ under tabular Markov decision processes. Particularly, we demonstrate that the error bound between estimated values and optimality values of OOD state-actions possesses the same magnitude as that of in-distribution ones, thereby indicating that the bias in value estimates is effectively mitigated. Empirically, our method achieves state-of-the-art performance on a wide range of tasks in the D4RL benchmark. Zhijian Wu, Dingjiang Huang, Shuigeng Zhou |
IJCAI | 4 |
| 2025 | A Prior-Driven Lightweight Network for Endoscopic Exposure Correction
Zhijian Wu, Hong Wang 0021, Dingjiang Huang, Yefeng Zheng 0001 |
MICCAI (11) | 4 |
| 2025 | Learning Spectral Diffusion Prior for Hyperspectral Image ReconstructionabstractHyperspectral image (HSI) reconstruction aims to recover 3D HSI from its degraded 2D measurements. Recently great progress has been made in deep learning-based methods, however, these methods often struggle to accurately capture high-frequency details of the HSI. To address this issue, this paper proposes a Spectral Diffusion Prior (SDP) that is implicitly learned from hyperspectral images using a diffusion model. Leveraging the powerful ability of the diffusion model to reconstruct details, this learned prior can significantly improve the performance when injected into the HSI model. To further improve the effectiveness of the learned prior, we also propose the Spectral Prior Injector Module (SPIM) to dynamically guide the model to recover the HSI details. We evaluate our method on two representative HSI methods: MST and BISRNet. Experimental results show that our method outperforms existing networks by about 0.5 dB, effectively improving the performance of HSI reconstruction. Zhijian Wu, Dingjiang Huang |
SMC | 3 |
| 2025 | Sparse personalized federated class-incremental learning
Youchao Liu, Dingjiang Huang |
Inf. Sci. | 2 |
| 2025 | FrePrompter: Frequency self-prompt for all-in-one image restoration
Zhijian Wu, Jun Li 0033, Dingjiang Huang |
Pattern Recognit. | 5 |
| 2025 | Adversarial Conservative Alternating Q-Learning for Credit Card Debt CollectionabstractDebt collection is utilized for risk control after credit card delinquency. The existing rule-based method tends to be myopic and non-adaptive due to the delayed feedback. Reinforcement learning (RL) has an inherent advantage in dealing with such task and can learn policies end-to-end. However, employing RL here remains difficult because of different interaction processes from standard RL and the notorious problem of optimistic estimations in the offline setting. To tackle these challenges, we first propose an Alternating Q-Learning (AQL) framework to adapt debt collection processes to comparable procedures in RL. Based on AQL, we further develop an Adversarial Conservative Alternating Q-Learning (ACAQL) to address the issue of overoptimistic estimations. Specifically, adversarial conservative value regularization is proposed to balance optimism and conservatism on Q-values of out-of-distribution actions. Furthermore, ACAQL utilizes the counterfactual action stitching to mitigate the overestimation by enhancing behavior data. Finally, we evaluate ACAQL on a real-world dataset created from Bank of Shanghai. Offline experimental results show that our approach outperforms state-of-the-art methods and effectively alleviates the optimistic estimation issue. Moreover, we conduct online A/B tests on the bank, and ACAQL achieves at least a$\emph {6\%}$improvement of the debt recovery rate, which yields tangible economic benefits. Jiapeng Zhu 0002, Lyu Ni, Jingyu Bi, Zhijian Wu, Jiajie Long, Mengyao Gao, Dingjiang Huang, Shuigeng Zhou |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2024 | Ultralight-weight Binary Neural Network with 1K Parameters for Image Super-ResolutionabstractImage super-resolution (SR) is a long-standing research in the computer vision community, which aims to reconstruct high-resolution (HR) images from their low-resolution (LR) counterparts. The past decade has witnessed impressive advances propelled by deep learning methods. However, the prohibitive model complexity hinders the deployment of deep networks in resource-constrained edge devices. This paper addresses this pain point by pushing the neural architecture to an extremely small size. We comprehensively renovate the modern SR network design including shallow feature extraction, deep feature extraction, and upscale reconstruction, and propose an ultralight-weight binary neural network (UBSR) with only 1K parameters for image SR. Especially, we rethink the design of binary convolution and design an efficient binary convolution block tailored for the SR task. Experimental results show that the proposed method achieves promising performance with desirable parameters and computational overhead. Notably, our UBSR requires only 48M OPs for processing an image and the model parameters are only 1K. Zhijian Wu, Dingjiang Huang |
ICME | 2 |
| 2024 | When Handcrafted Filter Meets CNN: A Lightweight Conv-Filter Mixer Network for Efficient Image Super-ResolutionabstractDue to their powerful representational ability, convolutional neural networks (CNN) have achieved great success in image super-resolution (SR). In the trained SR models such as EDSR, we observe that partial convolutions exhibit analogous characteristics compared to handcrafted filters which avoid parameters with much less computational cost. This inspires us to substitute the handcrafted filters for the learnable convolutions in the SR models, such that the network complexity and the computational overhead are significantly reduced. In this study, we propose a novel lightweight SR network dubbed as Conv-Filter Mixer (CFM). Specifically, our CFM encapsulates three kinds of computations: learnable convolution, integrated filter unit (IFU), and identity mapping. Among them, IFU consists of diverse handcrafted filters to efficiently extract primitive representations in a non-parametric manner, making the limited parameterized components of lightweight networks focus on learning abstract and intricate features. To further improve efficiency, we introduce channel splitting and shuffling structures to mix the features produced by heterogeneous components efficiently. Extensive experiments demonstrate that our CFM achieves state-of-the-art performance with fewer parameters and computational costs. Zhijian Wu, Dingjiang Huang |
ICMR | 3 |
| 2024 | Compacter: A Lightweight Transformer for Image RestorationabstractAlthough deep learning-based methods have made significant advances in the field of image restoration (IR), they often suffer from excessive model parameters. To tackle this problem, this work proposes a compact Transformer (Compacter) for lightweight image restoration by making several key designs. We employ the concepts of projection sharing, adaptive interaction, and heterogeneous aggregation to develop a novel Compact Adaptive Self-Attention (CASA). Specifically, CASA utilizes shared projection to generate Query, Key, and Value to simultaneously model spatial and channel-wise self-attention. The adaptive interaction process is then used to propagate and integrate global information from two different dimensions, thus enabling omnidirectional relational interaction. Finally, a depth-wise convolution is incorporated on Value to complement heterogeneous local information, enabling global-local coupling. Moreover, we propose a Dual Selective Gated Module (DSGM) to dynamically encapsulate the globality into each pixel for context-adaptive aggregation. Extensive experiments demonstrate that our Compacter achieves state-of-the-art performance for a variety of lightweight IR tasks with approximately 400K parameters. Zhijian Wu, Jun Li 0033, Dingjiang Huang |
ACM Multimedia | 4 |
| 2024 | RUN: Rethinking the UNet Architecture for Efficient Image RestorationabstractRecent advanced image restoration (IR) methods typically stack homogeneous operators hierarchically in the UNet architecture. To achieve higher accuracy, these models are now going deeper and more complex, making them resource-intensive. After comprehensively reviewing different operators within modern networks, we provide an in-depth analysis of their individual favorable properties and invent a novel efficient IR network by redesigning the UNet architecture (RUN) with heterogeneous operators. Specifically, we propose three heterogeneous operators for different relational interactions concerning the specificity of different hierarchical features of the UNet architecture. First, the spatial self-attention block (SSA Block) processes high-resolution top-level features by modeling pixel interactions from the spatial dimension. Second, the channel self-attention block (CSA Block) performs channel recalibration and information transmission for the bottom-level features with rich channels. Finally, a simple and efficient convolution block (Conv Block) is used to facilitate middle-order information propagation, which complements the self-attention mechanism to achieve local-global coupling. Based on these designs, our RUN enables more comprehensive information dissemination and interaction regardless of topological distance, thus achieving superior performance while maintaining desirable computational budgets. Extensive experiments show that our RUN achieves state-of-the-art results for a variety of IR tasks, including image deblurring, image denoising, image deraining, and low-light image enhancement. Zhijian Wu, Jun Li 0033, Chang Xu 0002, Dingjiang Huang, Steven C. H. Hoi |
IEEE Trans. Multim. | 4 |
| 2023 | Separable Modulation Network for Efficient Image Super-ResolutionabstractDeep learning-based models have demonstrated unprecedented success in image super-resolution (SR) tasks. However, more attention has been paid to lightweight SR models lately, due to the increasing demand for on-device inference. In this paper, we propose a novel Separable Modulation Network (SMN) for efficient image SR. The key parts of the SMN are the Separable Modulation Unit (SMU) and the Locality Self-enhanced Network (LSN). SMU enables global relational interactions but significantly eases the process by separating spatial modulation from channel aggregation, hence making the long-range interaction efficient. Specifically, spatial modulation extracts global contexts from spatial, and channel aggregation condenses all global context features into the channel modulator, ultimately the aggregated contexts are fused into the final features. In addition, LSN allows guiding the network to focus on more refined image attributes by encoding local contextual information. By coupling two complementary components, SMN can capture both short- and long-range contexts for accurate image reconstruction. Extensive experimental results demonstrate that our SMN achieves state-of-the-art performance among the existing efficient SR methods with less complexity. Zhijian Wu, Jun Li 0033, Dingjiang Huang |
ACM Multimedia | 3 |
| 2023 | SFHN: Spatial-Frequency Domain Hybrid Network for Image Super-ResolutionabstractDeep convolutional neural networks (CNNs) have demonstrated tremendous success in image super-resolution (SR). According to the frequency principle, the vanilla CNNs fit the target function from low to high frequencies during the training process. It implies an implicit bias that CNNs tend to fit the training data by a low-frequency function. This is detrimental to SR task which essentially recovers the missing high-frequency cues from degraded images. To address this issue, a novel spatial-frequency domain hybrid network (SFHN) is proposed for image SR in this paper. More specifically, it contains multiple spatial-frequency domain hybrid convolution blocks (SFBlocks) for extracting both spatial and frequency components. In particular, frequency information is complementary to spatial one, mitigating the disadvantage of the spatial CNNs in capturing high-frequency content resulting from the inherent bias of the networks. In contrast to convolution blocks in the previous SR methods, the proposed SFBlock is configured with additional frequency-domain convolution branch driven by the Fourier transform, which allows to process frequency content. In addition, we introduce a spectral loss based on Fast Fourier Transform (FFT) to fully leverage the performance of our SFHN, as it prevents the loss of important frequency content during training. Experimental results on public benchmarks demonstrate that our SFHN achieves promising performance superior to the state-of-the-art methods. Zhijian Wu, Jun Li 0033, Chang Xu 0002, Dingjiang Huang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Improving Chinese Named Entity Recognition by Large-Scale Syntactic Dependency GraphabstractNamed entity recognition (NER) isa preliminary task in natural language processing (NLP). Recognizing Chinese named entities from unstructured texts is challenging due to the lack of word boundaries. Even if performing Chinese Word Segmentation (CWS) could help to determine word boundaries, it is still difficult to determine which words should be clustered together for entity identification, since entities are often composed of multiple-segmented words. As dependency relationships between segmented words could help to determine entity boundaries, it is crucial to employ information related to syntactic dependency relationships to improve NER performance. In this paper, we propose a novel NER model to learn information about syntactic dependency graphs with graph neural networks, and merge learned information into the classic Bidirectional Long Short-Term Memory (BiLSTM) - Conditional Random Field (CRF) NER scheme. In addition, we extract various kinds of task-specific hidden information from multiple CWS and part-of-speech (POS) tagging tasks, to further improve the NER model. We finally leverage multiple self-attention components to integrate multiple kinds of extracted information for named entity identification. Experimental results on three public benchmark datasets show that our model outperforms the state-of-the-art baselines in most scenarios. Peng Zhu 0002, Dawei Cheng, Fangzhou Yang, Yifeng Luo, Dingjiang Huang, Weining Qian, Aoying Zhou |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | TTPNet: A Neural Network for Travel Time Prediction Based on Tensor Decomposition and Graph EmbeddingabstractTravel time prediction of a given trajectory plays an indispensable role in intelligent transportation systems. Although many prior researches have struggled for accurate prediction results, most of them achieve inferior performance due to insufficient feature extraction of travel speed and road network structure from the trajectory data, which confirms the challenges involved in this topic. To overcome those issues, we propose a novel neuralNetworkforTravelTimePredictionbased on tensor decomposition and graph embedding, namedTTPNet, which can extract travel speed and representation of road network structure effectively from historical trajectories, as well as predict the travel time with better accuracy. Specifically,TTPNetconsists of three components: the first module (Travel Speed Features Layer) leverages non-negative tensor decomposition to restore travel speed distributions on different roads in the previous hour, and integrates a CNN-RNN model to extract both long-term and short-term travel speed features of the query trajectory; the second module (Road Network Structure Features Layer) utilizes graph embedding to generate the representation of local and global road network structure; the last module (Deep LSTM Prediction Layer) completes the final predicting task. Empirical results over two real-world large-scale datasets show that our proposedTTPNetmodel can achieve significantly better performance and remarkable robustness. Yibin Shen, Cheqing Jin, Jiaxun Hua, Dingjiang Huang |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2019 | StrokeNet: A Neural Painting Environment
Ningyuan Zheng, Dingjiang Huang |
ICLR (Poster) | 3 |
| 2018 | Combination Forecasting Reversion Strategy for Online Portfolio SelectionabstractMachine learning and artificial intelligence techniques have been applied to construct online portfolio selection strategies recently. A popular and state-of-the-art family of strategies is to explore the reversion phenomenon through online learning algorithms and statistical prediction models. Despite gaining promising results on some benchmark datasets, these strategies often adopt a single model based on a selection criterion (e.g., breakdown point) for predicting future price. However, such model selection is often unstable and may cause unnecessarily high variability in the final estimation, leading to poor prediction performance in real datasets and thus non-optimal portfolios. To overcome the drawbacks, in this article, we propose to exploit the reversion phenomenon by using combination forecasting estimators and design a novel online portfolio selection strategy, named Combination Forecasting Reversion (CFR), which outputs optimal portfolios based on the improved reversion estimator. We further present two efficient CFR implementations based on online Newton step (ONS) and online gradient descent (OGD) algorithms, respectively, and theoretically analyze their regret bounds, which guarantee that the online CFR model performs as well as the best CFR model in hindsight. We evaluate the proposed algorithms on various real markets with extensive experiments. Empirical results show that CFR can effectively overcome the drawbacks of existing reversion strategies and achieve the state-of-the-art performance. Dingjiang Huang, Shunchang Yu, Bin Li 0027, Steven C. H. Hoi, Shuigeng Zhou |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2016 | Robust Median Reversion Strategy for Online Portfolio SelectionabstractOnline portfolio selection has attracted increasing attention from data mining and machine learning communities in recent years. An important theory in financial markets is mean reversion, which plays a critical role in some state-of-the-art portfolio selection strategies. Although existing mean reversion strategies have been shown to achieve good empirical performance on certain datasets, they seldom carefully deal with noise and outliers in the data, leading to suboptimal portfolios, and consequently yielding poor performance in practice. In this paper, we propose to exploit the reversion phenomenon by using robust$L_1$-median estimators, and design a novel online portfolio selection strategy named “Robust Median Reversion” (RMR), which constructs optimal portfolios based on the improved reversion estimator. We examine the performance of the proposed algorithms on various real markets with extensive experiments. Empirical results show that RMR can overcome the drawbacks of existing mean reversion algorithms and achieve significantly better results. Finally, RMR runs in linear time, and thus is suitable for large-scale real-time algorithmic trading applications. Dingjiang Huang, Junlong Zhou, Bin Li 0027, Steven C. H. Hoi, Shuigeng Zhou |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Semi-Universal Portfolios with Transaction Costs
Dingjiang Huang, Yan Zhu 0008, Bin Li 0027, Shuigeng Zhou, Steven C. H. Hoi |
IJCAI | 1 |
| 2013 | Robust Median Reversion Strategy for On-Line Portfolio Selection
Dingjiang Huang, Junlong Zhou, Bin Li 0027, Steven C. H. Hoi, Shuigeng Zhou |
IJCAI | 1 |