VLDB 2026 Research / reviewers in the wild / expert
Haitao Wang 0026
dblp:71/3863-26
· DBLP profile ↗
12ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-8394-6410ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VisRec: A Semi-Supervised Approach to Visibility Data Reconstruction in Radio AstronomyabstractRadio telescopes produce visibility data about celestial objects, but these data are sparse and noisy. As a result, images created on raw visibility data are of low quality. Recent studies have used deep learning models to reconstruct visibility data to get cleaner images. However, these methods rely on a substantial amount of labeled training data, which requires significant labeling effort from radio astronomers. Addressing this challenge, we propose VisRec, a model-agnostic semi-supervised learning approach to visibility data reconstruction in radio astronomy. Specifically, VisRec consists of both a supervised learning module and an unsupervised learning module. In the supervised learning module, we introduce a set of data augmentation functions to produce diverse visibility examples. In comparison, the unsupervised learning module in VisRec augments unlabeled data and uses reconstructions from non-augmented visibility as pseudo-labels for training. This hybrid approach allows VisRec to effectively leverage both labeled and unlabeled data. This way, VisRec performs well even when labeled data is scarce. Our evaluation results show that VisRec is applicable to various models, and outperforms all baseline methods in terms of reconstruction quality, robustness, and generalizability. Haitao Wang 0026, Qiong Luo 0001, Hejun Wu |
AAAI | 2 |
| 2025 | GalaxAlign: Mimicking Citizen Scientists' Multimodal Guidance for Galaxy Morphology AnalysisabstractGalaxy morphology analysis involves studying galaxies based on their shapes and structures. For such studies, fundamental tasks include identifying and classifying galaxies in astronomical images, as well as retrieving visually or structurally similar galaxies through similarity search. Existing methods either directly train domain-specific foundation models on large, annotated datasets or fine-tune vision foundation models on a smaller set of images. The former is effective but costly, while the latter is more resource-efficient but often yields lower accuracy. To address these challenges, we introduce GalaxAlign, a multimodal approach inspired by how citizen scientists identify galaxies in astronomical images by following textual descriptions and matching schematic symbols. Specifically, GalaxAlign employs a tri-modal alignment framework to align three types of data during fine-tuning: (1) schematic symbols representing galaxy shapes and structures, (2) textual labels for these symbols, and (3) galaxy images. By incorporating multimodal instructions, GalaxAlign eliminates the need for expensive pretraining and enhances the effectiveness of fine-tuning. Experiments on galaxy classification and similarity search demonstrate that our method effectively fine-tunes general pre-trained models for astronomical tasks by incorporating domain-specific multi-modal knowledge. Code is available at https://github.com/RapidsAtHKUST/GalaxAlign. Haitao Wang 0026, Qiong Luo 0001 |
ACM Multimedia | 2 |
| 2024 | Task-Aware Lipschitz Confidence Data Augmentation in Visual Reinforcement Learning From ImagesabstractVisual reinforcement learning is a technique that learns effective policies from image pixels. Data augmentation is widely adopted in visual reinforcement learning to improve the generalization of the learned policies as data augmentation increases data diversity. However, applying data augmentation to all pixels simultaneously results in a divergence in action distribution as well and degrades the training stability in turn. Additionally, existing methods compute task weights for each pixel and apply augmentation methods separately based on tasks. As a result, they require a significant amount of computational resources. To enhance both the training stability and computational efficiency in the computation of task weights, we propose a Task-Aware Lipschitz Confidence (TALC) data augmentation method for visual reinforcement tasks. TALC calculates the task-aware confidence of all pixels on the image at once, only enhancing low confidence pixels to increase data diversity. We have conducted experiments on DeepMind Control suite tasks and the results demonstrate that TALC not only improves the training efficiency, but also enhances the generalization ability during testing. Overall, TALC out-performs existing methods in most different visual control benchmarks. Haitao Wang 0026, Hejun Wu |
ICME | 2 |
| 2024 | ApmNet: Toward Generalizable Visual Continuous Control with Pre-trained Image Models
Haitao Wang 0026, Hejun Wu |
ECML/PKDD (3) | 1 |
| 2024 | Asynchronous Multi-Agent Reinforcement Learning for Collaborative Partial Charging in Wireless Rechargeable Sensor NetworksabstractOnline Scheduling for Partial charging with Multi-Mobile Chargers (OSPM) is critical for Wireless Rechargeable Sensor Networks (WRSNs) performing high-power monitoring tasks with a large number of simultaneous charging requests. However, existing studies for online scheduling assume full charging of sensors, leading to delays and inefficient resource utilization. Partially charging the sensors can improve scheduling efficiency and flexibility, but these studies focus on off-line scheduling, hindering dynamic decision-making. Multi-Agent Reinforcement Learning (MARL) is advantageous in online collaboration. Nevertheless, existing MARL methods assume synchronized actions, while Mobile Chargers (MCs) performing charging tasks asynchronously due to the difference in movement and charging times. On the other hand, hybrid actions are required to capture the simultaneous decision-making of MCs, involving sensor selection (discrete action) and energy allocation (continuous parameter). This introduces a circular dependency between a discrete action and its corresponding continuous parameter due to their interdependence. To deal with the above problems and address OSPM, we propose Asynchronous and Scalable Multi-agent Hybrid Proximal Policy Optimization (ASM-HPPO). The evaluation results not only indicate that our ASM-HPPO has advantages in terms of various performance metrics over existing schemes, but also demonstrate that our methods achieve higher stability and scalability. Yongheng Liang, Hejun Wu, Haitao Wang 0026 |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Improving Visual Reinforcement Learning with Discrete Information Bottleneck ApproachabstractContrastive learning has been used to learn useful low-dimensional state representations in visual reinforcement learning (RL). Such state representations substantially improve the sample efficiency of visual RL. Nevertheless, existing contrastive learning-based RL methods have the problem of unstable training. Such instability comes from the fact that contrastive learning requires an extremely large batch size (e.g., 4096 or larger), while current contrastive learning-based RL methods typically set a small batch size (e.g., 512). In this paper, we propose an approach of discrete information bottleneck (DIB) to address this problem. DIB applies the technique of discretization and information bottleneck to contrastive learning in representing the state with concise discrete representation. Using this discrete representation for policy learning results in more stable algorithm training and higher sample efficiency with a small batch size. We demonstrate the advantage of discrete state representation of DIB on several continuous control tasks in the DeepMind Control suite. In the experiments, DIB outperforms prior visual RL methods, both model-based and model-free, in terms of performance and sample efficiency. Haitao Wang 0026, Hejun Wu |
ECAI | 1 |
| 2023 | VMBRL3: A Simple Visual Model-Based Reinforcement Learning Framework for Continuous ControlabstractUnsupervised pre-training has demonstrated its potential for accurately constructing world models in visual model-based reinforcement learning (MBRL). However, such MBRL approaches exhibit limited generalizability, thereby limiting their practicality in diverse scenarios. These methods produce models that are restricted to the specific task they were trained on, and are not easily adaptable to other tasks. In this work, we introduce a powerful unsupervised pre-training reinforcement learning (RL) framework called VMBRL3, which improves the generalization ability of visual MBRL. VMBRL3 employs task-agnostic videos to pre-train both the autoencoder and world model without access to actions or rewards information. The fine-tuned world model can then be applied to a range of downstream reinforcement learning tasks, allowing for rapid adaptation to diverse environments and facilitating policy learning. We demonstrate that our framework significantly improves generalization ability in a variety of manipulation and locomotion tasks. Furthermore, VMBRL3 doubles the sample efficiency and overall performance compared to previous visual methods of MBRL. Haitao Wang 0026, Hejun Wu |
ECAI | 2 |
| 2021 | AMMASurv: Asymmetrical Multi-Modal Attention for Accurate Survival Analysis with Whole Slide Images and Gene Expression DataabstractThe use of multi-modal data such as the combination of whole slide images (WSIs) and gene expression data for survival analysis can lead to more accurate survival predictions. Previous multi-modal survival models are not able to efficiently excavate the intrinsic information within each modality. Moreover, previous methods regard the information from different modalities as similarly important so they cannot flexibly utilize the potential connection between the modalities. To address the above problems, we propose a new asymmetrical multi-modal method, termed as AMMASurv. Different from previous works, AMMASurv can effectively utilize the intrinsic information within every modality and flexibly adapts to the modalities of different importance. Encouraging experimental results demonstrate the superiority of our method over other state-of-the-art methods. Ziwang Huang, Haitao Wang 0026, Hejun Wu |
BIBM | 3 |
| 2021 | Integration of Patch Features Through Self-supervised Learning and Transformer for Survival Analysis on Whole Slide Images
Ziwang Huang, Haitao Wang 0026, Yuedong Yang, Hejun Wu |
MICCAI (8) | 4 |
| 2021 | Asymmetric Supervised Consistent and Specific Hashing for Cross-Modal RetrievalabstractHashing-based techniques have provided attractive solutions to cross-modal similarity search when addressing vast quantities of multimedia data. However, existing cross-modal hashing (CMH) methods face two critical limitations: 1) there is no previous work that simultaneously exploits the consistent or modality-specific information of multi-modal data; 2) the discriminative capabilities of pairwise similarity is usually neglected due to the computational cost and storage overhead. Moreover, to tackle the discrete constraints, relaxation-based strategy is typically adopted to relax the discrete problem to the continuous one, which severely suffers from large quantization errors and leads to sub-optimal solutions. To overcome the above limitations, in this article, we present a novel supervised CMH method, namely Asymmetric Supervised Consistent and Specific Hashing (ASCSH). Specifically, we explicitly decompose the mapping matrices into the consistent and modality-specific ones to sufficiently exploit the intrinsic correlation between different modalities. Meanwhile, a novel discrete asymmetric framework is proposed to fully explore the supervised information, in which the pairwise similarity and semantic labels are jointly formulated to guide the hash code learning process. Unlike existing asymmetric methods, the discrete asymmetric structure developed is capable of solving the binary constraint problem discretely and efficiently without any relaxation. To validate the effectiveness of the proposed approach, extensive experiments on three widely used datasets are conducted and encouraging results demonstrate the superiority of ASCSH over other state-of-the-art CMH methods. Min Meng 0001, Haitao Wang 0026, Jun Yu 0002, Jigang Wu |
IEEE Trans. Image Process. | 2 |
| 2019 | Robust Multi-View Hashing for Cross-Modal RetrievalabstractExisting hashing methods barely explore the information loss problem during learning the common semantic subspace, thus retrieval performance may be degraded. Besides, these methods mainly rely on the inter-modality or intra-modality correlations separately and fail to exploit the full structure reflected by these correlations. To address these problems, we present a novel cross-modal hashing method, namely Robust Multi-View Hashing (RMVH). To learn a robust latent semantic subspace, we enforce the learnt representations to well reconstruct original features such that more important information can be retained. To comprehensively exploit the relationship between representations of multiple modalities, we utilize Multi-View Learning to construct an affinity matrix to guide the learning of common latent semantic subspace, which can preserve both inter-modality and intra-modality similarities. Instead of relaxing the binary constraints, we leverage the label information to learn hash codes discretely which can avoid the large quantization error and preserve the semantic similarity. Experimental results on three benchmark datasets show that the proposed RMVH achieves superior performance compared with other state-of-the-art methods. Haitao Wang 0026, Min Meng 0001, Jigang Wu |
ICME | 1 |
| 2019 | Supervised Consistent and Specific HashingabstractMost existing methods seek for the common semantics using different projections for different modalities, which isolates the intrinsic relationships among different modalities. Besides, to avoid the large quantization error, some of them adopt the discrete cyclic coordinate descent schemes which are usually time-consuming. To address these issues, we present a novel hashing method, namely Supervised Consistent and Specific Hashing (SCSH), for cross-modal retrieval. We explicitly decompose the mapping matrices into consistent part and modality-specific ones. Specifically, consistency excavates the semantic shared by different modalities, whereas specificity captures private properties for each modality. Different from prior works, SCSH can discover the intrinsic semantic shared among different modalities more accurately. Moreover, by regressing the semantic labels to hash codes, SCSH can further promote the discriminative power of hash codes and significantly accelerate the hashing learning process. Extensive experiments on three widely used datasets demonstrate that the proposed SCSH outperforms other state-of-the-art methods. Haitao Wang 0026, Min Meng 0001, Jigang Wu |
ICME | 1 |