VLDB 2026 Research / reviewers in the wild / expert
Haiping Wu
dblp:09/6392
· DBLP profile ↗
18ranked-venue papers
9as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-authorComputer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth FusionabstractWe present Florence-VL, a new family of multimodal large language models (MLLMs) with enriched visual representations produced by Florence-2 [45], a generative vision foundation model. Unlike the widely used CLIP-style vision transformer [35] trained by contrastive learning, Florence-2 can capture different levels and aspects of visual features, which are more versatile to be adapted to diverse downstream tasks. We propose a novel feature-fusion architecture and an innovative training recipe that effectively integrates Florence-2’s visual features into pre-trained LLMs, such as Phi 3.5 and LLama 3. In particular, we propose "depth-breath fusion (DBFusion)" to fuse the visual features extracted from different depths and under multiple prompts. Our model training is composed of end-to-end pretraining of the whole model followed by finetuning of the projection layer and the LLM, on a carefully designed recipe of diverse open-source datasets that include high-quality image captions and instruction-tuning pairs. Our quantitative analysis and visualization of Florence-VL’s visual features show its advantages over popular vision encoders on vision-language alignment, where the enriched depth and breath play important roles. Florence-VL achieves significant improvements over existing state-of-the-art MLLMs across various multi-modal and vision-centric benchmarks covering general VQA, perception, hallucination, OCR, Chart, knowledge-intensive understanding, etc. To facilitate future research, our models and the complete training recipe are open-sourced. https://github.com/JiuhaiChen/Florence-VL Jiuhai Chen, Haiping Wu, Dianqi Li, Jianfeng Gao 0001, Tianyi Zhou 0001, Bin Xiao 0004 |
CVPR | 3 |
| 2024 | Florence-2: Advancing a Unified Representation for a Variety of Vision TasksabstractWe introduce Florence-2, a novel vision foundation model with a unified, prompt-based representation for various computer vision and vision-language tasks. While existing large vision models excel in transfer learning, they struggle to perform diverse tasks with simple instructions, a capability that implies handling the complexity of various spatial hierarchy and semantic granularity. Florence-2 was designed to take text-prompt as task instructions and generate desirable results in text forms, whether it be captioning, object detection, grounding or segmentation. This multi-task learning setup demands large-scale, high-quality annotated data. To this end, we co-developed FLD-5B that consists of 5.4 billion comprehensive visual annotations on 126 million images, using an iterative strategy of automated image annotation and model refinement. We adopted a sequence-to-sequence structure to train Florence-2 to perform versatile and comprehensive vision tasks. Extensive evaluations on numerous tasks demonstrated Florence-2 to be a strong vision foundation model contender with un-precedented zero-shot and fine-tuning capabilities. Bin Xiao 0004, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng 0001, Ce Liu 0001, Lu Yuan 0001 |
CVPR | 2 |
| 2022 | From Individual to Whole: Reducing Intra-class Variance by Feature Aggregation
Zhaoxiang Zhang 0001, Chuanchen Luo, Haiping Wu, Yuntao Chen, Naiyan Wang, Chunfeng Song |
Int. J. Comput. Vis. | 3 |
| 2021 | Self-Supervised Attention-Aware Reinforcement LearningabstractVisual saliency has emerged as a major visualization tool for interpreting deep reinforcement learning (RL) agents. However, much of the existing research uses it as an analyzing tool rather than an inductive bias for policy learning. In this work, we use visual attention as an inductive bias for RL agents. We propose a novel self-supervised attention learning approach which can 1. learn to select regions of interest without explicit annotations, and 2. act as a plug for existing deep RL methods to improve the learning performance. We empirically show that the self-supervised attention-aware deep RL methods outperform the baselines in the context of both the rate of convergence and performance. Furthermore, the proposed self-supervised attention is not tied with specific policies, nor restricted to a specific scene. We posit that the proposed approach is a general self-supervised attention module for multi-task learning and transfer learning, and empirically validate the generalization ability of the proposed method. Finally, we show that our method learns meaningful object keypoints highlighting improvements both qualitatively and quantitatively. Haiping Wu, Khimya Khetarpal, Doina Precup |
AAAI | 1 |
| 2021 | Contrastive Learning of Image Representations with Cross-Video Cycle-ConsistencyabstractRecent works have advanced the performance of self-supervised representation learning by a large margin. The core among these methods is intra-image invariance learning. Two different transformations of one image instance are considered as a positive sample pair, where various tasks are designed to learn invariant representations by comparing the pair. Analogically, for video data, representations of frames from the same video are trained to be closer than frames from other videos, i.e. intra-video invariance. However, cross-video relation has barely been explored for visual representation learning. Unlike intra-video invariance, ground-truth labels of cross-video relation is usually unavailable without human labors. In this paper, we propose a novel contrastive learning method which explores the cross-video relation by using cycle-consistency for general image representation learning. This allows to collect positive sample pairs across different video instances, which we hypothesize will lead to higher-level semantics. We validate our method by transferring our image representation to multiple downstream tasks including visual object tracking, image classification, and action recognition. We show significant improvement over state-of-the-art contrastive learning methods. Project page is available at https://happywu.github.io/cycle_contrast_video. Haiping Wu, Xiaolong Wang 0004 |
ICCV | 1 |
| 2021 | CvT: Introducing Convolutions to Vision TransformersabstractWe present in this paper a new architecture, named Convolutional vision Transformer (CvT), that improves Vision Transformer (ViT) in performance and efficiency by introducing convolutions into ViT to yield the best of both de-signs. This is accomplished through two primary modifications: a hierarchy of Transformers containing a new convolutional token embedding, and a convolutional Transformer block leveraging a convolutional projection. These changes introduce desirable properties of convolutional neural networks (CNNs) to the ViT architecture (i.e. shift, scale, and distortion invariance) while maintaining the merits of Transformers (i.e. dynamic attention, global context, and better generalization). We validate CvT by conducting extensive experiments, showing that this approach achieves state-of-the-art performance over other Vision Transformers and ResNets on ImageNet-1k, with fewer parameters and lower FLOPs. In addition, performance gains are maintained when pretrained on larger datasets (e.g. ImageNet-22k) and fine-tuned to downstream tasks. Pretrained on ImageNet-22k, our CvT-W24 obtains a top-1 accuracy of 87.7% on the ImageNet-1k val set. Finally, our results show that the positional encoding, a crucial component in existing Vision Transformers, can be safely re-moved in our model, simplifying the design for higher resolution vision tasks. Code will be released at https://github.com/microsoft/CvT. Haiping Wu, Bin Xiao 0004, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan 0001, Lei Zhang 0001 |
ICCV | 1 |
| 2020 | 3D Human Pose Estimation via Explicit Compositional Depth MapsabstractIn this work, we tackle the problem of estimating 3D human pose in camera space from a monocular image. First, we propose to use densely-generated limb depth maps to ease the learning of body joints depth, which are well aligned with image cues. Then, we design a lifting module from 2D pixel coordinates to 3D camera coordinates which explicitly takes the depth values as inputs, and is aligned with camera perspective projection model. We show our method achieves superior performance on large-scale 3D pose datasets Human3.6M and MPI-INF-3DHP, and sets the new state-of-the-art. Haiping Wu, Bin Xiao 0004 |
AAAI | 1 |
| 2020 | A novel axle temperature forecasting method based on decomposition, reinforcement learning optimization and neural network
Hui Liu 0023, Chengming Yu, Chengqing Yu, Haiping Wu |
Adv. Eng. Informatics | 5 |
| 2019 | Sequence Level Semantics Aggregation for Video Object DetectionabstractVideo objection detection (VID) has been a rising research direction in recent years. A central issue of VID is the appearance degradation of video frames caused by fast motion. This problem is essentially ill-posed for a single frame. Therefore, aggregating features from other frames becomes a natural choice. Existing methods rely heavily on optical flow or recurrent neural networks for feature aggregation. However, these methods emphasize more on the temporally nearby frames. In this work, we argue that aggregating features in the full-sequence level will lead to more discriminative and robust features for video object detection. To achieve this goal, we devise a novel Sequence Level Semantics Aggregation (SELSA) module. We further demonstrate the close relationship between the proposed method and the classic spectral clustering method, providing a novel view for understanding the VID problem. We test the proposed method on the ImageNet VID and the EPIC KITCHENS dataset and achieve new state-of-the-art results. Our method does not need complicated postprocessing methods such as Seq-NMS or Tubelet rescoring, which keeps the pipeline simple and clean. Haiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
ICCV | 1 |
| 2018 | Simple Baselines for Human Pose Estimation and Tracking
Bin Xiao 0004, Haiping Wu |
ECCV (6) | 2 |
| 2018 | Semantic binary coding for visual recognition via joint concept-attribute modelling
Xing Xu 0001, Haiping Wu, Yang Yang 0002, Fumin Shen, Ning Xie 0003, Yanli Ji |
Multim. Tools Appl. | 2 |
| 2017 | Analysis on multiple factors influencing the lifetime of IGBTs of electric vehicles convertersabstractInsulated gate bipolar transistors (IGBTs) are one of the most critical components in the Electric Vehicles (EVs) because of their complex running environment, high power dissipation and frequent junction temperature fluctuation. Therefore, it's essential to predict the lifetime of IGBTs by considering the actual factors in applications. The aim of this paper is to present a lifetime prediction method during a long-term driving cycle by considering multiple influence factors of IGBTs. The lifetime prediction model includes motor drive system, electro-thermal coupling model, and lifetime analysis. Then, taking China as an example, the multiple influence factors to the result of lifetime prediction of IGBTs such as the ambient temperature, the driving cycle, and the degradation of thermal resistance are taken into consideration. The researches show that the percentage growth to lifetime prediction for these three factors is λT=108.4% λC=124.9%, and λR=69.9% respectively. The predicted lifetime of IGBTs considering three influence factors in China is 3.66×105km, which is nearly 5% lower than its original lifetime that is 3.87×105km. Yalei Sang, Haiping Wu, Jiajian Li |
IECON | 5 |
| 2015 | Remote sensing big data computing: Challenges and opportunities
Yan Ma 0001, Haiping Wu, Lizhe Wang 0001, Bormin Huang, Rajiv Ranjan 0001, Albert Y. Zomaya, Wei Jie |
Future Gener. Comput. Syst. | 2 |
| 2007 | Automatic Program Segment Similarity Detection in Targeted Program Performance ImprovementabstractTargeted optimization of program segments can provide an additional program speedup over the highest default optimization level, such as -O3 in GCC. The key challenge is how to automatically search for performance sensitive program segments in a given code, to which a customized set of optimization compiler options could be applied. In this paper we propose a method for automatic detection of performance sensitive program segments based on program segment similarity. First we create a proxy segment template database trained over a set of random input programs. The compiler identifies program segments by correlating them to the pre-build proxy segment templates using the syntax structure and architecture-dependent behavior similarity. We argue that the identified program segments can be custom optimized to improve the overall program performance. The method is evaluated on the Intel XScale PXA255 platform using randomly selected benchmarks. The experimental results show that our method can provide additional speedups over the highest optimization level in GCC 3.3 (-O3) for an arbitrary set of applications. Haiping Wu, Eunjung Park, Mihailo Kaplarevic, Yingping Zhang, Murat Bolat, Xiaoming Li 0010, Guang R. Gao |
IPDPS | 1 |
| 2006 | A study of the on-chip interconnection network for the IBM Cyclops64 multi-core architectureabstractThe designs of high-performance processor architectures are moving toward the integration of a large number of multiple processing cores on a single chip. The IBM Cyclops-64 (C64) is a petaflop supercomputer built on multi-core system-on-a-chip technology. Each C64 chip employs a multistage pipelined crossbar switch as its on-chip interconnection network to provide high bandwidth and low latency communication between the 160 thread processing cores, the on-chip SRAM memory banks, and other components. In this paper, we present a study of the architecture and performance of the C64 on-chip interconnection network through simulation. Our experimental results provide observations on the network behavior: (1) Dedicated channels can be created between any output port to input port of the C64 crossbar with latency as low as 7 cycles. The C64 crossbar has the potential reach the full hardware bandwidth, and exhibit a non-blocking behavior; (2) The C64 crossbar is a stable network; (3) The network logic design appears to provide a reasonable opportunity for sharing the channel bandwidth between traffic in either direction; (4) A simple circular neighbor arbitration scheme can achieve competitive performance level comparing to the complex segmented LRU (least recently used) matrix arbitration scheme without losing the fairness. (5) Application-driven benchmarks provide comparable results to synthetic workloads. Yingping Zhang, Taikyeong T. Jeong, Haiping Wu, Ronny Nitzsche, Guang R. Gao |
IPDPS | 4 |
| 2005 | An MMSE maximal shortening equalizer for 10GBASE-T Ether networksabstractRecently, IEEE is interested in specifying the next generation 10GBASE-T Ethernet network. Due to the limitations of the channel characteristics of Category 6 (CAT-6) unshielded twisted pair (UTP) cables, capacity of the channel is only slightly above the proposed 10 Gbps data rate. This poses a great challenge to the design of equalizers. This paper addresses the problem of designing a minimum mean-square error (MMSE) maximal shortening equalizer with decision feedback cancellation. Unlike traditional approaches such as the decision feedback equalizers (DFEs), the shortening equalizer optimally focuses the channel impulse response to a small number of taps. If a subsequent maximum likelihood sequence estimation (MLSE) equalizer is used, the performance loss caused by symbol-by-symbol hard decision of traditional equalizers can be reduced. However, due to the complexity involved in the shortening equalizers is generally prohibitive; it is difficult to apply them to systems using modulation with large constellations. To overcome this difficulty, we propose an equalizer that combines DFE and the shortening equalizer. Thus, with acceptable complexity, the shortened signal-to-noise ratio (SSNR) can be significantly improved. Haiping Wu, Mohsen Kavehrad |
GLOBECOM | 1 |
| 2005 | Madd Operation Aware Redundancy EliminationabstractOn general purpose computer architectures, the optimization of redundancy elimination almost always improves the cycle count. We argue that a specific consideration should be taken when applying this optimization to embedded architectures that feature multiply-add(MADD) instruction. This paper presents a redundancy elimination algorithm with MADD operation aware consideration. It produces optimized results for both code size and cycle count. The algorithm is integrated into KylinC compiler, a compiler for embedded systems developed at the University of Delaware. Experimental results demonstrate that the cycle counts of the benchmark programs are reduced on average 8% and the code sizes are reduced on average 5.27%. Haiping Wu, Ziang Hu, Joseph B. Manzano, Guang R. Gao |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2004 | A fixed radio system using MIMO cancellation and equalization and its performance modelabstractA theory of equalization and cancellation in digital data transmission over a MEMO multipath fading channel is presented. The theory extends the classic equalization and cancellation theory into matrix forms to mitigate intersymbol interference (ISI) and cross-polarization interference (CPI) by means of minimizing the overall mean-square error (MSE). We evaluate the performance of several configurations for a sample propagation model with several snapshots of fading events in computer simulations. Both MSE and average error probability are evaluated and compared. We find that the decision feedback structure demonstrates a better performance than the linear equalizer structure. Furthermore, less amount of MSE can be translated into smaller error probability values only in cases of relatively large signal-to-noise ratios. Haiping Wu, Mohsen Kavehrad |
PIMRC | 1 |