EDBT 2026 Demo / reviewers in the wild / expert
Haidong Zhu
dblp:09/1951
· DBLP profile ↗
21ranked-venue papers
14as first author
13since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 12 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 8 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Large Language Models are Good Prompt Learners for Low-Shot Image ClassificationabstractLow-shot image classification, where training images are limited or inaccessible, has benefited from recent progress on pretrained vision-language (VL) models with strong generalizability. e.g. CLIP. Prompt learning methods built with VL models generate text features from the class names that only have confined class-specific information. Large Language Models (LLMs), with their vast en-cyclopedic knowledge, emerge as the complement. Thus, in this paper, we discuss the integration of LLMs to enhance pretrained VL models, specifically on low-shot classification. However, the domain gap between language and vision blocks the direct application of LLMs. Thus, we propose LLaMp, Large Language Models as Prompt learners, that produces adaptive prompts for the CLIP text encoder, establishing it as the connecting bridge. Experiments show that, compared with other state-of-the-art prompt learning methods, LLaMP yields better performance on both zero-shot generalization and few-shot image classification, over a spectrum of 11 datasets. Code will be made available at: https://github.com/zhaohengz/LLaMP. Zhaoheng Zheng, Jingmin Wei, Xuefeng Hu, Haidong Zhu, Ramakant Nevatia |
CVPR | 4 |
| 2024 | SEAS: ShapE-Aligned Supervision for Person Re-IdentificationabstractWe introduce SEAS, using ShapE-Aligned Supervision, to enhance appearance-based person re-identification. When recognizing an individual's identity, existing methods primarily rely on appearance, which can be influenced by the background environment due to a lack of body shape awareness. Although some methods attempt to incorporate other modalities, such as gait or body shape, they encode the additional modality separately, resulting in extra computational costs and lacking an inherent connection with appearance. In this paper, we explore the use of implicit 3-D body shape representations as pixel-level guidance to augment the extraction of identity features with body shape knowledge, in addition to appearance. Using body shape as supervision, rather than as input, provides shapeaware enhancements without any increase in computational cost and delivers coherent integration with pixel-wise appearance features. Moreover, for video-based person reidentification, we align pixel-level features across frames with shape awareness to ensure temporal consistency. Our results demonstrate that incorporating body shape as pixel-level supervision reduces rank-1 errors by 1.4% for framebased and by 2.5% for video-based re-identification tasks, respectively, and can also be generalized to other existing appearance-based person re-identification methods. Haidong Zhu, Pranav Budhwant, Zhaoheng Zheng, Ramakant Nevatia |
CVPR | 1 |
| 2024 | CaesarNeRF: Calibrated Semantic Representation for Few-Shot Generalizable Neural Rendering
Haidong Zhu, Tianyu Ding, Ilya Zharkov, Ramakant Nevatia, Luming Liang |
ECCV (6) | 1 |
| 2024 | CAILA: Concept-Aware Intra-Layer Adapters for Compositional Zero-Shot LearningabstractIn this paper, we study the problem of Compositional Zero-Shot Learning (CZSL), which is to recognize novel attribute-object combinations with pre-existing concepts. Recent researchers focus on applying large-scale Vision-Language Pre-trained (VLP) models like CLIP with strong generalization ability. However, these methods treat the pre-trained model as a black box and focus on pre- and post-CLIP operations, which do not inherently mine the semantic concept between the layers inside CLIP. We propose to dive deep into the architecture and insert adapters, a parameter-efficient technique proven to be effective among large language models, into each CLIP encoder layer. We further equip adapters with concept awareness so that concept-specific features of "object", "attribute", and "composition" can be extracted. We assess our method on four popular CZSL datasets, MIT-States, C-GQA, UT-Zappos, and VAW-CZSL, which shows state-of-the-art performance compared to existing methods on all of them. Zhaoheng Zheng, Haidong Zhu, Ramakant Nevatia |
WACV | 2 |
| 2024 | ShARc: Shape and Appearance Recognition for Person Identification In-the-wildabstractIdentifying individuals in unconstrained video settings is a valuable yet challenging task in biometric analysis due to variations in appearances, environments, degradations, and occlusions. In this paper, we present ShARc, a multimodal approach for video-based person identification in uncontrolled environments that emphasizes 3-D body shape, pose, and appearance. We introduce two encoders: a Pose and Shape Encoder (PSE) and an Aggregated Appearance Encoder (AAE). PSE encodes the body shape via binarized silhouettes, skeleton motions, and 3-D body shape, while AAE provides two levels of temporal appearance feature aggregation: attention-based feature aggregation and averaging aggregation. For attention-based feature aggregation, we employ spatial and temporal attention to focus on key areas for person distinction. For averaging aggregation, we introduce a novel flattening layer after averaging to extract more distinguishable information and reduce overfitting of attention. We utilize centroid feature averaging for gallery registration. We demonstrate significant improvements over existing state-of-the-art methods on public datasets, including CCVID, MEVID, and BRIAR. Haidong Zhu, Wanrong Zheng, Zhaoheng Zheng, Ramakant Nevatia |
WACV | 1 |
| 2023 | AG-ReID 2023: Aerial-Ground Person Re-identification Challenge ResultsabstractPerson re-identification (Re-ID) on aerial-ground platforms has emerged as an intriguing topic within computer vision, presenting a plethora of unique challenges. Highflying altitudes of aerial cameras make persons appear differently in terms of viewpoints, poses, and resolution compared to the images of the same person viewed from ground cameras. Despite its potential, few algorithms have been developed for person re-identification on aerial-ground data, mainly due to the absence of comprehensive datasets. In response, we have collected a large-scale dataset and organized the Aerial-Ground person Re-IDentification Challenge (AG-ReID2023) to foster advancements in the field. The dataset comprises 100,502 images with 1,615 unique identities, including 51,530 training images featuring 807 identities. The test set is divided into two subsets: Aerial to Ground (808 ids, 4,348 query images, 19,259 gallery images) and Ground to Aerial (808 ids, 4,151 query images, 21,214 gallery images). In addition, we manually annotate individuals with their matching IDs across cameras and provide 15 soft attribute labels. The AG-ReID2023 Challenge in conjunction with the 7thIEEE International Joint Conference on Biometrics (IJCB) has garnered interest from numerous institutes, resulting in the submission of five distinct algorithms. We provide an in-depth examination of the evaluation outcomes and present our findings from the contest. For additional details, kindly refer to the official website1.1https://agreid23.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Dana Michalski, Debayan Deb, Mahak Kothari, Manisha Saini, Dawei Du, Scott McCloskey, Gabriel Bertocco, Fernanda A. Andaló, Terrance E. Boult, Anderson Rocha 0001, Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia, Zaigham A. Randhawa, Sinan Sabri, Gianfranco Doretto |
IJCB | 18 |
| 2023 | GaitRef: Gait Recognition with Refined Sequential SkeletonsabstractIdentifying humans with their walking sequences, known as gait recognition, is a useful biometric understanding task as it can be observed from a long distance and does not require cooperation from the subject. Two common modalities used for representing the walking sequence of a person are silhouettes and joint skeletons. Silhouette sequences, which record the boundary of the walking person in each frame, may suffer from the variant appearances from carried-on objects and clothes of the person. Framewise joint detections are noisy and introduce some jitters that are not consistent with sequential detections. In this paper, we combine the silhouettes and skeletons and refine the framewise joint predictions for gait recognition. With temporal information from the silhouette sequences. We show that the refined skeletons can improve gait recognition performance without extra annotations. We compare our methods on four public datasets, CASIA-B, OUMVLP, Gait3D and GREW, and show state-of-the-art performance. Haidong Zhu, Wanrong Zheng, Zhaoheng Zheng, Ramakant Nevatia |
IJCB | 1 |
| 2023 | Multimodal Neural Radiance FieldabstractThis paper addresses the challenge of reconstructing a scene with a neural radiance field (NeRF) for robot vision and scene understanding using multiple modalities. Researchers have introduced the use of NeRF to represent an object for synthesizing and rendering novel views of complex scenes by optimizing a 3-D radiance field for ray casting and rendering for 2-D RGB images. However, using RGB images alone introduces additional geometry ambiguities with transparent objects or complex scenes and cannot accurately depict the 3-D shapes. We discuss and solve this problem and use multiple modalities as input for the same NeRF model to build a multimodal NeRF by incorporating point clouds and infrared image supervision to prevent such bias. In contrast to RGB images, infrared images and point clouds are typically taken by separate cameras that cannot be aligned with the RGB camera. We further introduce the alignment of different modalities based on point cloud registration to estimate the relative transformation matrices between them before training a NeRF model with multiple modalities. We evaluate our model on chosen scenes from the ScanNet and M2DGR datasets and demonstrate that it outperforms existing state-of-the-art methods. Haidong Zhu, Yuyin Sun, Jiajia Luo, Nan Qiao 0009, Ramakant Nevatia, Cheng-Hao Kuo |
ICRA | 1 |
| 2023 | Gait Recognition Using 3-D Human Body Shape InferenceabstractGait recognition, which identifies individuals based on their walking patterns, is an important biometric technique since it can be observed from a distance and does not require the subject’s cooperation. Recognizing a person’s gait is difficult because of the appearance variants in human silhouette sequences produced by varying viewing angles, carrying objects, and clothing. Recent research has produced a number of ways for coping with these variants. In this paper, we present the usage of inferring 3-D body shapes distilled from limited images, which are, in principle, invariant to the specified variants. Inference of 3-D shape is a difficult task, especially when only silhouettes are provided in a dataset. We provide a method for learning 3-D body inference from silhouettes by transferring knowledge from 3-D shape prior from RGB photos. We use our method on multiple existing state-of-the-art gait baselines and obtain consistent improvements for gait identification on two public datasets, CASIA-B and OUMVLP, on several variants and settings, including a new setting of novel views not seen during training. Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia |
WACV | 1 |
| 2022 | Self-Supervised Learning for Sentiment Analysis via Image-Text MatchingabstractThere is often a resemblance in the sentiment expressed in social media posts (text) and their accompanying images. In this paper, We leverage this sentiment congruence for self-supervised representation learning for sentiment analysis. By teaching the model to pair an image with its corresponding social media post, the model can learn a representation capturing sentiment features from the image and text without supervision. We then use the pre-trained encoder for feature extraction for sentiment analysis in downstream tasks. We show significant improvement and good transferability for sentiment classification in addition to robustness in performance when available data decreases on public datasets (B-T4SA and IMDb Movie Review). With this work, we demonstrate the effectiveness of self-supervised learning through cross-modal matching for sentiment analysis. Haidong Zhu, Zhaoheng Zheng, Mohammad Soleymani 0001, Ramakant Nevatia |
ICASSP | 1 |
| 2022 | OPEN: Order-preserving Pointcloud Encoder Decoder Network for Body Shape RefinementabstractImage-based 3-D human body shape estimation and reconstruction have shown significant improvement by using deep neural networks. Compared with reconstructing from a single image, reconstructing 3-D human body shapes from video or image sequences requires high precision and dense correspondences between the keypoints of the reconstructed shape sequence. Existing methods cannot achieve both high accuracy and keep the dense correspondence between different shapes after reconstruction. In this paper, we propose a method named Order-preserving Point cloud Encoder-decoder Network to refine the reconstructed human body shape from SMPL with the assistance of RGB images while preserving its original dense correspondence. We further introduce using 2-D RGB images as weak supervision when 3-D labels are not available. We assess our methods on the public dataset and show improved results compared with the baseline methods. Haidong Zhu, Ramakant Nevatia |
ICPR | 1 |
| 2022 | Temporal Shift and Attention Modules for Graphical Skeleton Action RecognitionabstractSkeletons, consisting of joint positions and connections between them, are an important representation for modeling human bodies in image frames. Compared with understanding RGB videos, recognizing actions from the skeletons removes the biases of background and body shapes. Researchers use spatial-temporal graphs to model the skeleton sequences. These methods weigh all frames in the sequence equally even though many of the frames may not be useful for action and prediction and dilute the influence of important frames. Also, the temporal graph focuses on understanding only the low-level feature of the joints for the motion of the skeleton. In this paper, we introduce two modules, temporal shift module and temporal attention module that can be added to graph convolution networks for skeleton action recognition. Temporal attention module focuses on keyframes for making predictions, and temporal shift module helps to exchange the high-level features between different frames along the temporal dimension besides the local patterns. We evaluate the two modules with two existing skeleton action recognition networks, ST-GCN and MS-G3D, on three public datasets and show better results than the original methods. Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia |
ICPR | 1 |
| 2021 | Utilizing Every Image Object for Semi-supervised Phrase GroundingabstractPhrase grounding models localize an object in the image given a referring expression. The annotated language queries available during training are limited, which also limits the variations of language combinations that a model can see during training. In this paper, we study the case applying objects without labeled queries for training the semi-supervised phrase grounding. We propose to use learned location and subject embedding predictors (LSEP) to generate the corresponding language embeddings for objects lacking annotated queries in the training set. With the assistance of the detector, we also apply LSEP to train a grounding model on images without any annotation. We evaluate our method based on MAttNet on three public datasets: RefCOCO, RefCOCO+, and RefCOCOg. We show that our predictors allow the grounding system to learn from the objects without labeled queries and improve accuracy by 34.9% relatively with the detection results. Haidong Zhu, Arka Sadhu, Zhaoheng Zheng, Ramakant Nevatia |
WACV | 1 |
| 2020 | Curriculum DeepSDF
Yueqi Duan, Haidong Zhu, He Wang 0010, Li Yi 0001, Ramakant Nevatia, Leonidas J. Guibas |
ECCV (8) | 2 |
| 2019 | Biologically-Constrained Graphs for Global Connectomics ReconstructionabstractMost current state-of-the-art connectome reconstruction pipelines have two major steps: initial pixel-based segmentation with affinity prediction and watershed transform, and refined segmentation by merging over-segmented regions. These methods rely only on local context and are typically agnostic to the underlying biology. Since a few merge errors can lead to several incorrectly merged neuronal processes, these algorithms are currently tuned towards over-segmentation producing an overburden of costly proofreading. We propose a third step for connectomics reconstruction pipelines to refine an over-segmentation using both local and global context with an emphasis on adhering to the underlying biology. We first extract a graph from an input segmentation where nodes correspond to segment labels and edges indicate potential split errors in the over-segmentation. In order to increase throughput and allow for large-scale reconstruction, we employ biologically inspired geometric constraints based on neuron morphology to reduce the number of nodes and edges. Next, two neural networks learn these neuronal shapes to further aid the graph construction process. Lastly, we reformulate the region merging problem as a graph partitioning one to leverage global context. We demonstrate the performance of our approach on four real-world connectomics datasets with an average variation of information improvement of 21.3%. Brian Matejek, Daniel Haehn, Haidong Zhu, Donglai Wei 0001, Toufiq Parag, Hanspeter Pfister |
CVPR | 3 |
| 2019 | Pick-and-Learn: Automatic Quality Evaluation for Noisy-Labeled Image Segmentation
Haidong Zhu, Jialin Shi, Ji Wu 0002 |
MICCAI (6) | 1 |
| 2012 | SISO APP Searches in Lattices With Tanner GraphsabstractAn efficient, low-complexity, soft-output detector for general lattices is presented, based on their Tanner graph (TG) representations. Closest-point searches in lattices can be performed as nonbinary belief propagation on associated TGs; soft-information output is naturally generated in the process; the algorithm requires no backtrack (cf. classic sphere decoding), and extracts extrinsic information. A lattice's coding gain enables equivalence relations between lattice points, which can be thereby partitioned in cosets. Total and extrinsic a posteriori probabilities at the detector's output further enable the use of soft detection information in iterative schemes. The algorithm is illustrated via two scenarios that transmit a 32-point, uncoded super-orthogonal (SO) constellation for multiple-input multiple-output (MIMO) channels, carved from an 8-dimensional nonorthogonal latticeD4⊕D4: it achieves maximum likelihood performance in quasistatic fading; and, performs close to interference-free transmission, and identically to list sphere decoding, in independent fading with coordinate interleaving and iterative equalization and detection. Latter scenario outperforms former despite absence of forward error correction coding-because the inherent lattice coding gain allows for the refining of extrinsic information. The lattice constellation is the same as the one employed in the SO space-time trellis codes first introduced for 2 × 2 MIMO by Ionescu et al., then independently by Jafarkhani and Seshadri. Algorithmic complexity is log-linear in lattice dimensionality versus cubic in classic sphere decoders. Dumitru Mihai Ionescu, Haidong Zhu |
IEEE Trans. Inf. Theory | 2 |
| 2005 | MIMO detection using Markov chain Monte Carlo techniques for near-capacity performanceabstractIn this paper, we develop a new soft-in soft-out (SISO) multiple-input multiple-output (MIMO) detection algorithm using the Markov chain Monte Carlo (MCMC) simulation techniques and study its performance when applied to a MIMO communication system. Comparison with the best MIMO detection algorithm in the current literature, sphere decoding, shows that the proposed detection algorithm can improve the gap between the present results and the capacity by as much as 2 dB. Haidong Zhu, Zhenning Shi, Behrouz Farhang-Boroujeny |
ICASSP (3) | 1 |
| 2005 | On performance of sphere decoding and Markov chain Monte Carlo detection methodsabstractIn a recent work, it has been found that the suboptimum detectors that are based on Markov chain Monte Carlo (MCMC) simulation techniques perform significantly better than their sphere decoding (SD) counterparts. In this letter, we explore the sources of this difference and show that a modification to existing sphere decoders can result in some improvement in their performance, even though they still fall short when compared with the MCMC detector. We also present a novel SD detector that is an exact realization of max-log-MAP detector. We call this exact max-log SD detector. Comparison of the results of this detector with those of the max-log version of the MCMC detector reveals that the latter is near optimal. Haidong Zhu, Behrouz Farhang-Boroujeny, Rong-Rong Chen |
IEEE Signal Process. Lett. | 1 |
| 2004 | Markov chain Monte Carlo techniques in iterative detectors: a novel approach based on Monte Carlo integrationabstractWe present a novel soft-in soft-out (SISO) detection scheme based on Markov-chain Monte-Carlo (MCMC) simulations. The proposed detector is applicable to both synchronous multiuser and multiple-input multiple-output (MIMO) communication systems. Unlike previous publications on the subject, we use Monte Carlo integration technique to arrive at the receiver structure. The proposed multiuser/MIMO detector is found to follow the Rao-Blackwell formulation and found to result in much more reliable systems than those reported so far in the literature. In particular, through computer simulations, we demonstrate that the proposed receiver can easily handle the cases of overloaded systems, i.e. when in a multiuser set-up the number of users is larger than the spreading gain, or when in a MIMO set-up there are more transmit antennas than receive antennas. Zhenning Shi, Haidong Zhu, Behrouz Farhang-Boroujeny |
GLOBECOM | 2 |
| 2004 | Capacity of pilot-aided MIMO communication systemsabstractIn this paper a MIMO system with transmit and receive antenna and the channel model with two methods of pilot-aided (PA) schemes are considered. The first scheme is pilot insertion (PI) scheme, where pilots are time multiplexed with data and the second scheme pilot embedding (PE) scheme, where pilots are added and transmitted concurrently with data. The study shows that PI scheme is better than PE scheme. To obtain the capacity the maximum-likelihood channel estimation is used. Then the discrete and continuous capacity of a MIMO channel with imperfect channel estimates is evaluated. Haidong Zhu, Rong-Rong Chen, Behrouz Farhang-Boroujeny |
ISIT | 1 |