EDBT 2026 Demo / reviewers in the wild / expert
Xianzhi Du
dblp:32/10268
· DBLP profile ↗
25ranked-venue papers
8as first author
16since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 5 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuningabstractWe present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture, MM1.5 adopts a data-centric approach to model training, systematically exploring the impact of diverse data mixtures across the entire model training lifecycle. This includes high-quality OCR data and synthetic captions for continual pre-training, as well as an optimized visual instruction-tuning data mixture for supervised fine-tuning. Our models range from 1B to 30B parameters, encompassing both dense and mixture-of-experts (MoE) variants, and demonstrate that careful data curation and training strategies can yield strong performance even at small scales (1B and 3B). Additionally, we introduce two specialized variants: MM1.5-Video, designed for video understanding, and MM1.5-UI, tailored for mobile UI understanding. Through extensive empirical studies and ablations, we provide detailed insights into the training processes and decisions that inform our final designs, offering valuable guidance for future research in MLLM development. Haotian Zhang 0005, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, Bowen Zhang 0002, Yanghao Li, Sam Dodge, Keen You, Aleksei Timofeev, Hong-You Chen, Jean-Philippe Fauconnier, Zhengfeng Lai, Haoxuan You |
ICLR | 8 |
| 2024 | VeCLIP: Improving CLIP Training via Visual-Enriched Captions
Zhengfeng Lai, Haotian Zhang 0005, Bowen Zhang 0002, Haoping Bai, Aleksei Timofeev, Xianzhi Du, Zhe Gan, Jiulong Shan, Chen-Nee Chuah, Yinfei Yang |
ECCV (42) | 7 |
| 2024 | MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang 0002, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Anton Belyi, Haotian Zhang 0005, Karanjeet Singh 0003, Doug Kang, Hongyu Hè, Max Schwarzer, Tom Gunter, Xiang Kong, Aonan Zhang, Nan Du 0002, Tao Lei 0001, Sam Wiseman, Mark Lee 0003, Ruoming Pang, Peter Grasch, Alexander Toshev, Yinfei Yang |
ECCV (29) | 8 |
| 2024 | Guiding Instruction-based Image Editing via Multimodal Large Language ModelsabstractInstruction-based image editing improves the controllability and flexibility of image manipulation via natural commands without elaborate descriptions or regional masks. However, human instructions are sometimes too brief for current methods to capture and follow. Multimodal large language models (MLLMs) show promising capabilities in cross-modal understanding and visual-aware response generation via LMs. We investigate how MLLMs facilitate edit instructions and present MLLM-Guided Image Editing (MGIE). MGIE learns to derive expressive instructions and provides explicit guidance. The editing model jointly captures this visual imagination and performs manipulation through end-to-end training. We evaluate various aspects of Photoshop-style modification, global photo optimization, and local editing. Extensive experimental results demonstrate that expressive instructions are crucial to instruction-based image editing, and our MGIE can lead to a notable improvement in automatic metrics and human evaluation while maintaining competitive inference efficiency. Tsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang, Yinfei Yang, Zhe Gan |
ICLR | 3 |
| 2024 | Compressing LLMs: The Truth is Rarely Pure and Never SimpleabstractDespite their remarkable achievements, modern Large Language Models (LLMs) encounter exorbitant computational and memory footprints. Recently, several works have shown significant success in *training-free* and *data-free* compression (pruning and quantization) of LLMs achieving 50-60\% sparsity and reducing the bit-width down to 3 or 4 bits per weight, with negligible perplexity degradation over the uncompressed baseline. As recent research efforts are focused on developing increasingly sophisticated compression methods, our work takes a step back, and re-evaluates the effectiveness of existing SoTA compression methods, which rely on a fairly simple and widely questioned metric, perplexity (even for dense LLMs). We introduce **K**nowledge-**I**ntensive **C**ompressed LLM Benchmar**K** **(LLM-KICK)**, a collection of carefully-curated tasks to re-define the evaluation protocol for compressed LLMs, which have significant alignment with their dense counterparts, and perplexity fail to capture subtle change in their true capabilities. LLM-KICK unveils many favorable merits and unfortunate plights of current SoTA compression methods: all pruning methods suffer significant performance degradation, sometimes at trivial sparsity ratios (*e.g.*, 25-30\%), and fail for N:M sparsity on knowledge-intensive tasks; current quantization methods are more successful than pruning; yet, pruned LLMs even at $\geq 50$\% sparsity are robust in-context retrieval and summarization systems; among others. LLM-KICK is designed to holistically access compressed LLMs' ability for language understanding, reasoning, generation, in-context retrieval, in-context summarization, *etc.* We hope our study can foster the development of better LLM compression methods. The reproduced codes are available at https://github.com/VITA-Group/llm-kick. Ajay Jaiswal, Zhe Gan, Xianzhi Du, Zhangyang Wang, Yinfei Yang |
ICLR | 3 |
| 2024 | MOFI: Learning Image Representations from Noisy Entity Annotated ImagesabstractWe present MOFI, Manifold OF Images, a new vision foundation model designed to learn image representations from noisy entity annotated images. MOFI differs from previous work in two key aspects: 1. pre-training data, and 2. training recipe. Regarding data, we introduce a new approach to automatically assign entity labels to images from noisy image-text pairs. Our approach involves employing a named entity recognition model to extract entities from the alt-text, and then using a CLIP model to select the correct entities as labels of the paired image. It's a simple, cost-effective method that can scale to handle billions of web-mined image-text pairs. Through this method, we have created Image-to-Entities (I2E), a new dataset with 1 billion images and 2 million distinct entities, covering rich visual concepts in the wild. Building upon the I2E dataset, we study different training recipes like supervised pre-training, contrastive pre-training, and multi-task learning. For constrastive pre-training, we treat entity names as free-form text, and further enrich them with entity descriptions. Experiments show that supervised pre-training with large-scale fine-grained entity labels is highly effective for image retrieval tasks, and multi-task training further improves the performance. The final MOFI model achieves 86.66\% mAP on the challenging GPR1200 dataset, surpassing the previous state-of-the-art performance of 72.19% from OpenAI's CLIP model. Further experiments on zero-shot and linear probe image classification also show that MOFI outperforms a CLIP model trained on the original image-text data, demonstrating the effectiveness of the I2E dataset in learning strong image representations. We release our code and model weights at https://github.com/apple/ml-mofi. Aleksei Timofeev, Chen Chen 0005, Bowen Zhang 0002, Kun Duan, Shuangning Liu, Yantao Zheng, Jonathon Shlens, Xianzhi Du, Yinfei Yang |
ICLR | 9 |
| 2024 | Ferret: Refer and Ground Anything Anywhere at Any GranularityabstractWe introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions. To unify referring and grounding in the LLM paradigm, Ferret employs a novel and powerful hybrid region representation that integrates discrete coordinates and continuous features jointly to represent a region in the image. To extract the continuous features of versatile regions, we propose a spatial-aware visual sampler, adept at handling varying sparsity across different shapes. Consequently, Ferret can accept diverse region inputs, such as points, bounding boxes, and free-form shapes. To bolster the desired capability of Ferret, we curate GRIT, a comprehensive refer-and-ground instruction tuning dataset including 1.1M samples that contain rich hierarchical spatial knowledge, with an additional 130K hard negative data to promote model robustness. The resulting model not only achieves superior performance in classical referring and grounding tasks, but also greatly outperforms existing MLLMs in region-based and localization-demanded multimodal chatting. Our evaluations also reveal a significantly improved capability of describing image details and a remarkable alleviation in object hallucination. Haoxuan You, Haotian Zhang 0005, Zhe Gan, Xianzhi Du, Bowen Zhang 0002, Liangliang Cao, Shih-Fu Chang, Yinfei Yang |
ICLR | 4 |
| 2024 | Empowering Unsupervised Domain Adaptation with Large-scale Pre-trained Vision-Language ModelsabstractUnsupervised Domain Adaptation (UDA) aims to leverage the labeled source domain to solve the tasks on the unlabeled target domain. Traditional UDA methods face the challenge of the tradeoff between domain alignment and semantic class discriminability, especially when a large domain gap exists between the source and target domains. The efforts of applying large-scale pre-training to bridge the domain gaps remain limited. In this work, we propose that Vision-Language Models (VLMs) can empower UDA tasks due to their training pattern with language alignment and their large-scale pre-trained datasets. For example, CLIP and GLIP have shown promising zero-shot generalization in classification and detection tasks. However, directly fine-tuning these VLMs into downstream tasks may be computationally expensive and not scalable if we have multiple domains that need to be adapted. Therefore, in this work, we first study an efficient adaption of VLMs to preserve the original knowledge while maximizing its flexibility for learning new knowledge. Then, we design a domain-aware pseudo-labeling scheme tailored to VLMs for domain disentanglement. We show the superiority of the proposed methods in four UDA-classification and two UDA-detection benchmarks, with a significant improvement (+9.9%) on DomainNet. Zhengfeng Lai, Haoping Bai, Haotian Zhang 0005, Xianzhi Du, Jiulong Shan, Yinfei Yang, Chen-Nee Chuah |
WACV | 4 |
| 2023 | AdaMV-MoE: Adaptive Multi-Task Vision Mixture-of-ExpertsabstractSparsely activated Mixture-of-Experts (MoE) is becoming a promising paradigm for multi-task learning (MTL). Instead of compressing multiple tasks’ knowledge into a single model, MoE separates the parameter space and only utilizes the relevant model pieces given task type and its input, which provides stabilized MTL training and ultra-efficient inference. However, current MoE approaches adopt a fixed network capacity (e.g., two experts in usual) for all tasks. It potentially results in the over-fitting of simple tasks or the under-fitting of challenging scenarios, especially when tasks are significantly distinctive in their complexity. In this paper, we propose an adaptive MoE framework for multi-task vision recognition, dubbed AdaMV-MoE. Based on the training dynamics, it automatically determines the number of activated experts for each task, avoiding the laborious manual tuning of optimal model size. To validate our proposal, we benchmark it on ImageNet classification and COCO object detection & instance segmentation which are notoriously difficult to learn in concert, due to their discrepancy. Extensive experiments across a variety of vision transformers demonstrate a superior performance of AdaMV-MoE, compared to MTL with a shared backbone and the recent state-of-the-art (SoTA) MTL MoE approach. Codes are available online: https://github.com/google-research/google-research/tree/master/moe_mtl. Tianlong Chen 0001, Xuxi Chen, Xianzhi Du, Abdullah Rashwan, Huizhong Chen, Zhangyang Wang, Yeqing Li |
ICCV | 3 |
| 2023 | ISAA: Boost Repair Process by Constructing the Degree Constrained Optimal Repair Tree for Erasure-coded SystemsabstractTo ensure data reliability, large-scale distributed systems usually adopt erasure codes to restore failed nodes. However, existing erasure-coded repair strategies will cause heavy network traffics, which will increase the repair time. In order to boost the repair process, we consider optimizing the repair path which can be abstracted to a repair tree. Moreover, we add a degree constraint to each node to avoid local congestion. In this paper, we study the degree constrained optimal repair tree, which is an NP-hard problem. Current methods cannot find the optimal solution in a short time in complex non-uniform bandwidth networks. To obtain the optimal repair tree, an improved simulated annealing algorithm (ISAA) based on the Prufer code representation is proposed in this paper. In addition, we simulate the repair process of erasure codes in a non-uniform bandwidth network and experiments show that the repair time reduction can reach up to 66.4% and 88.6% with ISAA over Repair Pipelining and Partial-Parallel-Repair. Xianzhi Du, Bing Zhu 0003, Zhihang Deng, Kenneth W. Shum, Weiping Wang 0003 |
ICPADS | 1 |
| 2022 | A Simple Single-Scale Vision Transformer for Object Detection and Instance Segmentation
Wuyang Chen 0001, Xianzhi Du, Lucas Beyer, Xiaohua Zhai, Tsung-Yi Lin, Huizhong Chen, Xiaodan Song, Zhangyang Wang, Denny Zhou |
ECCV (10) | 2 |
| 2022 | A Genetic Algorithm-based Construction of Fractional Repetition CodesabstractFractional repetition (FR) codes form a special family of minimum bandwidth regenerating codes, characterized by an uncoded exact repair process. This low-complexity repair process benefits from the two-layer encoding structure consisting of an outer maximum distance separable code and an inner repetition code. However, it sacrifices certain storage efficiency, or supported file size. In order to improve the supported file size, we present in this paper a genetic algorithm-based construction of FR codes. In particular, we first review the existence of FR codes and introduce a general construction of FR codes. We implement an iterative optimization process of mimicking biological evolution on constructed FR codes. Through simulations, it is shown that the supported file size can be improved for different network scales. Zhihang Deng, Bing Zhu 0003, Xianzhi Du, Kenneth W. Shum |
GLOBECOM | 3 |
| 2022 | Auto-scaling Vision Transformers without Training
Wuyang Chen 0001, Wei Huang 0034, Xianzhi Du, Xiaodan Song, Zhangyang Wang, Denny Zhou |
ICLR | 3 |
| 2022 | Provable Stochastic Optimization for Global Contrastive Learning: Small Batch Does Not Harm PerformanceabstractIn this paper, we study contrastive learning from an optimization perspective, aiming to analyze and address a fundamental issue of existing contrastive learning methods that either rely on a large batch size or a large dictionary of feature vectors. We consider a global objective for contrastive learning, which contrasts each positive pair with all negative pairs for an anchor point. From the optimization perspective, we explain why existing methods such as SimCLR require a large batch size in order to achieve a satisfactory result. In order to remove such requirement, we propose a memory-efficient Stochastic Optimization algorithm for solving the Global objective of Contrastive Learning of Representations, named SogCLR. We show that its optimization error is negligible under a reasonable condition after a sufficient number of iterations or is diminishing for a slightly different global contrastive objective. Empirically, we demonstrate that SogCLR with small batch size (e.g., 256) can achieve similar performance as SimCLR with large batch size (e.g., 8192) on self-supervised learning task on ImageNet-1K. We also attempt to show that the proposed optimization technique is generic and can be applied to solving other contrastive losses, e.g., two-way contrastive losses for bimodal contrastive learning. The proposed method is implemented in our open-sourced library LibAUC (www.libauc.org). Zhuoning Yuan, Yuexin Wu, Zi-Hao Qiu, Xianzhi Du, Lijun Zhang 0005, Denny Zhou, Tianbao Yang |
ICML | 4 |
| 2022 | Back Razor: Memory-Efficient Transfer Learning by Self-Sparsified BackpropagationabstractTransfer learning from the model trained on large datasets to customized downstream tasks has been widely used as the pre-trained model can greatly boost the generalizability. However, the increasing sizes of pre-trained models also lead to a prohibitively large memory footprints for downstream transferring, making them unaffordable for personal devices. Previous work recognizes the bottleneck of the footprint to be the activation, and hence proposes various solutions such as injecting specific lite modules. In this work, we present a novel memory-efficient transfer framework called Back Razor, that can be plug-and-play applied to any pre-trained network without changing its architecture. The key idea of Back Razor is asymmetric sparsifying: pruning the activation stored for back-propagation, while keeping the forward activation dense. It is based on the observation that the stored activation, that dominates the memory footprint, is only needed for backpropagation. Such asymmetric pruning avoids affecting the precision of forward computation, thus making more aggressive pruning possible. Furthermore, we conduct the theoretical analysis for the convergence rate of Back Razor, showing that under mild conditions, our method retains the similar convergence rate as vanilla SGD. Extensive transfer learning experiments on both Convolutional Neural Networks and Vision Transformers with classification, dense prediction, and language modeling tasks show that Back Razor could yield up to 97% sparsity, saving 9.2x memory usage, without losing accuracy. The code is available at: https://github.com/VITA-Group/BackRazor_Neurips22. Ziyu Jiang, Xuxi Chen, Xueqin Huang, Xianzhi Du, Denny Zhou, Zhangyang Wang |
NeurIPS | 4 |
| 2021 | Revisiting ResNets: Improved Training and Scaling StrategiesabstractNovel computer vision architectures monopolize the spotlight, but the impact of the model architecture is often conflated with simultaneous changes to training methodology and scaling strategies.Our work revisits the canonical ResNet and studies these three aspects in an effort to disentangle them. Perhaps surprisingly, we find that training and scaling strategies may matter more than architectural changes, and further, that the resulting ResNets match recent state-of-the-art models. We show that the best performing scaling strategy depends on the training regime and offer two new scaling strategies: (1) scale model depth in regimes where overfitting can occur (width scaling is preferable otherwise); (2) increase image resolution more slowly than previously recommended.Using improved training and scaling strategies, we design a family of ResNet architectures, ResNet-RS, which are 1.7x - 2.7x faster than EfficientNets on TPUs, while achieving similar accuracies on ImageNet. In a large-scale semi-supervised learning setup, ResNet-RS achieves 86.2% top-1 ImageNet accuracy, while being 4.7x faster than EfficientNet-NoisyStudent. The training techniques improve transfer performance on a suite of downstream tasks (rivaling state-of-the-art self-supervised algorithms) and extend to video classification on Kinetics-400. We recommend practitioners use these simple revised ResNets as baselines for future research. Irwan Bello, William Fedus, Xianzhi Du, Ekin Dogus Cubuk, Aravind Srinivas, Tsung-Yi Lin, Jonathon Shlens, Barret Zoph |
NeurIPS | 3 |
| 2020 | SpineNet: Learning Scale-Permuted Backbone for Recognition and LocalizationabstractConvolutional neural networks typically encode an input image into a series of intermediate features with decreasing resolutions. While this structure is suited to classification tasks, it does not perform well for tasks requiring simultaneous recognition and localization (e.g., object detection). The encoder-decoder architectures are proposed to resolve this by applying a decoder network onto a backbone model designed for classification tasks. In this paper, we argue encoder-decoder architecture is ineffective in generating strong multi-scale features because of the scale-decreased backbone. We propose SpineNet, a backbone with scale-permuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search. Using similar building blocks, SpineNet models outperform ResNet-FPN models by 3%+ AP at various scales while using 10-20% fewer FLOPs. In particular, SpineNet-190 achieves 52.1% AP on COCO, attaining the new state-of-the-art performance for single model object detection without test-time augmentation. SpineNet can transfer to classification tasks, achieving 5% top-1 accuracy improvement on a challenging iNaturalist fine-grained dataset. Code is at: https://github.com/tensorflow/tpu/tree/master/models/official/detection. Xianzhi Du, Tsung-Yi Lin, Pengchong Jin, Golnaz Ghiasi, Mingxing Tan, Yin Cui, Quoc V. Le, Xiaodan Song |
CVPR | 1 |
| 2020 | Efficient Scale-Permuted Backbone with Learned Resource Distribution
Xianzhi Du, Tsung-Yi Lin, Pengchong Jin, Yin Cui, Mingxing Tan, Quoc V. Le, Xiaodan Song |
ECCV (23) | 1 |
| 2019 | Boundary-sensitive Network for Portrait SegmentationabstractPortrait segmentation has gained more and more attractions in recent years due to the popularity of selfie images. Compared to general semantic segmentation problems, portrait segmentation focuses on facial areas with higher requirements especially over the boundaries. To improve the performance of portrait segmentation, we propose a boundary-sensitive deep neural network (BSN) for better accuracy among the portrait boundaries. BSN introduces three novel techniques. First, an individual boundary-sensitive mask is proposed by dilating the contour line and assigning the boundary pixels with multi-class labels. Second, a global boundary-sensitive mask is employed as a position sensitive prior to further constrain the overall shape of the segmentation map. Third, we train a boundary-sensitive attribute classifier jointly with the segmentation network to reinforce the network with semantic boundary shape information. We have evaluated BSN on the state-of-the-art public portrait segmentation datasets, i.e., the PFCN dataset, as well as the portrait images collected from other three popular image segmentation datasets: COCO, COCO-Stuff, and PASCAL VOC. Our method achieves the superior quantitative and qualitative performance over state-of-the-arts on the evaluated datasets, especially obtains better visualization effect on the portrait boundary region. Xianzhi Du, Xiaolong Wang 0006, Dawei Li 0006, Serafettin Tasci, Cameron Upright, Stephen Walsh, Larry Davis 0001 |
FG | 1 |
| 2019 | Multi-Task Learning of Depth from Tele and Wide Stereo Image PairsabstractIn this paper, we introduce the problem of estimating the real world depth of elements in a scene captured by two cameras with different field of views, where the first field of view (FOV) is a Wide FOV obtained by a lens with 1 × the optical zoom, and the second FOV is contained in the first FOV and corresponds to a tele zoom lens with 2 × the optical zoom. Traditional stereo matching techniques can estimate the stereo disparity, and hence the depth, in the overlapping FOV between both cameras only, which corresponds to the Tele FOV. We refer to the problem of estimating the disparity, or inverse depth, for the union of FOVs as `Tele-Wide disparity estimation'. We propose different deep learning solutions to establish baseline performances. We trained a single-image inverse-depth estimation (SIDE) network to estimate the inverse depth from the image corresponding to the Wide FOV only. We also trained a stereo image disparity estimation network to estimate the disparity for the overlapping Tele FOV only, and another tele-wide stereo matching network (TW-SMNet) for estimating the disparity for the union Wide FOV. We further propose an end-to-end multi-task tele-wide stereo matching deep neural network (MT-TW-SMNet) which attempts to do stereo matching in the overlapped Tele FOV and SIDE in the union Wide FOV. Experimental results on KITTI and the SceneFlow datasets establish baseline performances for the tele-wide stereo matching and demonstrate that multitask tele-wide stereo matching provides a reasonable solution to the Tele-Wide depth estimation problem. Mostafa El-Khamy, Xianzhi Du |
ICIP | 2 |
| 2017 | Cyber-physical system enabled nearby traffic flow modelling for autonomous vehiclesabstractWe propose a nearby traffic flow modelling solution based on built-in Cyber-Physical System (CPS) sensors of autonomous vehicles. Our goal is to enhance the offline route planning and driving decision adjustment based on the first-hand traffic information, especially during poor Internet connection moments. Specifically, our model helps to select the optimal speed on a road, the optimal distance for timing to brake, and the safe distance from other vehicles to keep. Moreover, our model can also assist neighboring autonomous vehicles by communicating required information through Ad-Hoc network communications or through a centralized cloud. In detail, we first focus on the unique characteristic of traffic flow (such as traffic rule, avoid collision behaviours), and then build a comprehensive model to handle multiple scenarios. Technically, our model uses density functions of velocities, the differential equation of traffic flows, and the traffic viscosity with information collected from the traffic flow, the distances between vehicles, the amount and density of vehicle, the instant velocity, the speed limit, and the momentum to analysis the the driving scene. We evaluate our model with real traffic data collected by in-vehicle CPS sensors to the proposed nearby traffic flow model. Results show that our work can accurately conduct offline estimation on nearby traffic signal influence, and reveal the correlations among velocity, density and (spatial and temporal) location to adjust route during runtime. Zhengyu Yang 0001, Siyu Huang, Xianzhi Du, Janki Bhimani, Ningfang Mi |
IPCCC | 4 |
| 2017 | Fused DNN: A Deep Neural Network Fusion Approach to Fast and Robust Pedestrian DetectionabstractWe propose a deep neural network fusion architecture for fast and robust pedestrian detection. The proposed network fusion architecture allows for parallel processing of multiple networks for speed. A single shot deep convolutional network is trained as a object detector to generate all possible pedestrian candidates of different sizes and occlusions. This network outputs a large variety of pedestrian candidates to cover the majority of ground-truth pedestrians while also introducing a large number of false positives. Next, multiple deep neural networks are used in parallel for further refinement of these pedestrian candidates. We introduce a soft-rejection based network fusion method to fuse the soft metrics from all networks together to generate the final confidence scores. Our method performs better than existing state-of-the-arts, especially when detecting small-size and occluded pedestrians. Furthermore, we propose a method for integrating pixel-wise semantic segmentation network into the network fusion architecture as a reinforcement to the pedestrian detector. The approach outperforms state-of-the-art methods on most protocols on Caltech Pedestrian dataset, with significant boosts on several protocols. It is also faster than all other methods. Xianzhi Du, Mostafa El-Khamy, Larry Davis 0001 |
WACV | 1 |
| 2015 | A graphical model approach for matching partial signaturesabstractIn this paper, we present a novel partial signature matching method using graphical models. Shape context features are extracted from the contour of signatures to capture local variations, and K-means clustering is used to build a visual vocabulary from a set of reference signatures. To describe the signatures, supervised latent Dirichlet allocation is used to learn the latent distributions of the salient regions over the visual vocabulary and hierarchical Dirichlet processes are implemented to infer the number of salient regions needed. Our work is evaluated on three datasets derived from the DS-I Tobacco signature dataset with clean signatures and the DS-II UMD dataset with signatures with different degradations. The results show the effectiveness of the approach for both the partial and full signature matching. Xianzhi Du, David S. Doermann, Wael Abd-Almageed |
CVPR | 1 |
| 2014 | Signature Matching Using Supervised Topic ModelsabstractIn this paper, we present a novel signature matching method based on supervised topic models. Shape Context features are extracted from signature shape contours which capture the local variations in signature properties. We then use the concept of topic models to learn the shape context features which correspond to individual authors. The approach consists of three primary steps. First, K-means is used to cluster shape context features to form term frequency histograms which correspond to a vocabulary for the set of signatures in the gallery. Second, a supervised topic model is used to construct an observation/author correspondence. Finally, the correspondence is used to classify query signatures and return the corresponding author. Two datasets are used to test our algorithm: DS-I Tobacco signature dataset with clean signatures and DS-II UMD dataset with noisy signatures. We demonstrate considerable improvement over state of the art methods. Xianzhi Du, David S. Doermann, Wael Abd-Almageed |
ICPR | 1 |
| 2013 | Large-Scale Signature Matching Using Multi-stage HashingabstractIn this paper, we propose a fast large-scale signature matching method based on locality sensitive hashing (LSH). Shape Context features are used to describe the structure of signatures. Two stages of hashing are performed to find the nearest neighbours for query signatures. In the first stage, we use M randomly generated hyper planes to separate shape context feature points into different bins, and compute a term-frequency histogram to represent the feature point distribution as a feature vector. In the second stage we again use LSH to categorize the high-level features into different classes. The experiments are carried out on two datasets - DS-I, a small dataset contains 189 signatures, and DS-II, a large dataset created by our group which contains 26,000 signatures. We show that our algorithm can achieve a high accuracy even when few signatures are collected from one same person and perform fast matching when dealing with a large dataset. Xianzhi Du, Wael Abd-Almageed, David S. Doermann |
ICDAR | 1 |