EDBT 2026 Demo / reviewers in the wild / expert
Guan Pang
dblp:99/8792
· DBLP profile ↗
18ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TLDR: Token-Level Detective Reward Model for Large Vision Language ModelsabstractAlthough reward models have been successful in improving multimodal large language models, the reward models themselves remain brutal and contain minimal information. Notably, existing reward models only mimic human annotations by assigning only one feedback to any text, no matter how long the text is. In the realm of multimodal language models, where models are required to process both images and texts, a naive reward model may learn implicit biases toward texts and become less grounded in images. In this paper, we propose a **T**oken-**L**evel **D**etective **R**eward Model (**TLDR**) to provide fine-grained annotations to each text token. We first introduce a perturbation-based method to generate synthetic hard negatives and their token-level labels to train TLDR models. Then we show the rich usefulness of TLDR models both in assisting off-the-shelf models to self-correct their generations, and in serving as a hallucination evaluation tool. We show that TLDR automatically trains a token-level likelihood optimization, and can improve the base model's performance significantly. Finally, we show that TLDR models can significantly speed up human annotation by 3 times to acquire a broader range of high-quality vision language data. Deqing Fu, Tong Xiao 0003, Wang Zhu 0001, Pengchuan Zhang, Guan Pang, Robin Jia, Lawrence Chen 0002 |
ICLR | 6 |
| 2024 | Layout-Agnostic Scene Text Image Synthesis with Diffusion ModelsabstractWhile diffusion models have significantly advanced the quality of image generation, their capability to accurately and coherently render text within these images remains a substantial challenge. Conventional diffusion-based methods for scene text generation are typically limited by their reliance on an intermediate layout output. This dependency often results in a constrained diversity of text styles and fonts, an inherent limitation stemming from the deterministic nature of the layout generation phase. To address these challenges, this paper introduces Scene TextGen, a novel diffusion-based model specifically designed to circumvent the need for a predefined layout stage. By doing so, Scene-TextGen facilitates a more natural and varied representation of text. The novelty of SceneTextGen lies in its integration of three key components: a character-level encoder for capturing detailed typographic properties, coupled with a character-level instance segmentation model and a word-level spotting model to address the issues of unwanted text generation and minor character inaccuracies. We validate the performance of our method by demonstrating improved character recognition rates on generated images across different public visual text datasets in comparison to both standard diffusion based methods and text specific methods. Qilong Zhangli, Jindong Jiang, Di Liu 0003, Licheng Yu, Xiaoliang Dai, Ankit Ramchandani, Guan Pang, Dimitris N. Metaxas, Praveen Krishnan |
CVPR | 7 |
| 2024 | LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
Bolin Lai, Xiaoliang Dai, Lawrence Chen 0002, Guan Pang, James M. Rehg, Miao Liu 0007 |
ECCV (9) | 4 |
| 2023 | TextStyleBrush: Transfer of Text Aesthetics From a Single ExampleabstractWe present a novel approach for disentangling the content of a text image from all aspects of its appearance. The appearance representation we derive can then be applied to new content, for one-shot transfer of the source style to new content. We learn this disentanglement in a self-supervised manner. Our method processes entire word boxes, without requiring segmentation of text from background, per-character processing, or making assumptions on string lengths. We show results in different text domains which were previously handled by specialized methods, e.g., scene text, handwritten text. To these ends, we make a number of technical contributions: (1) We disentangle the style and content of a textual image into a non-parametric, fixed-dimensional vector. (2) We propose a novel approach inspired by StyleGAN but conditioned over the example style at different resolution and content. (3) We present novel self-supervised training criteria which preserve both source style and target content using a pre-trained font classifier and text recognizer. Finally, (4) we also introduce Imgur5K, a new challenging dataset for handwritten word images. We offer numerous qualitative photo-realistic results of our method. We further show that our method surpasses previous work in quantitative tests on scene text and handwriting datasets, as well as in a user study. Praveen Krishnan, Rama Kovvuri, Guan Pang, Boris Vassilev, Tal Hassner |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer
Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin 0001, Guan Pang, David Jacobs 0001, Jia-Bin Huang 0001, Devi Parikh |
ECCV (17) | 5 |
| 2022 | MUGEN: A Playground for Video-Audio-Text Multimodal Understanding and GENeration
Thomas Hayes, Songyang Zhang 0004, Xi Yin 0001, Guan Pang, Sasha Sheng, Harry Yang, Songwei Ge, Qiyuan Hu, Devi Parikh |
ECCV (8) | 4 |
| 2021 | A Multiplexed Network for End-to-End, Multilingual OCRabstractRecent advances in OCR have shown that an end-to-end (E2E) training pipeline that includes both detection and recognition leads to the best results. However, many existing methods focus primarily on Latin-alphabet languages, often even only case-insensitive English characters. In this paper, we propose an E2E approach, Multiplexed Multilingual Mask TextSpotter, that performs script identification at the word level and handles different scripts with different recognition heads, all while maintaining a unified loss that simultaneously optimizes script identification and multiple recognition heads. Experiments show that our method outperforms the single-head model with similar number of parameters in end-to-end recognition tasks, and achieves state-of-the-art results on MLT17 and MLT19 joint text detection and script identification benchmarks. We believe that our work is a step towards the end-to-end trainable and scalable multilingual multi-purpose OCR system. Our code and model will be released. Jing Huang 0020, Guan Pang, Rama Kovvuri, Mandy Toh, Kevin J. Liang, Praveen Krishnan, Xi Yin 0001, Tal Hassner |
CVPR | 2 |
| 2021 | img2pose: Face Alignment and Detection via 6DoF, Face Pose EstimationabstractWe propose real-time, six degrees of freedom (6DoF), 3D face pose estimation without face detection or landmark localization. We observe that estimating the 6DoF rigid transformation of a face is a simpler problem than facial landmark detection, often used for 3D face alignment. In addition, 6DoF offers more information than face bounding box labels. We leverage these observations to make multiple contributions: (a) We describe an easily trained, efficient, Faster R-CNN–based model which regresses 6DoF pose for all faces in the photo, without preliminary face detection. (b) We explain how pose is converted and kept consistent between the input photo and arbitrary crops created while training and evaluating our model. (c) Finally, we show how face poses can replace detection bounding box training labels. Tests on AFLW2000-3D and BIWI show that our method runs at real-time and outperforms state of the art (SotA) face pose estimators. Remarkably, our method also surpasses SotA models of comparable complexity on the WIDER FACE detection benchmark, despite not been optimized on bounding box labels. Vitor Albiero, Xi Yin 0001, Guan Pang, Tal Hassner |
CVPR | 4 |
| 2021 | TextOCR: Towards Large-Scale End-to-End Reasoning for Arbitrary-Shaped Scene TextabstractA crucial component for the scene text based reasoning required for TextVQA and TextCaps datasets involve detecting and recognizing text present in the images using an optical character recognition (OCR) system. The current systems are crippled by the unavailability of ground truth text annotations for these datasets as well as lack of scene text detection and recognition datasets on real images disallowing the progress in the field of OCR and evaluation of scene text based reasoning in isolation from OCR systems. In this work, we propose TextOCR, an arbitrary-shaped scene text detection and recognition with 900k annotated words collected on real images from TextVQA dataset. We show that current state-of-the-art text-recognition (OCR) models fail to perform well on TextOCR and that training on TextOCR helps achieve state-of-the-art performance on multiple other OCR datasets as well. We use a TextOCR trained OCR model to create PixelM4C model which can do scene text based reasoning on an image in an end-to-end fashion, allowing us to revisit several design choices to achieve new state-of-the-art performance on TextVQA dataset. Amanpreet Singh, Guan Pang, Mandy Toh, Jing Huang 0020, Wojciech Galuba, Tal Hassner |
CVPR | 2 |
| 2020 | Mask TextSpotter v3: Segmentation Proposal Network for Robust Scene Text Spotting
Minghui Liao, Guan Pang, Jing Huang 0020, Tal Hassner, Xiang Bai |
ECCV (11) | 2 |
| 2019 | Improved Road Connectivity by Joint Learning of Orientation and SegmentationabstractRoad network extraction from satellite images often produce fragmented road segments leading to road maps unfit for real applications. Pixel-wise classification fails to predict topologically correct and connected road masks due to the absence of connectivity supervision and difficulty in enforcing topological constraints. In this paper, we propose a connectivity task called Orientation Learning, motivated by the human behavior of annotating roads by tracing it at a specific orientation. We also develop a stacked multi-branch convolutional module to effectively utilize the mutual information between orientation learning and segmentation tasks. These contributions ensure that the model predicts topologically correct and connected road masks. We also propose Connectivity Refinement approach to further enhance the estimated road networks. The refinement model is pre-trained to connect and refine the corrupted ground-truth masks and later fine-tuned to enhance the predicted road masks. We demonstrate the advantages of our approach on two diverse road extraction datasets SpaceNet and DeepGlobe. Our approach improves over the state-of-the-art techniques by 9% and 7.5% in road topology metric on SpaceNet and DeepGlobe, respectively. Anil Batra, Suriya Singh, Guan Pang, Saikat Basu, C. V. Jawahar, Manohar Paluri |
CVPR | 3 |
| 2019 | A Computer Vision Perspective on Analyzing and Synthesizing Geospatial DataabstractThe AI era sustains its foundations from the availability of large datasets. Especially geospatial datasets are very interesting from a computer vision perspective, as they enable us to understand the world we live in. Although many application domains arise from analyzing such big data, analysis itself is not enough for impacting lives. As its counter part, synthesis approaches are recently being developed for mimicking real-world data for completing and creating new worlds. In this paper, we will explore not only example analysis methods developed using large public datasets, but also some generative models to propose realistic and impactful solutions for going beyond observations. Ilke Demir, Guan Pang, Jing Huang 0020 |
IGARSS | 2 |
| 2018 | Self-Supervised Feature Learning for Semantic Segmentation of Overhead Imagery
Suriya Singh, Anil Batra, Guan Pang, Lorenzo Torresani, Saikat Basu, Manohar Paluri, C. V. Jawahar |
BMVC | 3 |
| 2016 | 3D point cloud object detection with multi-view convolutional neural networkabstractEfficient detection of three dimensional (3D) objects in point clouds is a challenging problem. Performing 3D descriptor matching or 3D scanning-window search with detector are both time-consuming due to the 3-dimensional complexity. One solution is to project 3D point cloud into 2D images and thus transform the 3D detection problem into 2D space, but projection at multiple viewpoints and rotations produce a large amount of 2D detection tasks, which limit the performance and complexity of the 2D detection algorithm choice. We propose to use convolutional neural network (CNN) for the 2D detection task, because it can handle all viewpoints and rotations for the same class of object together, as well as predicting multiple classes of objects with the same network, without the need for individual detector for each object class. We further improve the detection efficiency by concatenating two extra levels of early rejection networks with binary outputs before the multi-class detection network. Experiments show that our method has competitive overall performance with at least one-order of magnitude speed-up comparing with latest 3D point cloud detection methods. Guan Pang, Ulrich Neumann |
ICPR | 1 |
| 2015 | Fast and Robust Multi-view 3D Object Recognition in Point CloudsabstractRecognition of three dimensional (3D) objects in point clouds is a challenging problem. Existing methods often require prior segmentation or 3D descriptor training and matching, both time consuming and complex processes, especially for large-scale industrial or urban street data. We describe a new recognition approach that projects a 3D point cloud into several 2D depth images from multiple viewpoints, transforming the 3D recognition problem into a series of 2D detection problems. This method reduces complexity, stabilizes performance, and significantly speeds up the recognition process, without any requirement for object segmentation or detector training. Experiments validate the superiority of our method over several state-of-the-art methods on examples from industrial and street data scans. Guan Pang, Ulrich Neumann |
3DV | 1 |
| 2013 | Training-Based Object Recognition in Cluttered 3D Point CloudsabstractRecognition of three dimensional (3D) objects is a challenging problem, especially in cluttered or occluded scenes. Many existing methods focus on a specific type of object or scene, or require prior segmentation. We describe a robust and efficient general purpose 3D object recognition method that combines machine learning procedures with 3D local features, without a requirement for a priori object segmentation. Experiments validate our method on various object types from engineering and street data scans. Guan Pang, Ulrich Neumann |
3DV | 1 |
| 2013 | Estimation of camera pose with respect to terrestrial LiDAR dataabstractIn this paper, we present an algorithm that is to estimate the position of a hand-held camera with respect to terrestrial LiDAR data. Our input is a set of 3D range scans with intensities and one or a set of 2D uncalibrated camera images of the scene. The algorithm that automatically registers range scans and 2D images is composed of following steps. In the first step, we project the terrestrial LiDAR onto 2D images according to several preselected viewpoints. Intensity-based features such as SIFT are extracted from these projected images and these features are projected back onto the LiDAR data to obtain their 3D positions. In the second step, we estimate the initial pose of given 2D images from feature correspondences. In the third step, we refine the coarse camera pose obtained from the previous step through iterative matchings and optimization process. We presents results from experiments in several different urban settings. Wei Guan 0005, Suya You, Guan Pang |
WACV | 3 |
| 2013 | The Gixel array descriptor (GAD) for multimodal image matchingabstractFeature description and matching is a fundamental problem for many computer vision applications. However, most existing descriptors only work well on images of a single modality with similar texture. This paper presents a novel basic descriptor unit called a Gixel, which uses an additive scoring method to sample surrounding edge information. Several Gixels in a circular array create a powerful descriptor called the Gixel Array Descriptor (GAD), excelling in multi-modal image matching, especially when one of the images is edge-dominant with little texture. Experiments demonstrate the superiority of GAD on multi-modal matching, while maintaining a performance comparable to several state-of-the-art descriptors on single modality matching. Guan Pang, Ulrich Neumann |
WACV | 1 |