EDBT 2026 Demo / reviewers in the wild / expert
Mun Wai Lee
dblp:77/2305
· DBLP profile ↗
14ranked-venue papers
7as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Deep learning architectures and training · 40% Segmentation and scene understanding · 16% Face, body and person analysis · 14% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Database theory · 50% Knowledge graphs · 50% |
Topics — the 24 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention |
0.7 | 1 | 2023 | PaCa-ViT: Learning Patch-to-Cluster Attention in Vision Transformers · CVPR 2023 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.7 | 1 | 2023 | PaCa-ViT: Learning Patch-to-Cluster Attention in Vision Transformers · CVPR 2023 |
Computer vision › Face, body and person analysis
human pose estimation |
0.3 | 5 | 2009 | Human Pose Tracking in Monocular Sequence Using Multilevel Structured Models · IEEE Trans. Pattern Anal. Mach. Intell. 2009 A Model-Based Approach for Estimating Human 3D Poses in Static Images · IEEE Trans. Pattern Anal. Mach. Intell. 2006 Human Pose Tracking Using Multi-level Structured Models · ECCV (3) 2006 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2023 | PaCa-ViT: Learning Patch-to-Cluster Attention in Vision Transformers · CVPR 2023 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.2 | 1 | 2023 | PaCa-ViT: Learning Patch-to-Cluster Attention in Vision Transformers · CVPR 2023 |
Computer vision › Segmentation and scene understanding › semantic segmentation
transformer-based segmentation |
0.2 | 1 | 2023 | PaCa-ViT: Learning Patch-to-Cluster Attention in Vision Transformers · CVPR 2023 |
Computer vision › Vision and language
image captioning |
0.1 | 1 | 2010 | I2T: Image Parsing to Text Description · Proc. IEEE 2010 |
Computer vision › Segmentation and scene understanding
scene parsing |
0.1 | 1 | 2010 | I2T: Image Parsing to Text Description · Proc. IEEE 2010 |
Computer vision › Vision and language › image captioning › description generation
sentence generation from images |
0.1 | 1 | 2010 | I2T: Image Parsing to Text Description · Proc. IEEE 2010 |
Computer vision › Video understanding and tracking › multi-object tracking
multi-person tracking |
0.1 | 1 | 2009 | Human Pose Tracking in Monocular Sequence Using Multilevel Structured Models · IEEE Trans. Pattern Anal. Mach. Intell. 2009 |
Computer vision › 3D vision › pose estimation
pose tracking |
0.1 | 1 | 2009 | Human Pose Tracking in Monocular Sequence Using Multilevel Structured Models · IEEE Trans. Pattern Anal. Mach. Intell. 2009 |
Database theory
ontology-mediated queries |
0.1 | 1 | 2009 | Semantic video search using natural language queries · ACM Multimedia 2009 |
Knowledge graphs › knowledge graph querying
SPARQL query generation |
0.1 | 1 | 2009 | Semantic video search using natural language queries · ACM Multimedia 2009 |
Multimedia analysis and retrieval › video retrieval
semantic video retrieval |
0.1 | 1 | 2009 | Semantic video search using natural language queries · ACM Multimedia 2009 |
Multimedia analysis and retrieval
video retrieval |
0.1 | 1 | 2009 | Semantic video search using natural language queries · ACM Multimedia 2009 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.1 | 1 | 2008 | A rank constrained continuous formulation of multi-frame multi-target tracking problem · CVPR 2008 |
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation |
0.1 | 1 | 2006 | A Model-Based Approach for Estimating Human 3D Poses in Static Images · IEEE Trans. Pattern Anal. Mach. Intell. 2006 |
Computer vision › Face, body and person analysis › human pose estimation
human pose tracking |
0.1 | 1 | 2006 | Human Pose Tracking Using Multi-level Structured Models · ECCV (3) 2006 |
Machine learning › Probabilistic and Bayesian machine learning
structured models |
0.1 | 1 | 2006 | Human Pose Tracking Using Multi-level Structured Models · ECCV (3) 2006 |
Robotics › Autonomous driving
driving scene understanding |
0.0 | 1 | 2010 | I2T: Image Parsing to Text Description · Proc. IEEE 2010 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › representation language › knowledge representation formalisms
visual knowledge representation |
0.0 | 1 | 2010 | I2T: Image Parsing to Text Description · Proc. IEEE 2010 |
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction |
0.0 | 1 | 2009 | Human Pose Tracking in Monocular Sequence Using Multilevel Structured Models · IEEE Trans. Pattern Anal. Mach. Intell. 2009 |
Computer vision › 3D vision
3d reconstruction |
0.0 | 1 | 2006 | A Model-Based Approach for Estimating Human 3D Poses in Static Images · IEEE Trans. Pattern Anal. Mach. Intell. 2006 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.0 | 1 | 2004 | Proposal Maps Driven MCMC for Estimating Human Body Pose in Static Images · CVPR (2) 2004 |
Methods — techniques the papers use, named apart from their topics
clustering · 0.7attention mechanism · 0.7markov chain monte carlo · 0.2semantic stochastic context-free grammar · 0.2earley-stolcke parsing · 0.2RDF/OWL ontology · 0.2nonlinear optimization · 0.2proposal map · 0.1web ontology language · 0.1stochastic image grammar · 0.1and-or graph · 0.1data-driven MCMC · 0.1belief propagation · 0.1scanning window · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ArcGeo: Localizing Limited Field-of-View Images using Cross-view MatchingabstractCross-view matching techniques for image geo-localization attempt to match features in ground-level query images against a collection of satellite images to determine their positions of origin. We present ArcGeo, a novel cross-view image matching approach which introduces a batch-all angular margin loss and several train-time strategies including large-scale pretraining and FoV-based data augmentation. This allows our model to perform well even in challenging cases with limited field-of-view (FoV). Further, we evaluate multiple model architectures, data augmentation approaches and optimization strategies to train a deep cross-view matching network, specifically optimized for limited FoV cases. In low FoV experiments (FoV = 90°) our method improves top-1 image recall rate on the CVUSA dataset from 30.12% to 43.08%. We also demonstrate improved performance over the state-of-the-art techniques for panoramic cross-view retrieval, improving top-1 recall from 95.43% to 96.06% on the CVUSA dataset and from 64.52% to 79.88% on the CVACT test dataset. Lastly, we evaluate the role of large-scale pretraining for improved robustness. With appropriate pretraining on external data, our model improves top-1 recall dramatically to 66.83% for the FoV = 90° test case on CVUSA, an increase of over twice what is reported by existing approaches. Maxim Shugaev, Ilya Semenov, Kyle Ashley, Michael Klaczynski, Naresh P. Cuntoor, Mun Wai Lee, Nathan Jacobs |
WACV | 6 |
| 2023 | PaCa-ViT: Learning Patch-to-Cluster Attention in Vision TransformersabstractVision Transformers (ViTs) are built on the assumption of treating image patches as “visual tokens” and learn patch-to-patch attention. The patch embedding based tokenizer has a semantic gap with respect to its counterpart, the textual tokenizer. The patch-to-patch attention suffers from the quadratic complexity issue, and also makes it non-trivial to explain learned ViTs. To address these issues in ViT, this paper proposes to learn Patch-to-Cluster attention (PaCa) in ViT. Queries in our PaCa-ViT starts with patches, while keys and values are directly based on clustering (with a predefined small number of clusters). The clusters are learned end-to-end, leading to better tokenizers and inducing joint clustering-for-attention and attention-for-clustering for better and interpretable models. The quadratic complexity is relaxed to linear complexity. The proposed PaCa module is used in designing efficient and interpretable ViT backbones and semantic segmentation head networks. In experiments, the proposed methods are tested on ImageNet-1k image classification, MS-COCO object detection and instance segmentation and MIT-ADE20k semantic segmentation. Compared with the prior art, it obtains better performance in all the three benchmarks than the SWin [32] and the PVTs [47], [48] by significant margins in ImageNet-1k and MIT-ADE20k. It is also significantly more efficient than PVT models in MS-COCO and MIT-ADE20k due to the linear complexity. The learned clusters are semantically meaningful. Code and model checkpoints are available at https:/github.com/iVMCL/PaCaViT. Ryan Grainger, Thomas Paniagua, Naresh P. Cuntoor, Mun Wai Lee, Tianfu Wu 0001 |
CVPR | 5 |
| 2023 | HBRC-500: A Long Range Recognition Benchmark Dataset using Face and Whole-body ImageryabstractWhile biometric-face and whole-body-recognition technology have recently advanced and matured, there are increasing interests in enhanced long-range recognition capabilities. However, long-range recognition requires the use of large, specialized datasets to support research and development for next generation systems. Moreover, existing datasets are further limited by the types of modalities (face or whole-body), number of subjects, maximum standoff distance, clothing variability, or availability restrictions. For long-range recognition, low-quality probe (query) images, which are often acquired from extended standoff distances or aerial platforms, are matched against higher quality gallery images and frequently results in poor identification performance. To address the growing needs for relevant data sources, a large-scale biometric dataset was collected and curated for long-range biometric recognition. This dataset is comprised of more than 1.2 million outdoor and 250,000 indoor (face and whole-body) images from more than 250 subjects that were acquired using various high-end cameras, including Canon and Nikon DSLR cameras, surveillance cameras, specialized long-range face cameras, and UAV platforms. The primary goal of this dataset is to support the development of algorithms for face and whole-body recognition at extended standoff distances. The availability of such a dataset is crucial in advancing technology for recognition under challenging conditions such as atmospheric turbulence. Cedric Nimpa Fondje, Kshitij Nikhal, John Brennan Peace, Ryan Karl, Mun Wai Lee, Phillip Berkowitz, Katrina Gramzinski, Bridget Kennedy, Nkirukaegbunam Uzuegbunam, Victoria Ou, Tyler Barret, Oliver Arend, Wei Ming, Svetlana Semenova, Benjamin S. Riggan |
IJCB | 5 |
| 2010 | I2T: Image Parsing to Text DescriptionabstractIn this paper, we present an image parsing to text description (I2T) framework that generates text descriptions of image and video content based on image understanding. The proposed I2T framework follows three steps: 1) input images (or video frames) are decomposed into their constituent visual patterns by an image parsing engine, in a spirit similar to parsing sentences in natural language; 2) the image parsing results are converted into semantic representation in the form of Web ontology language (OWL), which enables seamless integration with general knowledge bases; and 3) a text generation engine converts the results from previous steps into semantically meaningful, human readable, and query-able text reports. The centerpiece of the I2T framework is an and-or graph (AoG) visual knowledge representation, which provides a graphical representation serving as prior knowledge for representing diverse visual patterns and provides top-down hypotheses during the image parsing. The AoG embodies vocabularies of visual elements including primitives, parts, objects, scenes as well as a stochastic image grammar that specifies syntactic relations (i.e., compositional) and semantic relations (e.g., categorical, spatial, temporal, and functional) between these visual elements. Therefore, the AoG is a unified model of both categorical and symbolic representations of visual knowledge. The proposed I2T framework has two objectives. First, we use semiautomatic method to parse images from the Internet in order to build an AoG for visual knowledge representation. Our goal is to make the parsing process more and more automatic using the learned AoG model. Second, we use automatic methods to parse image/video in specific domains and generate text reports that are useful for real-world applications. In the case studies at the end of this paper, we demonstrate two automatic I2T systems: a maritime and urban scene video surveillance system and a real-time automatic driving scene understanding system. Benjamin Z. Yao, Liang Lin 0004, Mun Wai Lee, Song-Chun Zhu |
Proc. IEEE | 4 |
| 2009 | Semantic video search using natural language queriesabstractRecent advances in computer vision and artificial intelligence algorithms have allowed automatic extraction of metadata from video. This metadata can be represented by using the RDF/OWL ontology which can encode scene objects and their relationships in an unambiguous and well-formed manner. The encoded data can be queried using SPARQL. However, SPARQL has a steep learning curve and cannot be directly utilized by a general user for video content search. In this paper, we propose a method to bridge this gap by automatically translating user provided natural language query into an ontology-based SPARQL query for semantic video search. The proposed method consists of three major steps. First, semantically labeled training corpus of natural language query sentences is used for learning the Semantic Stochastic Context Free Grammar (SSCFG). Second, given a user provided natural language query sentence, we use the Earley-Stolcke parsing algorithm to determine the maximum likelihood semantic parsing of the query sentence. This parsing infers the semantic meaning for each word in the query sentence from which the SPARQL query is constructed. Third, the SPARQL query is executed to retrieve relevant video segments from the RDF-OWL video content database. The method is evaluated by running natural language queries on surveillance videos from maritime and land-based domains, though the framework itself is general and extensible to search videos from other domains. Asaad Hakeem, Mun Wai Lee, Omar Javed, Niels Haering |
ACM Multimedia | 2 |
| 2009 | Human Pose Tracking in Monocular Sequence Using Multilevel Structured ModelsabstractTracking human body poses in monocular video has many important applications. The problem is challenging in realistic scenes due to background clutter, variation in human appearance and self-occlusion. The complexity of pose tracking is further increased when there are multiple people whose bodies may inter-occlude. We proposed a three-stage approach with multi-level state representation that enables a hierarchical estimation of 3D body poses. Our method addresses various issues including automatic initialization, data association, self and inter-occlusion. At the first stage, humans are tracked as foreground blobs and their positions and sizes are coarsely estimated. In the second stage, parts such as face, shoulders and limbs are detected using various cues and the results are combined by a grid-based belief propagation algorithm to infer 2D joint positions. The derived belief maps are used as proposal functions in the third stage to infer the 3D pose using data-driven Markov chain Monte Carlo. Experimental results on several realistic indoor video sequences show that the method is able to track multiple persons during complex movement including sitting and turning movements with self and inter-occlusion. Mun Wai Lee, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | A rank constrained continuous formulation of multi-frame multi-target tracking problemabstractThis paper presents a multi-frame data association algorithm for tracking multiple targets in video sequences. Multi-frame data association involves finding the most probable correspondences between target tracks and measurements (collected over multiple time instances) as well as handling the common tracking problems such as, track initiations and terminations, occlusions, and noisy detections. The problem is known to be NP-Hard for more than two frames. A rank constrained continuous formulation of the problem is presented that can be efficiently solved using nonlinear optimization methods. It is shown that the global and local extrema of the continuous problem respectively coincide with the maximum and the maximal solutions of the discrete counterpart. A scanning window based tracking algorithm is developed using the formulation that performs well under noisy conditions with frequent occlusions and multiple track initiations and terminations. The above claims are supported by experiments and quantitative evaluations using both synthetic and real data under different operating conditions. Khurram Shafique, Mun Wai Lee, Niels Haering |
CVPR | 2 |
| 2008 | Image transformation for object tracking in high-resolution videoabstractWe propose a new method for warping high-resolution images to efficiently track objects on the ground plane in real time. Recently, the emergence of high resolution video cameras (> 5 megapixels) has enabled surveillance over a much larger area using only a single camera. However, real-time processing of high resolution video for automatic detection and tracking of multiple targets is a challenge. When the surveillance camera covers greater depth of ground regions, due to perspective effect, the image size of a target varies significantly depending on the distance between the camera and the target. In this study, we propose a framework to transform high resolution images into warped images using a plane homography to make the target size uniform regardless of the position. The method not only reduces the number of pixels to be processed for speed-up, but also improves the tracking performance. We provide experimental results on object tracking in high-resolution maritime videos to demonstrate the validity of our method. Tae Eun Choe, Krishnan Ramnath, Mun Wai Lee, Niels Haering |
ICPR | 3 |
| 2006 | Human Pose Tracking Using Multi-level Structured Models
Mun Wai Lee, Ramakant Nevatia |
ECCV (3) | 1 |
| 2006 | A Model-Based Approach for Estimating Human 3D Poses in Static ImagesabstractEstimating human body poses in static images is important for many image understanding applications including semantic content extraction and image database query and retrieval. This problem is challenging due to the presence of clutter in the image, ambiguities in image observation, unknown human image boundary, and high-dimensional state space due to the complex articulated structure of the human body. Human pose estimation can be made more robust by integrating the detection of body components such as face and limbs, with the highly constrained structure of the articulated body. In this paper, a data-driven approach based on Markov chain Monte Carlo (DD-MCMC) is used, where component detection results generate state proposals for 3D pose estimation. To translate these observations into pose hypotheses, we introduce the use of "proposal maps," an efficient way of consolidating the evidence and generating 3D pose candidates during the MCMC search. Experimental results on a set of test images show that the method is able to estimate the human pose in static images of real scenes. Mun Wai Lee, Isaac Cohen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Proposal Maps Driven MCMC for Estimating Human Body Pose in Static Images
Mun Wai Lee, Isaac Cohen |
CVPR (2) | 1 |
| 2004 | Human Upper Body Pose Estimation in Static Images
Mun Wai Lee, Isaac Cohen |
ECCV (2) | 1 |
| 2003 | Pose-invariant face recognition using a 3D deformable model
Mun Wai Lee, Surendra Ranganath |
Pattern Recognit. | 1 |
| 2000 | Pose Invariant Face Recognition by Face Synthesis
Mun Wai Lee, Surendra Ranganath |
BMVC | 1 |