Huizhong Chen

dblp:05/10534 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
3since 2021 · last 2023
0000-0002-9160-6571ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Segmentation and scene understanding · 43% Image recognition and object detection · 21% Face, body and person analysis · 16%
Computer graphics and multimedia
3 papers
Multimedia analysis and retrieval · 42% Geometric modeling and processing · 32% Image and video processing · 26%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 66% Recommender systems · 34%

Topics — the 25 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
mixture of experts
0.712023
AdaMV-MoE: Adaptive Multi-Task Vision Mixture-of-Experts · ICCV 2023
Computer vision › Segmentation and scene understanding › image segmentation
multi-task segmentation
0.712023
DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation Model · NeurIPS 2023
Computer vision › Image recognition and object detection › visual recognition
multi-task visual recognition
0.712023
AdaMV-MoE: Adaptive Multi-Task Vision Mixture-of-Experts · ICCV 2023
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation
0.712023
DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation Model · NeurIPS 2023
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.712023
DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation Model · NeurIPS 2023
Computer vision › Segmentation and scene understanding
instance segmentation
0.612022
A Simple Single-Scale Vision Transformer for Object Detection and Instance Segmentation · ECCV (10) 2022
Computer vision › Image recognition and object detection
object detection
0.612022
A Simple Single-Scale Vision Transformer for Object Detection and Instance Segmentation · ECCV (10) 2022
Computer vision › Face, body and person analysis › facial attribute analysis
facial attribute recognition
0.222014
What's in a Name? First Names as Facial Attributes · CVPR 2013
The Hidden Sides of Names - Face Modeling with First Name Attributes · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Machine learning › Transfer learning and domain adaptation › multi-source learning
multi-dataset learning
0.212023
DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation Model · NeurIPS 2023
Computer vision › Face, body and person analysis
face modeling
0.212014
The Hidden Sides of Names - Face Modeling with First Name Attributes · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Computer vision › Face, body and person analysis › face recognition
face verification
0.212014
The Hidden Sides of Names - Face Modeling with First Name Attributes · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Computer vision › Face, body and person analysis › face recognition › face verification
unconstrained face verification
0.212014
The Hidden Sides of Names - Face Modeling with First Name Attributes · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Multimedia analysis and retrieval
video retrieval
0.212014
Multi-modal Language Models for Lecture Video Retrieval · ACM Multimedia 2014
Computer vision › Face, body and person analysis › face recognition › face annotation
face naming
0.212013
What's in a Name? First Names as Facial Attributes · CVPR 2013
Recommender systems › news recommendation
personalized news recommendation
0.212013
EigenNews: a personalized news video delivery platform · ACM Multimedia 2013
Image and video processing › image registration
deformable image registration
0.112012
Efficient Registration of Nonrigid 3-D Bodies · IEEE Trans. Image Process. 2012
Geometric modeling and processing › shape registration
nonrigid shape registration
0.112012
Efficient Registration of Nonrigid 3-D Bodies · IEEE Trans. Image Process. 2012
Geometric modeling and processing
shape registration
0.112012
Efficient Registration of Nonrigid 3-D Bodies · IEEE Trans. Image Process. 2012
Information retrieval › retrieval models
hybrid retrieval
0.112011
Combining image and text features: a hybrid approach to mobile book spine recognition · ACM Multimedia 2011
Information retrieval
retrieval models
0.112011
Combining image and text features: a hybrid approach to mobile book spine recognition · ACM Multimedia 2011
Content delivery and video streaming
video delivery
0.012013
EigenNews: a personalized news video delivery platform · ACM Multimedia 2013
Image and video processing
motion estimation
0.012012
Efficient Registration of Nonrigid 3-D Bodies · IEEE Trans. Image Process. 2012
Image and video processing › motion estimation
optical flow
0.012012
Efficient Registration of Nonrigid 3-D Bodies · IEEE Trans. Image Process. 2012
Information retrieval › image retrieval
mobile visual search
0.012011
Combining image and text features: a hybrid approach to mobile book spine recognition · ACM Multimedia 2011
Information retrieval
search engines
0.012011
Combining image and text features: a hybrid approach to mobile book spine recognition · ACM Multimedia 2011

Methods — techniques the papers use, named apart from their topics

vision transformer · 1.2weak supervision · 0.7text embedding · 0.7mixture of experts · 0.7mask proposal · 0.7single-scale feature pyramid · 0.6pairwise classifier · 0.4multimodal segmentation · 0.3text recognition · 0.2image feature matching · 0.2probabilistic ranking · 0.2probabilistic latent semantic analysis · 0.2latent variable model · 0.2adaboost · 0.2semantic attributes · 0.1phase-shift motion estimation · 0.1dual-tree complex wavelet transform · 0.1clothing parsing · 0.1
YearPublicationVenuePosition
2023 AdaMV-MoE: Adaptive Multi-Task Vision Mixture-of-Experts
abstract
Sparsely activated Mixture-of-Experts (MoE) is becoming a promising paradigm for multi-task learning (MTL). Instead of compressing multiple tasks’ knowledge into a single model, MoE separates the parameter space and only utilizes the relevant model pieces given task type and its input, which provides stabilized MTL training and ultra-efficient inference. However, current MoE approaches adopt a fixed network capacity (e.g., two experts in usual) for all tasks. It potentially results in the over-fitting of simple tasks or the under-fitting of challenging scenarios, especially when tasks are significantly distinctive in their complexity. In this paper, we propose an adaptive MoE framework for multi-task vision recognition, dubbed AdaMV-MoE. Based on the training dynamics, it automatically determines the number of activated experts for each task, avoiding the laborious manual tuning of optimal model size. To validate our proposal, we benchmark it on ImageNet classification and COCO object detection & instance segmentation which are notoriously difficult to learn in concert, due to their discrepancy. Extensive experiments across a variety of vision transformers demonstrate a superior performance of AdaMV-MoE, compared to MTL with a shared backbone and the recent state-of-the-art (SoTA) MTL MoE approach. Codes are available online: https://github.com/google-research/google-research/tree/master/moe_mtl.
Tianlong Chen 0001, Xuxi Chen, Xianzhi Du, Abdullah Rashwan, Huizhong Chen, Zhangyang Wang, Yeqing Li
ICCV6
2023 DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation Model
abstract
Observing the close relationship among panoptic, semantic and instance segmentation tasks, we propose to train a universal multi-dataset multi-task segmentation model: DaTaSeg. We use a shared representation (mask proposals with class predictions) for all tasks. To tackle task discrepancy, we adopt different merge operations and post-processing for different tasks. We also leverage weak-supervision, allowing our segmentation model to benefit from cheaper bounding box annotations. To share knowledge across datasets, we use text embeddings from the same semantic embedding space as classifiers and share all network parameters among datasets. We train DaTaSeg on ADE semantic, COCO panoptic, and Objects365 detection datasets. DaTaSeg improves performance on all datasets, especially small-scale datasets, achieving 54.0 mIoU on ADE semantic and 53.5 PQ on COCO panoptic. DaTaSeg also enables weakly-supervised knowledge transfer on ADE panoptic and Objects365 instance segmentation. Experiments show DaTaSeg scales with the number of training datasets and enables open-vocabulary segmentation through direct transfer. In addition, we annotate an Objects365 instance segmentation set of 1,000 images and release it as a public evaluation benchmark on https://laoreja.github.io/dataseg.
Xiuye Gu, Yin Cui, Jonathan Huang, Abdullah Rashwan, Xingyi Zhou, Golnaz Ghiasi, Weicheng Kuo, Huizhong Chen, Liang-Chieh Chen, David A. Ross
NeurIPS9
2022 A Simple Single-Scale Vision Transformer for Object Detection and Instance Segmentation
Wuyang Chen 0001, Xianzhi Du, Lucas Beyer, Xiaohua Zhai, Tsung-Yi Lin, Huizhong Chen, Xiaodan Song, Zhangyang Wang, Denny Zhou
ECCV (10)7
2015 Mobile Visual Search with Word-HOG Descriptors
abstract
Visual text information is a descriptive part of many images that can be used to perform mobile visual search (MVS) with particularly small queries. In this paper, we propose a system that uses word patch descriptors for retrieving images containing visual text. A random sampling method is used to find duplicate word patches in the database and reduce the database size. The system achieves comparable retrieval performance to state-of-the-art image feature-based systems for images of book covers, and performs better than state-of-the-art text-based retrieval systems for images of book pages. Using visual text to provide distinctive features, our system achieves more than 10-to-1 query size reduction for images of book covers and more than 16-to-1 query size reduction for images of book pages.
Sam S. Tsai, Huizhong Chen, David M. Chen, Bernd Girod
DCC2
2014 Efficient video search using image queries
abstract
We study the challenges of image-based retrieval when the database consists of videos. This variation of visual search is important for a broad range of applications that require indexing video databases based on their visual contents. We present new solutions to reduce storage requirements, while at the same time improving video search quality. The video database is preprocessed to find different appearances of the same visual elements, and build robust descriptors. Compression algorithms are developed to reduce system's storage requirements. We introduce a dataset of CNN broadcasts and queries that include photos taken with mobile phones and images of objects. Our experiments include pairwise matching and retrieval scenarios. We demonstrate one order of magnitude storage reduction and search quality improvements of up to 12% in mean average precision, compared to a baseline system that does not make use of our techniques.
André Araújo 0001, Mina Makar, Vijay Chandrasekhar 0001, David M. Chen, Sam S. Tsai, Huizhong Chen, Roland Angst, Bernd Girod
ICIP6
2014 Word-HOGs: Word histogram of oriented gradients for mobile visual search
abstract
Despite being a highly distinctive feature, the potential of text in images for mobile visual search has been largely neglected. Our research reported in this paper strives to improve mobile visual search by incorporating algorithms and compact representations tailored to visual text. We develop a new word patch descriptor, called Word Histogram of Oriented Gradients (Word-HOGs). The descriptor is based on gradient orientation histograms and can be compressed to a very small size using context-based arithmetic coding with lattice quantization. Because of its special structure, we can build image features from the descriptor and use these features for large-scale word patch matching. We show that the Word-HOG descriptor has a word patch matching performance that is better or comparable to the state-of-the-art approaches while being more efficient in feature counts, and it can be highly compressed with negligible loss in retrieval performance.
Sam S. Tsai, Huizhong Chen, David M. Chen, Bernd Girod
ICIP2
2014 Multi-modal Language Models for Lecture Video Retrieval
abstract
We propose Multi-modal Language Models (MLMs), which adapt latent variable techniques for document analysis to exploring co-occurrence relationships in multi-modal data. In this paper, we focus on the application of MLMs to indexing text from slides and speech in lecture videos, and subsequently employ a multi-modal probabilistic ranking function for lecture video retrieval. The MLM achieves highly competitive results against well established retrieval methods such as the Vector Space Model and Probabilistic Latent Semantic Analysis. When noise is present in the data, retrieval performance with MLMs is shown to improve with the quality of the spoken text extracted from the video.
Huizhong Chen, Matthew Cooper 0002, Dhiraj Joshi, Bernd Girod
ACM Multimedia1
2014 A Novel Saliency Detection Method for Lunar Remote Sensing Images
abstract
The saliency detection provides an alternative methodology to semantic image understanding in many applications, for example, content-based image retrieval. To detect saliency for lunar remote sensing images, this letter proposes a crater feature model by analyzing the relationship between local interest points and saliency of lunar images. Based on the model, we propose a novel saliency detection method for lunar images. Our method merges and combines the speed-up robust feature features of the highlight region and shadow region of an impact crater to get the candidate regions of interest (ROI). Then, a descriptive feature vector is generated for each ROI, and the resulting saliency regions are distinguished from false detected and inconspicuous ones through a support vector machine. The method has been put into test on Chang'e-1 and Chang'e-2 lunar image data, and confirmed to be able to detect the salient region of impact craters correctly, with results much better than those obtained by the classical saliency detection method.
Huizhong Chen, Ning Jing, Yongguang Chen
IEEE Geosci. Remote. Sens. Lett.1
2014 The Hidden Sides of Names - Face Modeling with First Name Attributes
abstract
This paper introduces the new idea of describing people using first names. We show that describing people in terms of similarity to a vector of possible first names is a powerful representation of facial appearance that can be used for a number of important applications, such as naming never-seen faces and building facial attribute classifiers. We build models for 100 common first names used in the US and for each pair, construct a pairwise first-name classifier. These classifiers are built using training images downloaded from the internet, with no additional user interaction. This gives our approach important advantages in building practical systems that do not require additional human intervention for data labeling. The classification scores from each pairwise name classifier can be used as a set of facial attributes to describe facial appearance. We show several surprising results. Our name attributes predict the correct first names of test faces at rates far greater than chance. The name attributes are applied to gender recognition and to age classification, outperforming state-of-the-art methods with all training images automatically gathered from the internet. We also demonstrate the powerful use of our name attributes for associating faces in images with names from caption, and the important application of unconstrained face verification.
Huizhong Chen, Andrew C. Gallagher, Bernd Girod
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 What's in a Name? First Names as Facial Attributes
abstract
This paper introduces a new idea in describing people using their first names, i.e., the name assigned at birth. We show that describing people in terms of similarity to a vector of possible first names is a powerful description of facial appearance that can be used for face naming and building facial attribute classifiers. We build models for 100 common first names used in the United States and for each pair, construct a pair wise first-name classifier. These classifiers are built using training images downloaded from the Internet, with no additional user interaction. This gives our approach important advantages in building practical systems that do not require additional human intervention for labeling. We use the scores from each pair wise name classifier as a set of facial attributes. We show several surprising results. Our name attributes predict the correct first names of test faces at rates far greater than chance. The name attributes are applied to gender recognition and to age classification, outperforming state-of-the-art methods with all training images automatically gathered from the Internet.
Huizhong Chen, Andrew C. Gallagher, Bernd Girod
CVPR1
2013 Content Based Retrieval for Lunar Exploration Image Databases
Huizhong Chen, Ning Jing, Yongguang Chen
DASFAA (2)1
2013 Face-graph matching for classifying groups of people
abstract
When people gather for a group photo, they are together for a social reason. Past work has shown that these social relationships affect how people position themselves in a group photograph. We propose classifying the type of group photo based on the spatial arrangement and the predicted attributes of the faces in the image. We propose a matching algorithm for finding images from a training set that have both similar arrangement of faces and attribute correspondence. We formulate the problem as a bipartite matching problem where the faces from each of the pair of images are nodes in the graph. Our work demonstrates that face arrangement, when combined with attribute (age and gender) correspondence, is a useful cue in capturing an approximate social essence of the group of people, and lets us understand why the group of people gathered for the photo.
Henry Shu, Andrew C. Gallagher, Huizhong Chen, Tsuhan Chen
ICIP3
2013 EigenNews: a personalized news video delivery platform
abstract
We demonstrate EigenNews, a personalized television news system. Upon visiting the EigenNews website, a user is shown a variety of news videos which have been automatically selected based on her individual preferences. These videos are extracted from 16 continually recorded television programs using a multimodal segmentation algorithm. Relevant metadata for each video are generated by linking videos to online news articles. Selected news videos can be watched in three different layouts and on various devices.
Matt C. Yu, Peter Vajda, David M. Chen, Sam S. Tsai, Maryam Daneshi, André Araújo 0001, Huizhong Chen, Bernd Girod
ACM Multimedia7
2012 Describing Clothing by Semantic Attributes
Huizhong Chen, Andrew C. Gallagher, Bernd Girod
ECCV (3)1
2012 Visual Text Features for Image Matching
abstract
We present a new class of visual text features that are based on text in camera phone images. A robust text detection algorithm locates individual text lines and feeds them to a recognition engine. From the recognized characters, we generate the visual text features in a way that resembles image features. We calculate their location, scale, orientation, and a descriptor that describes the character and word information. We apply visual text features to image matching. To disambiguate false matches, we developed a word-distance matching method. Our experiments with image that contain text show that the new visual text feature based image matching pipeline performs on par or better than a conventional image feature based pipeline while requiring less than 10 bits per feature. This is 4.5× smaller than state-of-the-art visual feature descriptors.
Sam S. Tsai, Huizhong Chen, David M. Chen, Vasu Parameswaran, Radek Grzeszczuk, Bernd Girod
ISM2
2012 Efficient Registration of Nonrigid 3-D Bodies
abstract
We present a novel method to perform an accurate registration of 3-D nonrigid bodies by using phase-shift properties of the dual-tree complex wavelet transform [Formula: see text]. Since the phases of [Formula: see text] coefficients change approximately linearly with the amount of feature displacement in the spatial domain, motion can be estimated using the phase information from these coefficients. The motion estimation is performed iteratively: first by using coarser level complex coefficients to determine large motion components and then by employing finer level coefficients to refine the motion field. We use a parametric affine model to describe the motion, where the affine parameters are found locally by substituting into an optical flow model and by solving the resulting overdetermined set of equations. From the estimated affine parameters, the motion field between the sensed and the reference data sets can be generated, and the sensed data set then can be shifted and interpolated spatially to align with the reference data set.
Huizhong Chen, Nick G. Kingsbury
IEEE Trans. Image Process.1
2011 Robust text detection in natural images with edge-enhanced Maximally Stable Extremal Regions
abstract
Detecting text in natural images is an important prerequisite. In this paper, we propose a novel text detection algorithm, which employs edge-enhanced Maximally Stable Extremal Regions as basic letter candidates. These candidates are then filtered using geometric and stroke width information to exclude non-text objects. Letters are paired to identify text lines, which are subsequently separated into words. We evaluate our system using the ICDAR competition dataset and our mobile document database. The experimental results demonstrate the excellent performance of the proposed method.
Huizhong Chen, Sam S. Tsai, Georg Schroth, David M. Chen, Radek Grzeszczuk, Bernd Girod
ICIP1
2011 Mobile visual search on printed documents using text and low bit-rate features
abstract
We present a novel mobile printed document retrieval system that utilizes both text and low bit-rate features. On the client phone, text are detected using an algorithm based on edge-enhanced Maximally Stable Extremal Regions. The title text image patch is rectified using a gradient based algorithm and recognized using Optical Character Recognition. Low bit-rate image features are extracted from the query image. Both text and compressed features are sent to a server. On the server, the title text is used for on-line search and the features are used for image-based comparison. The proposed system is capable of web-scale document retrieval using title text without the need of constructing a document image database. Using features for image-based comparison, we can reliably match retrieved documents to the query document. Last, by using text and low bit-rate features, we can reduce the transmitted query size significantly.
Sam S. Tsai, Huizhong Chen, David M. Chen, Georg Schroth, Radek Grzeszczuk, Bernd Girod
ICIP2
2011 Combining image and text features: a hybrid approach to mobile book spine recognition
abstract
Despite the successful use of local image features for large-scale object recognition, they are not effective in recognizing book spines on bookshelves. This is because some book spines contain only text components that do not yield distinguishing image features. To overcome this issue, we develop a new approach that combines a text-based spine recognition pipeline with an image feature-based spine recognition pipeline. The text within the book spine image is recognized and used as keywords to search a book spine text database. The image features of the book spine image are searched through a book spine image database. The search results of the two approaches are then carefully combined to form the final result. We implement the proposed hybrid book recognition pipeline used in a book inventory management system, and conduct extensive experiments to evaluate its performance. The experimental results show that while text-based or image feature-based systems only achieve a recall of 72%, the proposed hybrid system achieves a recall of ~91%.
Sam S. Tsai, David M. Chen, Huizhong Chen, Cheng-Hsin Hsu, Kyu-Han Kim, Jatinder Pal Singh, Bernd Girod
ACM Multimedia3
2011 The stanford mobile visual search data set
abstract
We survey popular data sets used in computer vision literature and point out their limitations for mobile visual search applications. To overcome many of the limitations, we propose the Stanford Mobile Visual Search data set. The data set contains camera-phone images of products, CDs, books, outdoor landmarks, business cards, text documents, museum paintings and video clips. The data set has several key characteristics lacking in existing data sets: rigid objects, widely varying lighting conditions, perspective distortion, foreground and background clutter, realistic ground-truth reference data, and query data collected from heterogeneous low and high-end camera phones. We hope that the data set will help push research forward in the field of mobile visual search.
Vijay Chandrasekhar 0001, David M. Chen, Sam S. Tsai, Ngai-Man Cheung, Huizhong Chen, Gabriel Takacs, Yuriy A. Reznik, Ramakrishna Vedantham, Radek Grzeszczuk, Jeff Bach, Bernd Girod
MMSys5