Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Lezi Wang

dblp:25/9997 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
2since 2021 · last 2024
0000-0001-7927-6621ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 48% Deep learning architectures and training · 16% Trustworthy machine learning · 14%
Theoretical computer science
2 papers
Mathematical optimization · 100%
Computer graphics and multimedia
2 papers
Computer animation and physical simulation · 60% Multimedia analysis and retrieval · 40%
Human-computer interaction and pervasive computing
1 paper
Immersive interaction · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization › constrained optimization
duality theory
0.722020
Dual Iterative Hard Thresholding · J. Mach. Learn. Res. 2020
Dual Iterative Hard Thresholding: From Non-convex Sparse Minimization to Non-smooth Concave Maximization · ICML 2017
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
iterative hard thresholding
0.722020
Dual Iterative Hard Thresholding · J. Mach. Learn. Res. 2020
Dual Iterative Hard Thresholding: From Non-convex Sparse Minimization to Non-smooth Concave Maximization · ICML 2017
Machine learning › Generative modeling
diffusion model
0.712023
Social Diffusion: Long-term Multiple Human Motion Anticipation · ICCV 2023
Computer animation and physical simulation
human motion prediction
0.712023
Social Diffusion: Long-term Multiple Human Motion Anticipation · ICCV 2023
Machine learning › Deep learning architectures and training
attention mechanism
0.412020
Learning Trailer Moments in Full-Length Movies with Co-Contrastive Attention · ECCV (18) 2020
Multimedia analysis and retrieval
video analysis
0.412020
Learning Trailer Moments in Full-Length Movies with Co-Contrastive Attention · ECCV (18) 2020
Mathematical optimization
sparse optimization
0.412020
Dual Iterative Hard Thresholding · J. Mach. Learn. Res. 2020
Computer vision › Image recognition and object detection
image classification
0.412019
Sharpen Focus: Learning With Attention Separability and Consistency · ICCV 2019
Machine learning › Trustworthy machine learning
interpretability
0.412019
Sharpen Focus: Learning With Attention Separability and Consistency · ICCV 2019
Machine learning › Generative modeling › cross-modal generation
story visualization
0.312018
Show Me a Story: Towards Coherent Neural Story Illustration · CVPR 2018
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.312018
Show Me a Story: Towards Coherent Neural Story Illustration · CVPR 2018
Computer vision › Face, body and person analysis
human pose estimation
0.212024
EgoBody3M: Egocentric Body Tracking on a VR Headset using a Diverse Dataset · ECCV (79) 2024

Methods — techniques the papers use, named apart from their topics

temporal convolutional network · 1.3order-invariant aggregation · 1.3diffusion model · 1.3co-contrastive attention · 0.9super-gradient ascent · 0.4cross-layer consistency · 0.4attention modeling · 0.4order embedding loss · 0.3encoder-decoder · 0.3GRU · 0.3supergradient method · 0.3stochastic optimization · 0.3projected gradient descent · 0.3
YearPublicationVenuePosition
2024 EgoBody3M: Egocentric Body Tracking on a VR Headset using a Diverse Dataset
Amy Zhao, Chengcheng Tang, Lezi Wang, Mihika Dave, Lingling Tao, Christopher D. Twigg, Robert Y. Wang
ECCV (79)3
2023 Social Diffusion: Long-term Multiple Human Motion Anticipation
abstract
We propose Social Diffusion, a novel method for short-term and long-term forecasting of the motion of multiple persons as well as their social interactions. Jointly forecasting motions for multiple persons involved in social activities is inherently a challenging problem due to the interdependencies between individuals. In this work, we leverage a diffusion model conditioned on motion histories and causal temporal convolutional networks to forecast individually and contextually plausible motions for all participants. The contextual plausibility is achieved via an order-invariant aggregation function. As a second contribution, we design a new evaluation protocol that measures the plausibility of social interactions which we evaluate on the Haggling dataset, which features a challenging social activity where people are actively taking turns to talk and switching their attention. We evaluate our approach on four datasets for multi-person forecasting where our approach outperforms the state-of-the-art in terms of motion realism and contextual plausibility.
Julian Tanke, Linguang Zhang, Amy Zhao, Chengcheng Tang, Yujun Cai, Lezi Wang, Po-Chen Wu, Juergen Gall, Cem Keskin
ICCV6
2020 Learning Trailer Moments in Full-Length Movies with Co-Contrastive Attention
Lezi Wang, Dong Liu 0002, Rohit Puri, Dimitris N. Metaxas
ECCV (18)1
2020 Dual Iterative Hard Thresholding
abstract
Iterative Hard Thresholding (IHT) is a popular class of first-order greedy selection methods for loss minimization under cardinality constraint. The existing IHT-style algorithms, however, are proposed for minimizing the primal formulation. It is still an open issue to explore duality theory and algorithms for such a non-convex and NP-hard combinatorial optimization problem. To address this issue, we develop in this article a novel duality theory for $\ell_2$-regularized empirical risk minimization under cardinality constraint, along with an IHT-style algorithm for dual optimization. Our sparse duality theory establishes a set of sufficient and/or necessary conditions under which the original non-convex problem can be equivalently or approximately solved in a concave dual formulation. In view of this theory, we propose the Dual IHT (DIHT) algorithm as a super-gradient ascent method to solve the non-smooth dual problem with provable guarantees on primal-dual gap convergence and sparsity recovery. Numerical results confirm our theoretical predictions and demonstrate the superiority of DIHT to the state-of-the-art primal IHT-style algorithms in model estimation accuracy and computational efficiency.
Xiao-Tong Yuan, Bo Liu 0005, Lezi Wang, Qingshan Liu 0001, Dimitris N. Metaxas
J. Mach. Learn. Res.3
2019 Distributed Inexact Newton-type Pursuit for Non-convex Sparse Learning
abstract
In this paper, we present a sample distributed greedy pursuit method for non-convex sparse learning under cardinality constraint. Given the training samples uniformly randomly partitioned across multiple machines, the proposed method alternates between local inexact sparse minimization of a Newton-type approximation and centralized global results aggregation. Theoretical analysis shows that for a general class of convex functions with Lipschitze continues Hessian, the method converges linearly with contraction factor scaling inversely to the local data size; whilst the communication complexity required to reach desirable statistical accuracy scales logarithmically with respect to the number of machines for some popular statistical learning models. For nonconvex objective functions, up to a local estimation error, our method can be shown to converge to a local stationary sparse solution with sub-linear communication complexity. Numerical results demonstrate the efficiency and accuracy of our method when applied to large-scale sparse learning tasks including deep neural nets pruning
Bo Liu 0005, Xiao-Tong Yuan, Lezi Wang, Qingshan Liu 0001, Junzhou Huang, Dimitris N. Metaxas
AISTATS3
2019 Attention-based Facial Behavior Analytics inSocial Communication
Lezi Wang, Chongyang Bai, Maksim Bolonkin, Judee K. Burgoon, Norah E. Dunbar, V. S. Subrahmanian, Dimitris N. Metaxas
BMVC1
2019 Sharpen Focus: Learning With Attention Separability and Consistency
abstract
Recent developments in gradient-based attention modeling have seen attention maps emerge as a powerful tool for interpreting convolutional neural networks. Despite good localization for an individual class of interest, these techniques produce attention maps with substantially overlapping responses among different classes, leading to the problem of visual confusion and the need for discriminative attention. In this paper, we address this problem by means of a new framework that makes class-discriminative attention a principled part of the learning process. Our key innovations include new learning objectives for attention separability and cross-layer consistency, which result in improved attention discriminability and reduced visual confusion. Extensive experiments on image classification benchmarks show the effectiveness of our approach in terms of improved classification accuracy, including CIFAR-100 (+3.33%), Caltech-256 (+1.64%), ImageNet (+0.92%), CUB-200-2011 (+4.8%) and PASCAL VOC2012 (+5.73%).
Lezi Wang, Ziyan Wu 0001, Srikrishna Karanam, Kuan-Chuan Peng, Rajat Vikram Singh, Bo Liu 0005, Dimitris N. Metaxas
ICCV1
2019 A coupled encoder-decoder network for joint face detection and landmark localization
Lezi Wang, Xiang Yu 0002, Thirimachos Bourlai, Dimitris N. Metaxas
Image Vis. Comput.1
2018 Show Me a Story: Towards Coherent Neural Story Illustration
abstract
We propose an end-to-end network for visual illustration of a sequence of sentences forming a story. At the core of our model is the ability to model the inter-related nature of the sentences within a story, as well as the ability to learn coherence to support reference resolution. The framework takes the form of an encoder-decoder architecture, where sentences are encoded using a hierarchical two-level sentence-story GRU, combined with an encoding of coherence, and sequentially decoded using a predicted feature representation into a consistent illustrative image sequence. We optimize all parameters of our network in an end-to-end fashion with respect to order embedding loss, encoding entailment between images and sentences. Experiments on the VIST storytelling dataset [9] highlight the importance of our algorithmic choices and efficacy of our overall model.
Hareesh Ravi, Lezi Wang, Carlos Muñiz 0001, Leonid Sigal, Dimitris N. Metaxas, Mubbasir Kapadia
CVPR2
2017 A Coupled Encoder-Decoder Network for Joint Face Detection and Landmark Localization
abstract
Face detection and landmark localization have been extensively investigated and are the prerequisite for many face applications, such as face recognition and 3D face reconstruction. Most existing methods achieve success on only one of the two problems. In this paper, we propose a coupled encoder-decoder network to jointly detect faces and localize facial key points. The encoder and decoder generate response maps for facial landmark localization. Moreover, we observe that the intermediate feature maps from the encoder and decoder have strong power in describing facial regions, which motivates us to build a unified framework by coupling the feature maps for multi-scale cascaded face detection. Experiments on face detection show strongly competitive results against the existing methods on two public benchmarks. The landmark localization further shows consistently better accuracy than state-of-the-arts on three face-in-the-wild databases.
Lezi Wang, Xiang Yu 0002, Dimitris N. Metaxas
FG1
2017 Dual Iterative Hard Thresholding: From Non-convex Sparse Minimization to Non-smooth Concave Maximization
abstract
Iterative Hard Thresholding (IHT) is a class of projected gradient descent methods for optimizing sparsity-constrained minimization models, with the best known efficiency and scalability in practice. As far as we know, the existing IHT-style methods are designed for sparse minimization in primal form. It remains open to explore duality theory and algorithms in such a non-convex and NP-hard setting. In this article, we bridge the gap by establishing a duality theory for sparsity-constrained minimization with $\ell_2$-regularized objective and proposing an IHT-style algorithm for dual maximization. Our sparse duality theory provides a set of sufficient and necessary conditions under which the original NP-hard/non-convex problem can be equivalently solved in a dual space. The proposed dual IHT algorithm is a super-gradient method for maximizing the non-smooth dual objective. An interesting finding is that the sparse recovery performance of dual IHT is invariant to the Restricted Isometry Property (RIP), which is required by all the existing primal IHT without sparsity relaxation. Moreover, a stochastic variant of dual IHT is proposed for large-scale stochastic optimization. Numerical results demonstrate that dual IHT algorithms can achieve more accurate model estimation given small number of training data and have higher computational efficiency than the state-of-the-art primal IHT-style algorithms.
Bo Liu 0005, Xiao-Tong Yuan, Lezi Wang, Qingshan Liu 0001, Dimitris N. Metaxas
ICML3
2013 An efficient graph-based visual reranking
abstract
The state of the art in query expansion is mainly based on the spatial information. These methods achieve high performance, however, suffer from huge computation and memory. The objective of this paper is to perform visual reranking in near-real time regardless of the spatial information. We explore a graph-based method proposed as our confident sample detection baseline, which has been proved successful in achieving high precision. In addition, a novel maximum-kernel-based metric function is introduced to rerank the images in the initial result. We evaluated the method on the standard Paris dataset and a new Francelandmark dataset. Our experiments demonstrate that the algorithm has great value on practicality because of its good performance, easy implementation, and high computational efficiency.
Hongliang Bai, Lezi Wang, Shusheng Cen, Jian Zhao 0001
ICASSP4
2013 A semantic graph-based algorithm for image search reranking
abstract
Image search reranking has become a widely-used approach to significantly boost retrieval performance in the state-of-art content-based image retrieval system. Most of the methods merely rely on matching visual distances between query and initial results or among initial results to detect confident samples relevant to query. However, they may fail to rerank due to the existence of a huge gap between low-level visual features and high-level semantic concepts. In this paper, we propose to detect reliable relevant samples based on a semantic image graph of labeled auxiliary dataset and Markov random walk algorithm. A graph-based rerank method is then presented to propagate the scores of detected confident samples to the rest. Our method is evaluated on the standard Paris dataset and a France dataset introduced by us. The performance is demonstrated to match or exceed the state-of-art.
Hongliang Bai, Lezi Wang, Shusheng Cen, Jian Zhao 0001
ICASSP4
2013 Interactive Video Retrieval Using Combination of Semantic Index and Instance Search
Hongliang Bai, Lezi Wang, Kun Tao
MMM (2)2
2012 Contented-Based Large Scale Web Audio Copy Detection
abstract
The exponential growth of web videos brings content based copy detection into a crucial issue. Besides the image information, audio also plays an important role in copy detection. In this paper, the audio-based copy detection framework is introduced. Three contributions are presented: (1) the band energy difference based feature is improved by adding multi-scale information, which extends the candidate feature sets, (2) a conditional entropy based method is used to select 16 ordinal relations to generate a more compact and robust feature combination among the random $C_{91}^{16}\approx2.6\times10^{17}$ combinations, (3) the result-based fusion strategy is introduced to recall the missed true positives. The proposed algorithm outperforms the traditional coarse fingerprints, shown by experiments conducted in the TRECVID 2011 Content-based Copy Detection (CCD) database.
Lezi Wang, Hongliang Bai, Jiwei Zhang 0001, Wei Liu 0100
ICME1
2011 TV program segmentation using multi-modal information fusion
abstract
A TV program segmentation algorithm is presented by the fusion of the multi-modal information in the large-scale videos. As "Inter-Programs" are generally inserted into the TV videos repeatedly, the macro structures of the videos can be effectively and automatically generated by identifying the video-audio features of the special sequences. The Electronic Program Guide (EPG) is used to organize the structures into the programs. Three sections are included in the algorithm, namely, the video-based non-supervised duplicate sequence detection, the audio-based special clip retrieval and the EPG-based 24-hour program segmentation. The algorithm has been tested in 60-day different-type TV videos. The F-measures of the multi-modal fusion and video-based duplicated sequence detection achieve the rates of over 98% and 96% respectively. These results show that the proposed method is highly efficient and effective for the TV Program segmentation.
Hongliang Bai, Lezi Wang, Jiwei Zhang 0001, Kun Tao, Xiaofu Chang
ICMR2