Li Weng

dblp:59/3954 · DBLP profile ↗
← Back
23ranked-venue papers
16as first author
8since 2021 · last 2025
0000-0002-7540-7652ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 10 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorSecurity and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Fusion of Global and Local Features with Multi-Inverted Indices for Image Retrieval
abstract
Feature fusion is an effective technique for improving image retrieval performance. Although the more feature types, the better accuracy, complexity also increases. Applications in practice typically afford a limited number of feature types. Due to the strong complementarity, global and local features form a natural combination for fusion applications. However, the two kinds of features are intrinsically different in nature, thus cannot be fused in a straightforward way. In this work, we propose an integrated image retrieval and feature fusion framework for global and local features. It is based on inverted index fusion, a technique for efficient image retrieval. The core idea is to rank candidates by weighted voting during candidate selection, which is named pre-ranking. This procedure takes place before re-ranking, and is potentially superior to conventional late fusion. Experiments on two public datasets show that the light-weight pre-ranking stage significantly contributes to accuracy, and brings substantial improvement when used together with re-ranking. Our method is a robust and versatile technique for image retrieval in the big data era.
Li Weng, Qianneng Wang, Bingya Wu
CBMI1
2025 SUMAC '25: 7th Workshop on analySis, Understanding and proMotion of heritAge Contents: Advances in Machine Learning, Signal Processing, Multimodal Techniques and Human-machine Interaction
abstract
SUMAC 2025 is the 7th edition of the workshop on analySis, Understanding and proMotion of heritAge Contents. It is held in Dublin, Ireland, on 27 October and is co-located with the 33rd ACM International Conference on Multimedia. The workshop's objective is to present and discuss the latest and most significant trends, challenges, and advances in the fields of machine learning, signal processing, multimodal techniques, and human-machine interaction. The workshop is dedicated to the valorization of cultural heritage, with an emphasis on unlocking and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field.
Valérie Gouet-Brunet, Edgar Roman-Rangel, Li Weng
ACM Multimedia3
2024 Reduced-Reference Learning for Target Localization in Deep Brain Stimulation
abstract
This work proposes a supervised machine learning method for target localization in deep brain stimulation (DBS). DBS is a recognized treatment for essential tremor. The effects of DBS significantly depend on the precise implantation of electrodes. Recent research on diffusion tensor imaging shows that the optimal target for essential tremor is related to the dentato-rubro-thalamic tract (DRTT), thus DRTT targeting has become a promising direction. The tractography-based targeting is more accurate than conventional ones, but still too complicated for clinical scenarios, where only structural magnetic resonance imaging (sMRI) data is available. In order to improve efficiency and utility, we consider target localization as a non-linear regression problem in a reduced-reference learning framework, and solve it with convolutional neural networks (CNNs). The proposed method is an efficient two-step framework, and consists of two image-based networks: one for classification and the other for localization. We model the basic workflow as an image retrieval process and define relevant performance metrics. Using DRTT as pseudo groundtruths, we show that individualized tractography-based optimal targets can be inferred from sMRI data with high accuracy. For two datasets of 280×220/272×227 (0.7/0.8 mm slice thickness) sMRI input, our model achieves an average posterior localization error of 2.3/1.2 mm, and a median of 1.7/1.02 mm. The proposed framework is a novel application of reduced-reference learning, and a first attempt to localize DRTT from sMRI. It significantly outperforms existing methods using 3D-CNN, anatomical and DRTT atlas, and may serve as a new baseline for general target localization problems.
Li Weng, Zhoule Zhu, Kaixin Dai, Junming Zhu, Hemmings Wu
IEEE Trans. Medical Imaging1
2023 SUMAC '23: 5th Workshop on the analySis, Understanding and proMotion of heritAge Contents: Advances in Machine Learning, Signal Processing, Multimodal Techniques and Human-machine Interaction
abstract
SUMAC 2023 is the fifth edition of the workshop on analySis, Understanding and proMotion of heritAge Contents. It is held in Ottawa, Canada on November 2, 2023 and is co-located with the 31st ACM International Conference on Multimedia. The workshop's objective is to present and discuss the latest and most significant trends, challenges and advances in the fields of machine learning, signal processing, multimodal techniques and human-machine interaction. The workshop is dedicated to the valorization of cultural heritage, with the emphasis on the unlocking of and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field. The complete SUMAC'23 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3581783.3610949.
Valérie Gouet-Brunet, Ronak Kosti, Li Weng
ACM Multimedia3
2023 Edge-Guided Recurrent Positioning Network for Salient Object Detection in Optical Remote Sensing Images
abstract
Optical remote sensing images (RSIs) have been widely used in many applications, and one of the interesting issues about optical RSIs is the salient object detection (SOD). However, due to diverse object types, various object scales, numerous object orientations, and cluttered backgrounds in optical RSIs, the performance of the existing SOD models often degrade largely. Meanwhile, cutting-edge SOD models targeting optical RSIs typically focus on suppressing cluttered backgrounds, while they neglect the importance of edge information which is crucial for obtaining precise saliency maps. To address this dilemma, this article proposes an edge-guided recurrent positioning network (ERPNet) to pop-out salient objects in optical RSIs, where the key point lies in the edge-aware position attention unit (EPAU). First, the encoder is used to give salient objects a good representation, that is, multilevel deep features, which are then delivered into two parallel decoders, including: 1) an edge extraction part and 2) a feature fusion part. The edge extraction module and the encoder form a U-shape architecture, which not only provides accurate salient edge clues but also ensures the integrality of edge information by extra deploying the intraconnection. That is to say, edge features can be generated and reinforced by incorporating object features from the encoder. Meanwhile, each decoding step of the feature fusion module provides the position attention about salient objects, where position cues are sharpened by the effective edge information and are used to recurrently calibrate the misaligned decoding process. After that, we can obtain the final saliency map by fusing all position attention cues. Extensive experiments are conducted on two public optical RSIs datasets, and the results show that the proposed ERPNet can accurately and completely pop-out salient objects, which consistently outperforms the state-of-the-art SOD models.
Xiaofei Zhou 0003, Kunye Shen, Li Weng, Runmin Cong, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Cybern.3
2022 SUMAC '22: 4th ACM International workshop on Structuring and Understanding of Multimedia heritAge Contents
abstract
SUMAC 2022 is the fourth edition of the workshop on Structuring and Understanding of Multimedia heritAge Contents. It is held in Lisboa, Portugal on October 10th, 2022 and is co-located with the 30th ACM International Conference on Multimedia. Its objective is to present and discuss the latest and most significant trends and challenges in the analysis, structuring and understanding of multimedia contents dedicated to the valorization of heritage, with the emphasis on the unlocking of and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field. The complete SUMAC'22 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3552464
Valérie Gouet-Brunet, Ronak Kosti, Li Weng
ACM Multimedia3
2021 SUMAC'21: 3rd Workshop on Structuring and Understanding of Multimedia heritAge Contents
abstract
SUMAC 2021 is the third edition of the workshop on Structuring and Understanding of Multimedia heritAge Contents. It is held in Chengdu, China on October 20th, 2021 and is co-located with the 29th ACM International Conference on Multimedia. Its objective is to present and discuss the latest and most significant trends and challenges in the analysis, structuring and understanding of multimedia contents dedicated to the valorization of heritage, with the emphasis on the unlocking of and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field.
Valérie Gouet-Brunet, Margarita Khokhlova, Ronak Kosti, Li Weng
ACM Multimedia4
2021 Semantic signatures for large-scale visual localization
Li Weng, Valérie Gouet-Brunet, Bahman Soheilian
Multim. Tools Appl.1
2018 Semantic Signatures for Urban Visual Localization
abstract
Visual localization is a useful alternative to standard localization techniques. In a typical scenario, features are extracted from images captured by cameras and compared with geo-referenced databases. Location information is then inferred from the matching results. Conventional schemes mainly use low-level visual features. They offer good accuracy but suffer from scalability issues. In order to assist localization in large urban areas, this work explores a different path by utilizing high-level semantic information. It is found that object information in a street view can facilitate localization. A novel descriptor scheme called “semantic signature” is proposed to summarize this information. A semantic signature consists of type and angle information of visible objects at a spatial location. Several metrics and protocols are proposed for signature comparison and retrieval. They illustrate different trade-offs between accuracy and complexity. Extensive simulation results confirm the potential of the proposed scheme in large-scale applications.
Li Weng, Bahman Soheilian, Valérie Gouet-Brunet
CBMI1
2016 A feature fusion framework for hashing
abstract
A hash algorithm converts data into compact strings. In the multimedia domain, effective hashing is the key to large-scale similarity search in high-dimensional feature space. A limit of existing hashing techniques is that they typically use single features. In order to improve search performance, it is necessary to utilize multiple features. Due to the compactness requirement, concatenation of hash values from different features is not an optimal solution. Thus a fusion process is desired. In this paper, we solve the multiple feature fusion problem by a hash bit selection framework. Given multiple features, we derive an n-bit hash value of improved performance compared with hash values of the same length computed from each individual feature. The framework utilizes a feature-independent hash algorithm to generate a sufficient number of bits from each feature, and selects n bits from the hash bit pool by leveraging pair-wise label information. The metric bit reliability is used for ranking the bits. It is estimated by bit-level hypothesis testing. In addition, we also take into account the dependence among bits. A weighted graph is constructed for refined bit selection, where the bit reliability is used as vertex weights and the mutual information among hash bits is used as edge weights. We demonstrate our framework with LSH. Extensive experiments confirm that our method is effective, and outperforms several state-of-the-art methods.
I-Hong Jhuo, Li Weng, Wen-Huang Cheng, D. T. Lee
ICPR2
2016 Privacy-Preserving Outsourced Media Search
abstract
This work proposes a privacy-protection framework for an important application called outsourced media search. This scenario involves a data owner, a client, and an untrusted server, where the owner outsources a search service to the server. Due to lack of trust, the privacy of the client and the owner should be protected. The framework relies on multimedia hashing and symmetric encryption. It requires involved parties to participate in a privacy-enhancing protocol. Additional processing steps are carried out by the owner and the client: (i) before outsourcing low-level media features to the server, the owner has to one-way hash them, and partially encrypt each hash-value; (ii) the client completes the similarity search by re-ranking the most similar candidates received from the server. One-way hashing and encryption add ambiguity to data and make it difficult for the server to infer contents from database items and queries, so the privacy of both the owner and the client is enforced. The proposed framework realizes trade-offs among strength of privacy enforcement, quality of search, and complexity, because the information loss can be tuned during hashing and encryption. Extensive experiments demonstrate the effectiveness and the flexibility of the framework.
Li Weng, Laurent Amsaleg, Teddy Furon
IEEE Trans. Knowl. Data Eng.1
2015 Supervised Multi-scale Locality Sensitive Hashing
abstract
LSH is a popular framework to generate compact representations of multimedia data, which can be used for content based search. However, the performance of LSH is limited by its unsupervised nature and the underlying feature scale. In this work, we propose to improve LSH by incorporating two elements - supervised hash bit selection and multi-scale feature representation. First, a feature vector is represented by multiple scales. At each scale, the feature vector is divided into segments. The size of a segment is decreased gradually to make the representation correspond to a coarse-to-fine view of the feature. Then each segment is hashed to generate more bits than the target hash length. Finally the best ones are selected from the hash bit pool according to the notion of bit reliability, which is estimated by bit-level hypothesis testing.
Li Weng, I-Hong Jhuo, Miaojing Shi, Meng Sun 0001, Wen-Huang Cheng, Laurent Amsaleg
ICMR1
2015 A Privacy-Preserving Framework for Large-Scale Content-Based Information Retrieval
abstract
We propose a privacy protection framework for large-scale content-based information retrieval. It offers two layers of protection. First, robust hash values are used as queries to prevent revealing original content or features. Second, the client can choose to omit certain bits in a hash value to further increase the ambiguity for the server. Due to the reduced information, it is computationally difficult for the server to know the client's interest. The server has to return the hash values of all possible candidates to the client. The client performs a search within the candidate list to find the best match. Since only hash values are exchanged between the client and the server, the privacy of both parties is protected. We introduce the concept oftunable privacy, where the privacy protection level can be adjusted according to a policy. It is realized through hash-based piecewise inverted indexing. The idea is to divide a feature vector into pieces and index each piece with a subhash value. Each subhash value is associated with an inverted index list. The framework has been extensively tested using a large image database. We have evaluated both retrieval performance and privacy-preserving performance for a particular content identification application. Two different constructions of robust hash algorithms are used. One is based on random projections; the other is based on the discrete wavelet transform. Both algorithms exhibit satisfactory performance in comparison with state-of-the-art retrieval schemes. The results show that the privacy enhancement slightly improves the retrieval performance. We consider the majority voting attack for estimating the query category and identification. Experiment results show that this attack is a threat when there are near-duplicates, but the success rate decreases with the number of omitted bits and the number of distinct items.
Li Weng, Laurent Amsaleg, April Morton, Stéphane Marchand-Maillet
IEEE Trans. Inf. Forensics Secur.1
2014 Image auto-annotation by exploiting web information
abstract
We consider the image auto-annotation problem by exploiting information from Internet. Given a collection of semantically similar images and a keyword that accurately describes these images, our goal is to find a set of popular tags to annotate each image, conforming to those used for similar images found on the web. We propose a novel framework to exploit classification based learning and bipartitioning clustering algorithms for extracting meaningful tags from semantical images on the web. Specifically, we adopt multiple kernel learning (MKL) to first select relevant images with their associated tags, which are obtained from the web based on keyword search, and then build a bipartite graph to model the relation between related tags and images. Finally, we partition over the bipartite graph to produce a set of significant tags for each image. We evaluate our proposed method by using the colorful Natural Scene and Events datasets to generate related images and tags from the Flickr website. The experimental results show that our proposed method has superior performance compared with baseline methods.
I-Hong Jhuo, Li Weng
ICIP2
2012 Robust Image Content Authentication with Tamper Location
abstract
We propose a novel image authentication system by combining perceptual hashing and robust watermarking. An image is divided into blocks. Each block is represented by a compact hash value. The hash value is embedded in the block. The authenticity of the image can be verified by re-computing hash values and comparing them with the ones extracted from the image. The system can tolerate a wide range of incidental distortion, and locate tampered areas as small as 1/64 of an image. In order to have minimal interference, we design both the hash and the watermark algorithms in the wavelet domain. The hash is formed by the sign bits of wavelet coefficients. The lattice-based QIM watermarking algorithm ensures a high payload while maintaining the image quality. Extensive experiments confirm the good performance of the proposal, and show that our proposal significantly outperforms a state-of-the-art algorithm.
Li Weng, Geert Braeckman, Ann Dooms, Bart Preneel, Peter Schelkens
ICME1
2011 Image Distortion Estimation by Hash Comparison
Li Weng, Bart Preneel
MMM (1)1
2010 A novel video hash algorithm
abstract
Perceptual hashing is an emerging solution for identification and authentication of multimedia content. In this work, a video hash algorithm is proposed. This algorithm computes a 180-bit hash value for videos of arbitrary lengths. The hash value can resist common signal processing and slight geometric distortion. The basic mechanism of the algorithm is to compute and accumulate frame hash values. A frame hash algorithm is designed by combining semi-global and local features. Semi-global features are extracted by computing several statistics from image blocks. Local features are extracted by computing a compact edge density map around stable feature points. The good performance of the new algorithm has been demonstrated by experiments.
Li Weng, Bart Preneel
ACM Multimedia1
2010 From Image Hashing to Video Hashing
Li Weng, Bart Preneel
MMM1
2009 Shape-based features for image hashing
abstract
Perceptual hashing is a solution for identification and authentication of multimedia content. The key of this technique is the extraction of proper features. In this paper, two features are proposed for natural image hashing. They are based on the description of shapes, in terms of contours and regions. The contour-based feature is formed by edge detection. The region-based feature is formed by the angular radial transform. Simulation results show that they have good robustness and discriminability. Compared to some other features, better ROC performance is achieved.
Li Weng, Bart Preneel
ICME1
2007 Attacking Some Perceptual Image Hash Algorithms
abstract
Perceptual hashing is an emerging solution for multimedia content authentication. Due to their robustness, such techniques might not work well when malicious attack is perceptually insignificant. We designed an experiment and verified that some state-of-the-art image hash algorithms could not distinguish small malicious distortion and some authentic distortion. We proposed an enhancement framework as a remedy. It suggests extracting information from the content and combining it with the secret key to generate the perceptual hash, so that perceptually insignificant information can be protected.
Li Weng, Bart Preneel
ICME1
2006 Using Space and Attribute Partitioned Partial Replicas for Data Subsetting and Aggregation Queries
abstract
Partial replication is one type of optimization to speed up execution of queries submitted to large datasets. In partial replication, a portion of the dataset is extracted, re-organized, and re-distributed across the storage system. In this paper we investigate methods for efficient execution of queries when replicas of a dataset exist; we assume the replicas have already been created and do not target the replica creation problem. We propose a cost model and algorithm for combined use of space partitioned and attribute partitioned replicas for executing data subsetting range queries. We extend the cost model and propose a greedy algorithm to address range queries with aggregation operations. The extended replica selection algorithm allows uneven partitioning of replicas across storage nodes. Different replicas can be partitioned across different subsets of storage nodes. We have implemented these techniques as part of an automatic data virtualization system and have evaluated the benefits of our techniques using this system. We demonstrate the efficacy of the algorithms on parallel machines using queries on datasets from oil reservoir simulation studies and satellite data processing applications
Li Weng, Ümit V. Çatalyürek, Tahsin M. Kurç, Gagan Agrawal, Joel H. Saltz
ICPP1
2005 Servicing range queries on multidimensional datasets with partial replicas
abstract
Partial replication is one type of optimization to speed up execution of queries submitted to large datasets. In partial replication, a portion of the dataset is extracted, re-organized, and re-distributed across the storage system. The objective is to reduce the volume of I/O and increase I/O parallelism for different types of queries and for the portions of the dataset that are likely to be accessed frequently. When multiple partial replicas of a dataset exist, query execution plan should be generated so as to use the best combination of subsets of partial replicas (and possibly the original dataset) to minimize query execution time. In this paper, we present a compiler and runtime approach for range queries submitted against distributed scientific datasets. A heuristic algorithm is proposed to choose the set of replicas to reduce query execution. We show the efficiency of the proposed method using datasets and queries in oil reservoir simulation studies on a cluster machine.
Li Weng, Ümit V. Çatalyürek, Tahsin M. Kurç, Gagan Agrawal, Joel H. Saltz
CCGRID1
2004 An Approach for Automatic Data Virtualization
Li Weng, Gagan Agrawal, Ümit V. Çatalyürek, Tahsin M. Kurç, Sivaramakrishnan Narayanan, Joel H. Saltz
HPDC1