Yicheng Song

dblp:95/4262 · DBLP profile ↗
← Back
19ranked-venue papers
11as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 KP-Agent: Keyword Pruning in Sponsored Search Advertising via LLM-Powered Contextual Bandits
Hou-Wan Long, Yicheng Song, Tianshu Sun
CIKM2
2025 Scoring, Remember, and Reference: Catching Camouflaged Objects in Videos
abstract
Video Camouflaged Object Detection (VCOD) aims to segment objects whose appearances closely resemble their surroundings, posing a challenging and emerging task. Existing vision models often struggle in such scenarios due to the indistinguishable appearance of camouflaged objects and the insufficient exploitation of dynamic information in videos. To address these challenges, we propose an end-to-end VCOD framework inspired by human memory-recognition, which leverages historical video information by integrating memory reference frames for camouflaged sequence processing. Specifically, we design a dual-purpose decoder that simultaneously generates predicted masks and scores, enabling reference frame selection based on scores while introducing auxiliary supervision to enhance feature extraction.Furthermore, this study introduces a novel reference-guided multilevel asymmetric attention mechanism, effectively integrating long-term reference information with short-term motion cues for comprehensive feature extraction. By combining these modules, we develop the Scoring, Remember, and Reference (SRR) framework, which efficiently extracts information to locate targets and employs memory guidance to improve subsequent processing. With its optimized module design and effective utilization of video data, our model achieves significant performance improvements, surpassing existing approaches by 10% on benchmark datasets while requiring fewer parameters (54M) and only a single pass through the video. The code will be made publicly available.
Yu'ang Feng, Shuyong Gao, Fuzhen Yan, Yicheng Song, Lingyi Hong
ICCV4
2024 Landmine Detection Based on Riemannian Phasor Quaternion Self-Organizing Map
abstract
Ground penetrating radar (GPR) based landmine detection has advantages such as high safety and high efficiency. There are various methods to process data acquired from GPR systems. A common method is the Riemannian quaternion self-organizing map (RQSOM), which effectively enables self-organization of polarization data in quaternion form for visualization. However, RQSOM does not take into account the phase information of scattering components. Studies have shown that phase information is correlated with polarization information and can be effectively integrated into a new form, phasor quaternion (PQ). To visualize PQ-type feature vectors, in this paper, we propose a novel algrithm, Riemannian phasor quaternion self-organizing map (RPQSOM). RPQSOM utilizes the geometric characteristics of the phasor part to effectively learn useful information in PQ, thereby achieving superior visualization performance. We conduct experiments for visualizing a mock landmine. The experimental results demonstrate that RPQSOM fulfills our design objectives and gets better visualization results compared to RQSOM.
Yicheng Song, Ryo Natsuaki, Akira Hirose 0001
IGARSS1
2023 Landmine Detection Based on Generalized Riemannian Quaternion Self-Organizing Map
abstract
Ground penetrating radar (GPR) based landmine detection has advantages such as high safety and high efficiency. There are various methods to process the data obtained from GPR systems. One of the common methods is Riemannian quaternion self-organizing map (RQSOM), which can effectively make the polarization data self-organize for visualization. However, RQSOM cannot take into account the spatial degree of polarization (DoP) of the extracted data. Spatial DoP contains useful information for landmine visualization. To overcome the limitation, in this paper, we propose a novel algorithm, generalized Riemannian quaternion self-organizing map (GRQSOM), which utilize both polarization and spatial DoP during self-organization. Thus, better visualization performance can be obtained. We conduct experiments for the visualization of a mock landmine. The experimental results show that GRQSOM achieves our design and gets better visualization results compared with RQSOM.
Yicheng Song, Ryo Natsuaki, Akira Hirose 0001
IGARSS1
2023 Towards End-to-End Unsupervised Saliency Detection with Self-Supervised Top-Down Context
abstract
Unsupervised salient object detection aims to detect salient objects without using supervision signals eliminating the tedious task of manually labeling salient objects. To improve training efficiency, end-to-end methods for USOD have been proposed as a promising alternative. However, current solutions rely heavily on noisy handcraft labels and fail to mine rich semantic information from deep features. In this paper, we propose a self-supervised end-to-end salient object detection framework via top-down context. Specifically, motivated by contrastive learning, we exploit the self-localization from the deepest feature to construct the location maps which are then leveraged to learn the most instructive segmentation guidance. Further considering the lack of detailed information in deepest features, we exploit the detail-boosting refiner module to enrich the location labels with details. Moreover, we observe that due to lack of supervision, current unsupervised saliency models tend to detect non-salient objects that are salient in some other samples of corresponding scenarios. To address this widespread issue, we design a novel Unsupervised Non-Salient Suppression (UNSS) method developing the ability to ignore non-salient objects. Extensive experiments on benchmark datasets demonstrate that our method achieves leading performance among the recent end-to-end methods and most of the multi-stage solutions. The code is available.
Yicheng Song, Shuyong Gao, Haozhe Xing, Yiting Cheng 0001, Yan Wang 0068
ACM Multimedia1
2023 Riemannian Quaternion Self-Organizing Map to Overcome Degree-of-Polarization Error in Polarimetric Ground-Penetrating Radar
abstract
Ground penetrating radar (GPR) based landmine detection has advantages such as high safety and high efficiency. There are various methods to process the data obtained from GPR systems. One of the commonly used methods is to visualize the quaternion-type polarization data by quaternion self-organizing map (QSOM). However, QSOM can not take into account the geometric property of the polarization data. Then, degree-of-polarization (DoP) error is introduced in the self-organization process of QSOM, which leads to unsatisfactory visualization results. To overcome the limitation, in this paper, we propose a novel processing, Riemannian quaternion self-organizing map (RQSOM), which takes into account the geometric property to eliminate the DoP error with the help of logarithmic and exponential maps in Riemannian geometry. We analyze its basic dynamics and compare it with that of QSOM. We conduct experiments to visualize a mock plastic landmine with QSOM and RQSOM. The experimental results show that RQSOM realizes expected self-organization dynamics, thereby achieving better visualization results compared to QSOM.
Yicheng Song, Ryo Natsuaki, Akira Hirose 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Proposal of Detection of Subsurface Objects with Model-Based Homogeneity to Extend Compressed Sensing
abstract
This paper proposes model-based homogeneity (MBH) to extend compressed sensing (CS) for detection of subsurface objects. The scattered waves at landmines have features spe-cific to their material and structure. The features spread over a landmine, showing its shape. We calculate the MBH value from spatial distribution of scattering feature vectors by using a model having the landmine shape. We can utilize CS with MBH because the result of MBH calculation is sparse in general. This method eliminates clutter and reduces mea-surement points. Experiments demonstrate that this method makes the time for detection one-twentieth of that in a conventional method.
Ryuta Imai, Yicheng Song, Ryo Natsuaki, Akira Hirose 0001
IGARSS2
2022 Uncovering Characteristic Response Paths of a Population
abstract
We propose an approach for uncovering characteristic response paths of a population from an individual-level multivariate time series data set. The approach is based on a model that accommodates a set of arbitrary distributions for endogenous variables and interstep intervals and variables. The model enables reliable estimation of individual-level parameters by uncovering and statistically pooling clusters of similar individuals. We show that using such a model one can distribute the response of an outcome variable to an impulse over all possible preceding activity sequences. When a few such sequences explain most of the response, they describe the population’s characteristic response paths from the impulse to the outcome. We apply the proposed approach to a customer touchpoint data set from a large multichannel specialty retailer. This application uncovers six customer segments, each with unique characteristic paths to purchase. These paths provide insights into the behavior of customers and the optimal over-time communication strategy for different customer segments. Summary of Contribution: Uncovering users’ paths through physical and virtual spaces has been of considerable interest in the computing and operations research domain. The existing research suggests a demand for visualizing the primary paths of agents through geographic, online, and activity spaces. Thus far, most of the research have developed approaches that are unique to specific domains providing insight into the domain in the process. There is a need for a general statistically robust approach that can be applied to a broad range of domains to uncover variable sequences that lead to outcomes of interest. We propose a computational approach to uncover characteristic response paths of a population from an individual-level multivariate time series data set. The approach is based on a statistical model that accommodates arbitrary and mixed set of distributions for the endogenous variables, accommodates intersession intervals and variables, and reliably estimates individuals’ parameters through statistical pooling by uncovering clusters of similar members. These features make the proposed model suitable for a large variety of real-world datasets. We show that using such a model one can extract characteristic paths over possible activity sequences starting from an impulse leading up to a target variable of interest.
Yicheng Song, Nachiketa Sahoo, Shuba Srinivasan, Chrysanthos Dellarocas
INFORMS J. Comput.1
2022 Model-Based Homogeneity to Extend Compressed Sensing for Ground Penetrating Radar
abstract
This paper proposes model-based homogeneity (MBH) to extend compressed sensing (CS) for landmine-detection ground penetrating radar (GPR). Conventional CS methods have difficulty in distinguishing landmines from clutter since it principally pays attention to signal magnitude. In contrast, our method visualizes landmines based on homogeneity of high-dimensional scattering features in a spatial model. It realizes both the exclusion of clutter and the reduction of measurement points. Experiments demonstrate that the total measurement and processing time is reduced to one-twentieth of a conventional dense measurement case. We also investigate the influence of model size and number of landmines on the performance. The proposed method is capable of visualizing any objects having respective shapes by configuring corresponding models.
Ryuta Imai, Yicheng Song, Ryo Natsuaki, Akira Hirose 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Phasor-Quaternion Self-Organizing-Map-Based Ground Penetrating Radar Systems
abstract
Visualization by ground penetrating radar (GPR) systems has a wide demand in various application fields. Our conventional visualization method mainly extracts feature vectors based on the frequency and spatial correlation of scattering parameters and then uses complex-valued self-organizing maps (CSOM) in unsupervised grouping to achieve visualization. However, this method sometimes shows problems of unclear boundary and deformed shape of a detected object. In this article, we propose a novel algorithm, phasor-quaternion self-organizing map (PQSOM), based on the Poincare vectors combined with phase, which is organized in phasor-quaternion (PQ) form. By considering both the polarization state and the phase information as respective counterparts, we achieve a more accurate grouping result. Through visualization experiments of underground landmines, we find that, compared with the previous methods, the boundary clarity and the shape of objects are improved.
Yicheng Song, Akira Hirose 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 Proposal of a Ground Penetrating Radar System Utilizing Polarization Information by Using Phasor-Quaternion Self-Organizing Map
abstract
Previously we proposed a ground penetrating radar (GPR) system employing complex-valued self-organizing maps (CSOM) to deal with feature vectors based on the frequency and spatial correlation of scattering parameters. It performs unsupervised grouping to achieve visualization. However, this method sometimes shows problems of unclear boundary and shape distortion of a detected object. In this paper, we propose a novel method, phasor quaternion self-organizing map (PQSOM), based on the Poincare vectors combined with phase, which are organized in phasor quaternion (PQ) form. By considering both the polarization state and the phase information as respective counterparts, PQSOM is able to self-organize more effectively, which in turn leads to better grouping performance. Through visualization experiments of landmines, we find that PQSOM has better grouping results compared with CSOM and QSOM.
Yicheng Song, Akira Hirose 0001
IGARSS1
2014 A Unified Geolocation Framework for Web Videos
abstract
In this article, we propose a unified geolocation framework to automatically determine where on the earth a web video was shot. We analyze different social, visual, and textual relationships from a real-world dataset and find four relationships with apparent geography clues that can be used for web video geolocation. Then, the geolocation process is formulated as an optimization problem that simultaneously takes the social, visual, and textual relationships into consideration. The optimization problem is solved by an iterative procedure, which can be interpreted as a propagation of the geography information among the web video social network. Extensive experiments on a real-world dataset clearly demonstrate the effectiveness of our proposed framework, with the geolocation accuracy higher than state-of-the-art approaches.
Yicheng Song, Yongdong Zhang 0001, Juan Cao 0001, Jinhui Tang 0001, Xingyu Gao 0001, Jintao Li 0001
ACM Trans. Intell. Syst. Technol.1
2012 Web Video Geolocation by Geotagged Social Resources
abstract
This paper considers the problem of web video geolocation: we hope to determine where on the Earth a web video was taken. By analyzing a 6.5-million geotagged web video dataset, we observe that there exist inherent geography intimacies between a video with its relevant videos (related videos and same-author videos). This social relationship supplies a direct and effective cue to locate the video to a particular region on the earth. Based on this observation, we propose an effective web video geolocation algorithm by propagating geotags among the web video social relationship graph. For the video that have no geotagged relevant videos, we aim to collect those geotagged relevant images that are content similar with the video (share some visual or textual information with the video) as the cue to infer the location of the video. The experiments have demonstrated the effectiveness of both methods, with the geolocation accuracy much better than state-of-the-art approaches. Finally, an online web video geolocation system: Video2Locatoin (V2L) is developed to provide public access to our algorithm.
Yicheng Song, Yongdong Zhang 0001, Juan Cao 0001, Tian Xia 0002, Wu Liu 0005, Jintao Li 0001
IEEE Trans. Multim.1
2011 Web video retagging
Zhineng Chen, Juan Cao 0001, Tian Xia 0002, Yicheng Song, Yongdong Zhang 0001, Jintao Li 0001
Multim. Tools Appl.4
2010 Web video categorization based on Wikipedia categories and content-duplicated open resources
abstract
This paper presents a novel approach for web video categorization by leveraging Wikipedia categories (WikiCs) and open resources describing the same content as the video, i.e., content-duplicated open resources (CDORs). Note that current approaches only collect CDORs within one or a few media forms and ignore CDORs of other forms. We explore all these resources by utilizing WikiCs and commercial search engines. Given a web video, its discriminative Wikipedia concepts are first identified and classified. Then a textual query is constructed and from which CDORs are collected. Based on these CDORs, we propose to categorize web videos in the space spanned by WikiCs rather than that spanned by raw tags. Experimental results demonstrate the effectiveness of both the proposed CDOR collection method and the WikiC voting categorization algorithm. In addition, the categorization model built based on both WikiCs and CDORs achieves better performance compared with the models built based on only one of them as well as state-of-the-art approach.
Zhineng Chen, Juan Cao 0001, Yicheng Song, Yongdong Zhang 0001, Jintao Li 0001
ACM Multimedia3
2010 Tag transformer
abstract
Human annotations (titles and tags) of web videos facilitate most web video applications. However, the raw tags are noisy, sparse and structureless, which limit the effectiveness of tags. In this paper, we propose a tag transformer schema to solve these problems. We first eliminate those imprecise and meaningless tags with Wikipedia, and then transform the remaining tags to the Wikipedia category set to gather a precise, complete and structural description of the tags. Our experimental results on web video categorization demonstrate the superiority of the transformed space. We also apply tag transformer into the first study of using Wikipedia category system to structurally recommend the related videos. The online user study of the demo system suggests that our method could bring fantastic experience to the web users.
Yicheng Song, Juan Cao 0001, Zhineng Chen, Yongdong Zhang 0001, Jintao Li 0001
ACM Multimedia1
2010 Context-oriented web video tag recommendation
abstract
Tag recommendation is a common way to enrich the textual annotation of multimedia contents. However, state-of-the-art recommendation methods are built upon the pair-wised tag relevance, which hardly capture the context of the web video, i.e., when who are doing what at where. In this paper we propose the context-oriented tag recommendation (CtextR) approach, which expands tags for web videos under the context-consistent constraint. Given a web video, CtextR first collects the multi-form WWW resources describing the same event with the video, which produce an informative and consistent context; and then, the tag recommendation is conducted based on the obtained context. Experiments on an 80,031 web video collection show CtextR recommends various relevant tags to web videos. Moreover, the enriched tags improve the performance of web video categorization.
Zhineng Chen, Juan Cao 0001, Yicheng Song, Junbo Guo, Yongdong Zhang 0001, Jintao Li 0001
WWW3
2009 Google challenge: incremental-learning for web video categorization on robust semantic feature space
abstract
With the advent of video sharing websites, the amount of videos on the internet grows rapidly. Web video categorization is an efficient methodology to organize the huge amount of data. In this paper, we propose an effective web video categorization algorithm for the large scale dataset. It includes two factors: 1) For the great diversity of web videos, we develop an effective semantic feature space called Concept Collection for Web Video Categorization (CCWV-CD) to represent web videos, which consists of concepts with small semantic gap and high distinguishing ability. Meanwhile, the online Wikipedia API is employed to diffuse the concept correlations in this space. 2) We propose an incremental support vector machine with fixed number of support vectors (n-ISVM) to fit the large scale incremental learning problem in web video categorization. Extensive experiments are conducted on the dataset of 80024 most representative videos on YouTube demonstrate that the semantic space with Wikipedia prorogation is more representative for web videos, and n-ISVM outperforms other algorithms in efficiency when performs the incremental learning.
Yicheng Song, Yongdong Zhang 0001, Xu Zhang 0028, Juan Cao 0001, Jintao Li 0001
ACM Multimedia1
2008 MAPLE: a Maintenance Approach for Pattern-enabLed rEconfiguration of SOA-based Enterprise Application
Songlin Hu 0001, Jiuming Tian, Yicheng Song
SEKE4