Xiaodan Song

dblp:93/3688 · DBLP profile ↗
← Back
39ranked-venue papers
15as first author
14since 2021 · last 2026
0000-0002-8049-1828ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 9 first-author · 10 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 13 · 7 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorComputer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Dynamic Spatio-Temporal Compression Ratio Learning and Frequency-Aware Semantic Compression for Video Imaging
abstract
Snapshot compressive imaging (SCI) and video compressive sensing (VCS) typically use fixed, globally uniform compression ratios that ignore spatio-temporal heterogeneity. We present D-STCRL, a reinforcement-learned framework that unifies adaptive sensing and semantic transmission under an explicit rate-distortion-energy objective. The pipeline comprises: (i) a Ratio Generation Network predicts per-patch ratio maps via spatio-temporal attention and 3D frequency cues; (ii) a Programmable Sensing Model emulates pixel-wise variable exposure through differentiable binary gating under a global budget; and (iii) a Frequency-Aware Swin decoder with a low-rank prior restores temporally consistent frames. A multi-objective policy gradient couples the ratio policy with reconstruction and JSCC, yielding stable training. On the NFS benchmark, D-STCRL improves PSNR by$2-3 ~\text{dB}$over fixed-ratio SCI at the same sampling budget; under 10 dB AWGN it surpasses a CRL baseline by$0.6-1.0 ~\text{dB}$while reducing transmitted symbols by up to 15 %. These results unify content-adaptive sensing and efficient transmission for next-generation cameras. Our code, configs and reproducible pipelines will be released upon acceptance.
Haixiong Li, Dahua Gao, Xiaodan Song, Guangming Shi
DCC3
2026 Cross-Component Attention Network for In-Loop Filtering in Versatile Video Coding
abstract
In this paper, we propose a cross-component attention network (CCA-Net) for in-loop filtering to leverage the strengths of both separate and shared models for luma and chroma components while exploit their cross correlation. Fig. 1 shows the overall network. We adopt a structure akin to multi-task learning and introduce a cross-component attention (CCA) module to guide chroma filtering with luma by modeling the correlation as a linear combination of chroma and luma features with adaptive weights. Experimental results demonstrate that the proposed method achieves$\{0.09 \%, 4.03 \%, 3.37 \%\}$bitrate savings for$\{\mathrm{Y}, \mathrm{U}, \mathrm{V}\}$, compared with the shared model while maintaining similar complexity. Compared with the separate models, the proposed method only has$\{0.35 \%, 0.2 \%, 0.59 \%\}$performance loss for$\{\mathrm{Y}, \mathrm{U}, \mathrm{V}\}$, but significantly reduces the computation and parameters by 50 %.
Xiaodan Song, Fan Cai, Haixiong Li, Yuansheng Wu, Xuguang Zuo
DCC1
2026 PMCF: A Progressive Multi-Level Collaborative Framework for Face Forgery Detection
Hongning Li, Zengzhang Li, Haijie Du, Jiawei Zhang 0011, Xiaodan Song, Qingqi Pei
IEEE Signal Process. Lett.5
2026 CMANet: Context-Aware Mutual Attention Network for Referring Image Segmentation
Xiong Pan, Xuemei Xie, Jianxiu Yang, Xiaodan Song, Guangming Shi
IEEE Trans. Multim.4
2025 Affine Transformation-Based Generative Face Video Compression
abstract
In this paper, we propose a generative face video compression framework based on affine transformations to better represent large movements without parameter transmission. It mainly consists of an encoder and decoder, and our encoder is similar to the one in [1]. Intra frame are compressed by the existing encoder, while subsequent inter frames are compressed into compact inter frame features. In the decoder, feature alignment is first established to map the decoded intra frame and inter frame features into the same domain. The aligned features are then combined with the appearance features extracted by the appearance encoder from the intra frame and fed into the coarse-fine affine transform module to establish motion estimation and compensation. The coarse affine transform focuses on global motion, while the fine affine transform deals with local motion, such as lip motion. Finally, the transformed features are fed into the image generation module to obtain the final reconstruction results.
Xihua Lin, Xiaodan Song, Xuguang Zuo, Dahua Gao, Xuemei Xie, Guangming Shi
DCC2
2025 A Transmitter-Model Unaware Generative Image Compression Framework for Semantic Communication
abstract
Unlike traditional bit-level data transmission methods, semantic communication focuses on conveying the meaning behind the data. Though promising results have been achieved, existing end-to-end learning-based semantic communication frameworks often require a synchronization of deep models between the transmitter and the receiver. Such design leads to tens of thousands models to be stored at receiver since different manufactures may optimize their own models. To address this problem, we propose a novel model-unaware generative image compression framework for semantic communication. It features at employing human-understandable multi-modality representations as an intermediate layer to enhance information transmission efficiency and semantic consistency. Our framework introduces a mask-based rate-distortion optimization module, which effectively removes low-relevance information for image generation and reduces the bit rate while maintaining semantic consistency. Experimental results demonstrate that the framework can still reconstruct high-quality images at very low bit rates, showcasing its potential for applications in modern communication systems.
Rongcan Zheng, Xiaodan Song, Xuguang Zuo, Minxi Yang, Dahua Gao, Xuemei Xie
ICASSP2
2025 A LLM-guided hybrid Mamba-Transformer architecture for part-to-whole motion synthesis
Fuming Wang, Dahua Gao, Xunliang Huang, Xiaodan Song
Comput. Vis. Image Underst.5
2025 Prompt-Based Invertible Mapping Alignment for Unsupervised Domain Adaptation
abstract
Large pre-trained vision-language models (VLMs) like CLIP have shown great potential for solving the unsupervised domain adaptation (UDA) problem. Existing prompt learning for UDA based on the unsupervised-trained VLMs requires distribution alignment between source and target domains in the common space for both vision and language branches. However, it is difficult for rough cross-domain alignment to maintain the discriminative semantic structure of both domains. Besides, the coarse features with non-informative noises due to ignoring the pseudo-label noises may cause failures to concentrate on precise semantics alignment. In this work, we propose a Prompt-Based Invertible Mapping Alignment (PIMA) method to incorporate discriminative domain knowledge into prompt learning, which is featured with refined cross-domain alignment in two separate space with a well-kept structure. Specifically, we design an invertible neural network-based homeomorphism mapping, and then achieve distribution alignment through such invertible mapping for connecting source and target visual feature space, which can preserve the data semantic structure. For better semantic alignment in vision-language space, we develop cross-modal implicit contrastive learning module to regularize non-informative features, which aims to find the low-rankness of implicit representation space. We conducted extensive experiments on three benchmark datasets to prove the advantages of our proposed PIMA over state-of-the-art methods.
Xiaodan Song, Xuemei Xie
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Mosic: Multimodal Semantic Integrated Communication for Health Monitoring in Iot Scenarios
abstract
Monitoring multimodal signals provides a more comprehensive understanding of health conditions compared to singlemode monitoring. In the face of the significant volumes of multimodal signals, existing IoT health monitoring systems primarily focus on high-fidelity signal transmission by encoding multimodal signals separately. However, due to the lack of consideration for the downstream applications and correlation between multimodal signals, a portion of bandwidth resources is wasted on task-irrelevant information and intermodal redundancy. To address this issue, we propose the Multimodal Semantic Integration Communication (MoSIC) framework composed of three levels: At the sensor level, multiple wearable sensors collect and send different modal signals to a mobile terminal; at the mobile terminal level, the terminal employs deep source-channel joint encoding for the received multimodal signals, extracting single-modal embedded features using a backbone network, and obtaining cross-modal features through a feature fusion network with contrastive constrain; at the cloud level, a decoding network symmetric to the encoding network reconstructs the multimodal signals, which are then used for downstream applications such as human activity recognition. MoSIC focuses on semantically integrating multimodal signals for downstream applications, resulting in improved encoding and transmission efficiency. It also reduces the radio-frequency power consumption and bandwidth requirements.
Minxi Yang, Dahua Gao, Xiaodan Song, Guangming Shi
ICASSP5
2024 SG2SC: A Generative Semantic Communication Framework for Scene Understanding-Oriented Image Transmission
abstract
In recent years, semantic communication based on deep learning for source-channel joint encoding has garnered significant attention. It utilizes network models trained end-to-end to represent signals as embedding vectors and has demonstrated superior performance compared to traditional methods. However, due to the significant disparity between embedding vectors and human language, it can be challenging to succinctly capture abstract semantics such as scenes. In this paper, we introduce the Scene Graph-based Generative Semantic Communication (SG2SC) framework, built upon structured semantics and conditional generative models for image transmission. SG2SC aims to faithfully convey abstract semantics like scenes. It begins by detecting object categories, spatial attributes, and inter-category relationships in the image, representing scene semantics in a graph structure. Subsequently, it employs graph neural networks for scene graph encoding, decoding, and transmission, and finally utilizes a conditional diffusion model for semantic decoding. Benefiting from its concise graph structure semantics, SG2SC outperforms traditional method, semantic communication based on deep joint source-channel coding, and segmentation-based generative semantic communication in terms of noise resistance and encoding efficiency.
Minxi Yang, Dahua Gao, Feng Xie 0009, Xiaodan Song, Guangming Shi
ICASSP5
2024 GRA-Net: Group response attention for deep learning
Zhenyuan Wang, Xuemei Xie, Xiaodan Song, Jianxiu Yang
Neurocomputing3
2022 Token Dropping for Efficient BERT Pretraining
abstract
Le Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu, Xinying Song, Xiaodan Song, Denny Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Le Hou, Richard Yuanzhe Pang, Tianyi Zhou 0001, Yuexin Wu, Xinying Song, Xiaodan Song, Denny Zhou
ACL (1)6
2022 A Simple Single-Scale Vision Transformer for Object Detection and Instance Segmentation
Wuyang Chen 0001, Xianzhi Du, Lucas Beyer, Xiaohua Zhai, Tsung-Yi Lin, Huizhong Chen, Xiaodan Song, Zhangyang Wang, Denny Zhou
ECCV (10)9
2022 Auto-scaling Vision Transformers without Training
Wuyang Chen 0001, Wei Huang 0034, Xianzhi Du, Xiaodan Song, Zhangyang Wang, Denny Zhou
ICLR4
2020 MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
abstract
Natural Language Processing (NLP) has recently achieved great success by using huge pre-trained models with hundreds of millions of parameters.However, these models suffer from heavy model sizes and high latency such that they cannot be deployed to resourcelimited mobile devices.In this paper, we propose MobileBERT for compressing and accelerating the popular BERT model.Like the original BERT, MobileBERT is task-agnostic, that is, it can be generically applied to various downstream NLP tasks via simple fine-tuning.Basically, MobileBERT is a thin version of BERT LARGE , while equipped with bottleneck structures and a carefully designed balance between self-attentions and feed-forward networks.To train MobileBERT, we first train a specially designed teacher model, an invertedbottleneck incorporated BERT LARGE model.Then, we conduct knowledge transfer from this teacher to MobileBERT.Empirical studies show that MobileBERT is 4.3× smaller and 5.5× faster than BERT BASE while achieving competitive results on well-known benchmarks.On the natural language inference tasks of GLUE, MobileBERT achieves a GLUE score of 77.7 (0.6 lower than BERT BASE ), and 62 ms latency on a Pixel 4 phone.On the SQuAD v1.1/v2.0question answering task, MobileBERT achieves a dev F1 score of 90.0/79.2(1.5/2.1 higher than BERT BASE ).
Zhiqing Sun, Hongkun Yu 0001, Xiaodan Song, Yiming Yang 0002, Denny Zhou
ACL3
2020 SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization
abstract
Convolutional neural networks typically encode an input image into a series of intermediate features with decreasing resolutions. While this structure is suited to classification tasks, it does not perform well for tasks requiring simultaneous recognition and localization (e.g., object detection). The encoder-decoder architectures are proposed to resolve this by applying a decoder network onto a backbone model designed for classification tasks. In this paper, we argue encoder-decoder architecture is ineffective in generating strong multi-scale features because of the scale-decreased backbone. We propose SpineNet, a backbone with scale-permuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search. Using similar building blocks, SpineNet models outperform ResNet-FPN models by 3%+ AP at various scales while using 10-20% fewer FLOPs. In particular, SpineNet-190 achieves 52.1% AP on COCO, attaining the new state-of-the-art performance for single model object detection without test-time augmentation. SpineNet can transfer to classification tasks, achieving 5% top-1 accuracy improvement on a challenging iNaturalist fine-grained dataset. Code is at: https://github.com/tensorflow/tpu/tree/master/models/official/detection.
Xianzhi Du, Tsung-Yi Lin, Pengchong Jin, Golnaz Ghiasi, Mingxing Tan, Yin Cui, Quoc V. Le, Xiaodan Song
CVPR8
2020 Efficient Scale-Permuted Backbone with Learned Resource Distribution
Xianzhi Du, Tsung-Yi Lin, Pengchong Jin, Yin Cui, Mingxing Tan, Quoc V. Le, Xiaodan Song
ECCV (23)7
2020 BigNAS: Scaling up Neural Architecture Search with Big Single-Stage Models
Pengchong Jin, Hanxiao Liu, Gabriel Bender, Pieter-Jan Kindermans, Mingxing Tan, Thomas S. Huang, Xiaodan Song, Ruoming Pang, Quoc V. Le
ECCV (7)8
2020 Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Yang You 0001, Sashank J. Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, Cho-Jui Hsieh
ICLR7
2020 Go Wide, Then Narrow: Efficient Training of Deep Thin Networks
abstract
For deploying a deep learning model into production, it needs to be both accurate and compact to meet the latency and memory constraints. This usually results in a network that is deep (to ensure performance) and yet thin (to improve computational efficiency). In this paper, we propose an efficient method to train a deep thin network with a theoretic guarantee. Our method is motivated by model compression. It consists of three stages. First, we sufficiently widen the deep thin network and train it until convergence. Then, we use this well-trained deep wide network to warm up (or initialize) the original deep thin network. This is achieved by layerwise imitation, that is, forcing the thin network to mimic the intermediate outputs of the wide network from layer to layer. Finally, we further fine tune this already well-initialized deep thin network. The theoretical guarantee is established by using the neural mean field analysis. It demonstrates the advantage of our layerwise imitation approach over backpropagation. We also conduct large-scale empirical experiments to validate the proposed method. By training with our method, ResNet50 can outperform ResNet101, and BERT base can be comparable with BERT large, when ResNet101 and BERT large are trained under the standard training procedures as in the literature.
Denny Zhou, Mao Ye 0006, Tianjian Meng, Mingxing Tan, Xiaodan Song, Quoc V. Le, Qiang Liu 0001, Dale Schuurmans
ICML6
2018 A Practical Convolutional Neural Network as Loop Filter for Intra Frame
abstract
Loop filters are used in video coding to remove artifacts or improve performance. Recent advances in deploying convolutional neural network (CNN) to replace traditional loop filters show large gains but with problems for practical application. First, different model is used for frames encoded with different quantization parameter (QP), respectively. It is expensive for hardware. Second, float points operation in CNN leads to inconsistency between encoding and decoding across different platforms. Third, redundancy within CNN model consumes precious computational resources. This paper proposes a CNN as the loop filter for intra frames and proposes a scheme to solve the above problems. It aims to design a single CNN model with low redundancy to adapt to decoded frames with different qualities and ensure consistency. To adapt to reconstructions with different qualities, both reconstruction and QP are taken as inputs. After training, the obtained model is compressed to reduce redundancy. To ensure consistency, dynamic fixed points (DFP) are adopted in testing CNN. Parameters in the compressed model are first quantized to DFP and then used for inference of CNN. Outputs of each layer in CNN are computed by DFP operations. Experimental results on JEM 7.0 report 3.14%,5.21 %, 6.28% BD-rate savings for luma and two chroma components with all intra configuration when replacing all traditional filters.
Xiaodan Song, Jiabao Yao, Lulu Zhou, Xiaoyang Wu 0007, Di Xie, Shiliang Pu
ICIP1
2018 Unequal Error Protection for Scalable Video Storage in the Cloud
abstract
Redundancy is necessary for a storage system to achieve reliability. Frequent errors in large-scale storage systems, for example, cloud, make it desirable to reduce the cost of recovery. Among all types of data in cloud storage, videos generally occupy significant amounts of space due to high volumes and the rapid development of video sharing and video-on-demand services. Unlike general data, videos can tolerate a certain level of quality degradation. This paper investigates multilayer video representations, such as scalable videos and simulcast streaming, and proposes an unequal error protection scheme based on local reconstruction codes (LRC) for video storage. By providing less protection for less important layers or video copies, a better tradeoff between storage and repair cost is achieved. Both theoretical and simulation results show that such a tradeoff can be achieved over the LRC with equal error protection, though the recovered video quality might be slightly lower in rare cases.
Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001
IEEE Trans. Multim.1
2017 Distributed Compressive Sensing for Cloud-Based Wireless Image Transmission
abstract
We consider efficient image transmission via time-varying channels. To improve the performance, we propose a new distributed compressive sensing (CS) scheme that can leverage similar images in the cloud. It is featured by channel SNR and bandwidth scalability, high efficiency, and low encoding complexity. For each image, a compressed thumbnail is first transmitted after forward error correction (FEC) and modulation to retrieve similar images and generate a side information (SI) in the cloud. The residual image after subtracting the decompressed thumbnail is then coded and transmitted by CS through a very dense constellation without FEC. The linearly and ratelessly generated CS measurements make it capable of achieving both graceful quality degradation (GD) with the channel SNR and bandwidth scalability in a universal scheme. A mode decision and transform-domain power allocation are introduced for better bandwidth usage and protection against channel errors. At the decoder, a two-step CS decoding is performed to recover the residual signal, where both the local and nonlocal correlations within the image and that with the SI are exploited. Simulations on landmark images and an AWGN channel show that the received image quality gracefully increases with the channel SNR and bandwidth. Furthermore, it outperforms existing schemes both subjectively and objectively by up to 11 dB gains compared with the state-of-the-art transmission scheme with GD, i.e. SoftCast.
Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001
IEEE Trans. Multim.1
2015 Unequal error protection for scalable video storage in the cloud
abstract
Redundancy is necessary for a storage system to recover from errors. The frequent errors in large-scale systems, e.g. cloud, make it desired to reduce the recovery cost. Among all kinds of data stored in the cloud, video takes a large portion due to its large data volume. The other characteristic of video is that a certain distortion can be tolerated. This paper investigates using scalable video representation and unequal error protection scheme to reduce the storage and recovery costs in the cloud. By introducing more protection for the base layer and less on the enhancement layers, it can achieve a better tradeoff between storage and reconstruction costs although the reliability for the enhancement layer sacrifices a little. Simulation results based on local reconstruction codes (LRC) show that comparing with the existing (12, 2, 2) LRC code in Windows Azure Storage, the reconstruction cost can be reduced from 6x to 3x at the same storage cost at the expense of possible video quality loss.
Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001
ICME1
2015 Compressive sensing based image transmission with side information at the decoder
abstract
This paper proposes a distributed compressive sensing (CS) scheme for robust image transmission over unknown or time-varying channels with highly correlated images at the decoder. A compressed thumbnail is first transmitted after digital forward error correction (FEC) and modulation to retrieve highly correlated images and generate a side information (SI) at the decoder. The current residual image after subtracting the decompressed thumbnail is then coded and transmitted by CS through a very dense constellation without FEC. The linear representation of the residual signal by CS measurements and rateless sampling makes it able to achieve graceful degradation and bandwidth scalability without channel feedback. Moreover, a transform-domain power allocation is employed before random sampling to protect against channel errors. At the decoder, both the nonlocal correlations within the original image and the correlation with the SI are exploited in CS decoding via a low-rank regulation on similar patches. After CS decoding, a block-wise minimum-mean-square-error (MMSE) reconstruction using the SI is further performed in the spatial domain to enhance the reconstruction quality. Simulations on landmark images and an unknown Gaussian channel show that an up to 10 dB gain is achieved at low channel SNRs compared with the state-of-the-art uncoded image transmission scheme, i.e. SoftCast, when highly correlated images are available at the decoder.
Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001
VCIP1
2015 Cloud-Based Distributed Image Coding
abstract
With multimedia flourishing on the Web, it is easy to find similar images for a query, especially landmark images. Traditional image coding, such as JPEG, cannot exploit correlations with external images. Existing vision-based approaches are able to exploit such correlations by reconstructing from local descriptors but cannot ensure the pixel-level fidelity of the reconstruction. In this paper, a cloud-based distributed image coding (Cloud-DIC) scheme is proposed to exploit external correlations for mobile photo uploading. For each input image, a thumbnail is transmitted to retrieve correlated images and reconstruct it in the cloud by geometrical and illumination registrations. Such a reconstruction serves as the side information (SI) in the Cloud-DIC. The image is then compressed by a transform-domain syndrome coding to correct the disparity between the original image and the SI. Once a bitplane is received in the cloud, an iterative refinement process is performed between the final reconstruction and the SI. Moreover, a joint encoder/decoder mode decision at block, frequency, and bitplane levels is proposed to adapt to different correlations. Experimental results on a landmark image database show that the Cloud-DIC can largely enhance the coding efficiency both subjectively and objectively, with up to 5-dB gains and 70% bits saving over JPEG with arithmetic coding, and perform comparably at low bitrates with the intra coding of the High Efficiency Video Coding standard with a much lower encoder complexity.
Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2014 Cloud-based distributed image coding
abstract
This paper proposes a cloud-based distributed image coding scheme (Cloud-DIC) to exploit the strong correlations with external partial-duplicate images in the cloud. It features both high coding efficiency and low encoder complexity, which makes it suitable for photo sharing on mobile devices. To get the side information in the cloud, a thumbnail of the current image is transmitted to retrieve highly correlated images and reconstruct through geometrical registration and adaptive patched-based stitching. The current image is then compressed by a transform-domain syndrome coding, bitplane by bitplane. Once a bitplane is received, the decoded high-quality image is further used to refine the side information in the cloud, which will benefit the coding of following bitplanes and the reconstruction. Experimental results on a landmark image database show that it can largely enhance the coding efficiency both subjectively and objectively with up to 5 dB gains and 58% bits saving over JPEG.
Xiaodan Song, Xiulian Peng, Jizheng Xu, Feng Wu 0001
ICIP1
2009 On evolutionary spectral clustering
abstract
Evolutionary clustering is an emerging research area essential to important applications such as clustering dynamic Web and blog contents and clustering data streams. In evolutionary clustering, a good clustering result should fit the current data well, while simultaneously not deviate too dramatically from the recent history. To fulfill this dual purpose, a measure of temporal smoothness is integrated in the overall measure of clustering quality. In this article, we propose two frameworks that incorporate temporal smoothness in evolutionary spectral clustering. For both frameworks, we start with intuitions gained from the well-known k -means clustering problem, and then propose and solve corresponding cost functions for the evolutionary spectral clustering problems. Our solutions to the evolutionary spectral clustering problems provide more stable and consistent clustering results that are less sensitive to short-term noises while at the same time are adaptive to long-term cluster drifts. Furthermore, we demonstrate that our methods provide the optimal solutions to the relaxed versions of the corresponding evolutionary k -means clustering problems. Performance experiments over a number of real and synthetic data sets illustrate our evolutionary spectral clustering methods provide more robust clustering results that are not sensitive to noise and can adapt to data drifts.
Yun Chi, Xiaodan Song, Dengyong Zhou, Koji Hino, Belle L. Tseng
ACM Trans. Knowl. Discov. Data2
2008 Learning multiple graphs for document recommendations
abstract
The Web offers rich relational data with different semantics. In this paper, we address the problem of document recommendation in a digital library, where the documents in question are networked by citations and are associated with other entities by various relations. Due to the sparsity of a single graph and noise in graph construction, we propose a new method for combining multiple graphs to measure document similarities, where different factorization strategies are used based on the nature of different graphs. In particular, the new method seeks a single low-dimensional embedding of documents that captures their relative similarities in a latent space. Based on the obtained embedding, a new recommendation framework is developed using semi-supervised learning on graphs. In addition, we address the scalability issue and propose an incremental algorithm. The new incremental method significantly improves the efficiency by calculating the embedding for new incoming documents only. The new batch and incremental methods are evaluated on two real world datasets prepared from CiteSeer. Experiments demonstrate significant quality improvement for our batch method and significant efficiency improvement with tolerable quality loss for our incremental method.
Shenghuo Zhu, Kai Yu 0001, Xiaodan Song, Belle L. Tseng, Hongyuan Zha, C. Lee Giles
WWW4
2007 Identifying opinion leaders in the blogosphere
abstract
Opinion leaders are those who bring in new information, ideas, and opinions, then disseminate them down to the masses, and thus influence the opinions and decisions of others by a fashion of word of mouth. Opinion leaders capture the most representative opinions in the social network, and consequently are important for understanding the massive and complex blogosphere. In this paper, we propose a novel algorithm called InfluenceRank to identify opinion leaders in the blogosphere. The InfluenceRank algorithm ranks blogs according to not only how important they are as compared to other blogs, but also how novel the information they can contribute to the network. Experimental results indicate that our proposed algorithm is effective in identifying influential opinion leaders.
Xiaodan Song, Yun Chi, Koji Hino, Belle L. Tseng
CIKM1
2007 Summarization System by Identifying Influential Blogs
Xiaodan Song, Yun Chi, Koji Hino, Belle L. Tseng
ICWSM1
2007 Evolutionary spectral clustering by incorporating temporal smoothness
abstract
Evolutionary clustering is an emerging research area essential to important applications such as clustering dynamic Web and blog contents and clustering data streams. In evolutionary clustering, a good clustering result should fit the current data well, while simultaneously not deviate too dramatically from the recent history. To fulfill this dual purpose, a measure of temporal smoothness is integrated in the overall measure of clustering quality. In this paper, we propose two frameworks that incorporate temporal smoothness in evolutionary spectral clustering. For both frameworks, we start with intuitions gained from the well-known k-means clustering problem, and then propose and solve corresponding cost functions for the evolutionary spectral clustering problems. Our solutions to the evolutionary spectral clustering problems provide more stable and consistent clustering results that are less sensitive to short-term noises while at the same time are adaptive to long-term cluster drifts. Furthermore, we demonstrate that our methods provide the optimal solutions to the relaxed versions of the corresponding evolutionary k-means clustering problems. Performance experiments over a number of real and synthetic data sets illustrate our evolutionary spectral clustering methods provide more robust clustering results that are not sensitive to noise and can adapt to data drifts.
Yun Chi, Xiaodan Song, Dengyong Zhou, Koji Hino, Belle L. Tseng
KDD2
2007 Structural and temporal analysis of the blogosphere through community factorization
abstract
The blogosphere has unique structural and temporal properties since blogs are typically used as communication media among human individuals. In this paper, we propose a novel technique that captures the structure and temporal dynamics of blog communities. In our framework, a community is a set of blogs that communicate with each other triggered by some events (such as a news article). The community is represented by its structure and temporal dynamics: a community graph indicates how often one blog communicates with another, and a community intensity indicates the activity level of the community that varies over time. Our method, community factorization, extracts such communities from the blogosphere, where the communication among blogs is observed as a set of subgraphs (i.e., threads of discussion). This community extraction is formulated as a factorization problem in the framework of constrained optimization, in which the objective is to best explain the observed interactions in the blogosphere over time. We further provide a scalable algorithm for computing solutions to the constrained optimization problems. Extensive experimental studies on both synthetic and real blog data demonstrate that our technique is able to discover meaningful communities that are not detectable by traditional methods.
Yun Chi, Shenghuo Zhu, Xiaodan Song, Jun'ichi Tatemura, Belle L. Tseng
KDD3
2007 Information flow modeling based on diffusion rate for prediction and ranking
abstract
Information flows in a network where individuals influence each other. The diffusion rate captures how efficiently the information can diffuse among the users in the network. We propose an information flow model that leverages diffusion rates for: (1) prediction . identify where information should flow to, and (2) ranking . identify who will most quickly receive the information. For prediction, we measure how likely information will propagate from a specific sender to a specific receiver during a certain time period. Accordingly a rate-based recommendation algorithm is proposed that predicts who will most likely receive the information during a limited time period. For ranking, we estimate the expected time for information diffusion to reach a specific user in a network. Subsequently, a DiffusionRank algorithm is proposed that ranks users based on how quickly information will flow to them. Experiments on two datasets demonstrate the effectiveness of the proposed algorithms to both improve the recommendation performance and rank users by the efficiency of information flow.
Xiaodan Song, Yun Chi, Koji Hino, Belle L. Tseng
WWW1
2006 Modeling Evolutionary Behaviors for Community-based Dynamic Recommendation
abstract
We exploit dynamic patterns from both documents' and users' aspects to build models for recommendation. We propose a Community-Based Dynamic Recommendation (CBDR) scheme to make recommendations by taking content semantics, evolutionary patterns, and user communities into consideration. A Time-Sensitive Adaboost algorithm is proposed to build adaptive user models for ranking document candidates based on leveraging dynamic factors such as freshness, popularity, and other attributes. Our experimental results on a large online application system demonstrate the recommendation usefulness of the CBDR scheme is 259% better than the collaborative filtering, 126% better than the community-based static recommendation algorithm, and 106% better than the optimal global recommendation bound.
Xiaodan Song, Ching-Yung Lin, Belle L. Tseng, Ming-Ting Sun
SDM1
2006 Personalized recommendation driven by information flow
abstract
We propose that the information access behavior of a group of people can be modeled as an information flow issue, in which people intentionally or unintentionally influence and inspire each other, thus creating an interest in retrieving or getting a specific kind of information or product. Information flow models how information is propagated in a social network. It can be a real social network where interactions between people reside; it can be, moreover, a virtual social network in that people only influence each other unintentionally, for instance, through collaborative filtering. We leverage users' access patterns to model information flow and generate effective personalized recommendations. First, an early adoption based information flow (EABIF) network describes the influential relationships between people. Second, based on the fact that adoption is typically category specific, we propose a topic-sensitive EABIF (TEABIF) network, in which access patterns are clustered with respect to the categories. Once an item has been accessed by early adopters, personalized recommendations are achieved by estimating whom the information will be propagated to with high probabilities. In our experiments with an online document recommendation system, the results demonstrate that the EABIF and the TEABIF can respectively achieve an improved (precision, recall) of (91.0%, 87.1%) and (108.5%, 112.8%) compared to traditional collaborative filtering, given an early adopter exists.
Xiaodan Song, Belle L. Tseng, Ching-Yung Lin, Ming-Ting Sun
SIGIR1
2005 Speech-Based Visual Concept Learning Using Wordnet
abstract
Modeling visual concepts using supervised or unsupervised machine learning approaches are becoming increasing important for video semantic indexing, retrieval, and filtering applications. Naturally, videos include multimodality data such as audio, speech, visual and text, which are combined to infer therein the overall semantic concepts. However, in the literature, most researches were conducted within only one single domain. In this paper we propose an unsupervised technique that builds context-independent keyword lists for desired visual concept modeling using WordNet. Furthermore, we propose an Extended Speech-based Visual Concept (ESVC) model to reorder and extend the above keyword lists by supervised learning based on multimodality annotation. Experimental results show that the context-independent models can achieve comparable performance compared to conventional supervised learning algorithms, and the ESVC model achieves about 53% and 28.4% improvement in two testing subsets of the TRECVID 2003 corpus over a state-of-the-art speech-based video concept detection algorithm.
Xiaodan Song, Ching-Yung Lin, Ming-Ting Sun
ICME1
2005 Modeling and predicting personal information dissemination behavior
abstract
In this paper, we propose a new way to automatically model and predict human behavior of receiving and disseminating information by analyzing the contact and content of personal communications. A personal profile, called CommunityNet, is established for each individual based on a novel algorithm incorporating contact, content, and time information simultaneously. It can be used for personal social capital management. Clusters of CommunityNets provide a view of informal networks for organization management. Our new algorithm is developed based on the combination of dynamic algorithms in the social network field and the semantic content classification methods in the natural language processing and machine learning literatures. We tested CommunityNets on the Enron Email corpus and report experimental results including filtering, prediction, and recommendation capabilities. We show that the personal behavior and intention are somewhat predictable based on these models. For instance, "to whom a person is going to send a specific email" can be predicted by one's personal social network and content analysis. Experimental results show the prediction accuracy of the proposed adaptive algorithm is 58% better than the social network-based predictions, and is 75% better than an aggregated model based on Latent Dirichlet Allocation with social network enhancement. Two online demo systems we developed that allow interactive exploration of CommunityNet are also discussed.
Xiaodan Song, Ching-Yung Lin, Belle L. Tseng, Ming-Ting Sun
KDD1
2002 Unsupervised Mumford-Shah energy based hybrid of texture and nontexture image segmentation
abstract
In this paper, we extend Mumford-Shah energy, the most general model for image segmentation on the assumption that all visible surface patches have intensity which is slowly varying plus noise, to unsupervised texture segmentation by considering multiple channels including both colors and Gabor wavelets representations. The Mumford-Shah energy in the paper thus formalizes a tradeoff between region homogeneity in color and texture field and edge compactness. Numerical results are presented for synthesized and real images to reveal the validity of the proposed algorithm.
Fei Liu 0017, Xiaodan Song, Yupin Luo, Dongcheng Hu
ICIP (2)2