VLDB 2026 Research / reviewers in the wild / expert
Zhijie Shen
dblp:46/2026
· DBLP profile ↗
31ranked-venue papers
15as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 10 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Computer networks · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AT-Field: Rethinking the Games in Adversarial TrainingabstractAdversarial training is often modeled as a two-player zero-sum game, relying on strong assumptions that limit its practical guidance. In this paper, we instead analyze the interactions between training samples and show that even the fundamental objective—minimizing training loss—may not converge. To address this, we propose AT-Field, an adversarial training framework guided by sample-wise game-theoretic relationships. Specifically, we prove that training samples across different batches can form a none-potential game, where gradient descent induces cyclic behaviors, preventing convergence. By strategically searching and grouping these samples within the same batch, AT-Field transforms none-potential games into exact potential games, which are more effectively optimized using gradient-based methods. Experiments demonstrate that AT-Field integrates seamlessly with existing adversarial training techniques, enhancing both accuracy and robustness. Yixiao Xu, Mohan Li, Zhijie Shen, Yuan Liu 0002, Zhihong Tian 0001 |
AAAI | 3 |
| 2026 | Revisiting 360 Depth Estimation With PanoGabor: A New Fusion PerspectiveabstractDepth estimation from a monocular 360 image is important to the perception of the entire 3D environment. However, the inherent distortion and large field of view (FoV) in 360 images pose great challenges for this task. To this end, existing mainstream solutions typically introduce additional perspective-based 360 representations (e.g., Cubemap) to achieve effective feature extraction. Nevertheless, regardless of the introduced representations, they eventually need to be unified into the equirectangular projection (ERP) format for the subsequent depth estimation, which inevitably reintroduces additional distortions. In this work, we propose an oriented-distortion-aware Gabor Fusion framework (PGFuse) to address the above challenges. First, we introduce Gabor filters that analyze texture in the frequency domain, extending the receptive fields and enhancing depth cues. To address the reintroduced distortions, we design a latitude-aware distortion representation to generate customized, distortion-aware Gabor filters (PanoGabor filters). Furthermore, we design a channel-wise and spatial-wise unidirectional fusion module (CS-UFM) that integrates the proposed PanoGabor filters to unify other representations into the ERP format, delivering effective and distortion-aware features. Considering the orientation sensitivity of the Gabor transform, we further introduce a spherical gradient constraint to stabilize this sensitivity. Experimental results on three popular indoor 360 benchmarks demonstrate the superiority of the proposed PGFuse to existing state-of-the-art solutions. Code and models will be available at https://github.com/zhijieshen-bjtu/PGFuse. Zhijie Shen, Chunyu Lin, Lang Nie, Kang Liao, Weisi Lin, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Revisiting Monocular 3D Object Detection With Depth Thickness FieldabstractMonocular 3D object detection is challenging due to the lack of accurate depth. However, existing depth-assisted solutions still exhibit inferior performance, whose reason is universally acknowledged as the unsatisfactory accuracy of monocular depth estimation models. In this paper, we revisit monocular 3D object detection from the depth perspective and formulate an additional issue as the limited 3D structure-aware capability of existing depth representations (e.g., depth one-hot encoding or depth distribution). To address this issue, we introduce a novel Depth Thickness Field approach to embed clear 3D structures of the scenes. Specifically, we present MonoDTF, a scene-to-instance depth-adapted network comprising a Scene-Level Depth Retargeting (SDR) module and an Instance-Level Spatial Refinement (ISR) module. The former retargets traditional depth representations to the proposed depth thickness field, incorporating the scene-level perception of 3D structures. The latter refines the voxel space with the guidance of instances, enhancing the 3D instance-aware capability of the depth thickness field and thus improving detection accuracy. Extensive experiments on the KITTI and Waymo datasets demonstrate our superiority to existing state-of-the-art (SoTA) methods and the universality when equipped with different depth estimation models. The source codes are available at https://github.com/QiuDeZhang/MonoDTF. Qiude Zhang, Chunyu Lin, Zhijie Shen, Lang Nie, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | PanoKernel: Large Distortion-Aware Kernel for Panoramic Depth Perception
Zhiqiang Yan 0001, Zhijie Shen, Jiayi Yuan 0003, Xiang Li 0041, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | PACM: Position-Aware Cross-Modality Decoder for Handwritten Mathematical Expression Recognition
Zhijie Shen, Can Ma, Yaqiang Wu, Yu Zhou 0015 |
ICDAR (1) | 3 |
| 2025 | PerturbCTC: Improving Alignment in Scene Text Recognition with Feature Perturbation Based CTC
Zhijie Shen, Yaqiang Wu, Gangyan Zeng, Dongbao Yang, Yu Zhou 0015 |
ICDAR (4) | 3 |
| 2025 | Class-Agnostic Region-of-Interest Matching in Document Images
Demin Zhang, Jiahao Lyu 0002, Zhijie Shen, Yu Zhou 0015 |
ICDAR (4) | 3 |
| 2025 | Tree-Based Approach for Time-Independent Diffusion Network Inference
Weikai Jing, Chao Gao 0001, Kefeng Fan, Hailong Cheng, Zhijie Shen, Zhen Wang 0004 |
KSEM (1) | 6 |
| 2025 | A Consortium Blockchain-Based Edge Task Offloading Method for Connected Autonomous VehiclesabstractIn recent years, the proliferation of Connected Autonomous Vehicles (CAV) has revolutionized the transportation industry. However, these vehicles often face limitations in terms of local computing resources, leading to the need for offloading interactive-intensive application tasks to servers for processing. Traditional paradigm has its limitations in meeting the demands of massive task processing. The combination of Web3.0 and edge computing offers users high-reliable, low-latency, and highly flexible services. Nevertheless, the new paradigm also presents its own challenges such as ensuring privacy data protection, and reducing the time and energy costs associated with task offloading. To tackle these challenges, an edge task offloading framework based on consortium blockchain for CAVs has been developed. Within this framework, a consortium blockchain-based interaction-intensive task offloading method, called CBIToMe, has been designed. CBIToMe specifically addresses the multi-stage nature of interactive-intensive CAV tasks and aims to minimize task completion time and offloading costs, particularly when the waiting time for interaction is uncertain. Additionally, CBIToMe effectively utilizes consortium blockchain technology to safeguard the CAV privacy data. Results from experiments conducted in various scenarios demonstrate that CBIToMe outperforms three representative methods, showcasing its superior performance. Bowen Liu 0002, Hao Tian 0012, Zhijie Shen, Yueyue Xu, Wan-Chun Dou |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2025 | SGFormer: Spherical Geometry Transformer for 360° Depth EstimationabstractPanoramic distortion poses a significant challenge in 360° depth estimation, particularly pronounced at the north and south poles. Existing methods either adopt a bi-projection fusion strategy to remove distortions or model long-range dependencies to capture global structures, resulting in either unclear structure or insufficient local perception. In this paper, we propose a spherical geometry transformer, named SGFormer, to address the above issues, with an innovative step to integrate spherical geometric priors into vision transformers. To this end, we retarget the transformer decoder to a spherical prior decoder (termed SPDecoder), which endeavors to uphold the integrity of spherical structures during decoding. Concretely, we leverage bipolar reprojection, circular rotation, and curve local embedding to preserve the spherical characteristics of equidistortion, continuity, and surface distance, respectively. Furthermore, we present a query-based global conditional position embedding to compensate for spatial structure at varying resolutions. It not only boosts the global perception of spatial position but also sharpens the depth structure across different patches. Finally, we conduct extensive experiments on popular benchmarks, demonstrating our superiority over state-of-the-art solutions. Our code will be made publicly athttps://github.com/iuiuJaon/SGFormer. Junsong Zhang, Zisong Chen, Chunyu Lin, Zhijie Shen, Lang Nie, Kang Liao, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | An effective data placement strategy for IIoT applicationsabstractSummary With the rapid development of the Internet of things (IoT) and mobile communication technology, the amount of data related to industrial Internet of things (IIoT) applications has shown a trend of explosive growth, and hence edge‐cloud collaborative environment becomes one of the most popular paradigms to place the IIoT applications data. However, edge servers are often heterogeneous and capacity limited while having lower access delay, so there is a contradiction between capacity and latency while using edge storage. Additionally, when IIoT applications deployed crossing edge regions, the impact of data replication and data privacy should not be ignored. These factors often pose challenges to proposing an effective data placement strategy to take full advantage of edge storage. To address these challenges, an effective data placement strategy for IIoT applications is designed in this article. We first analyze the data access time and data placement cost in an edge‐cloud collaborative environment, with the consideration of data replication and data privacy. Then, we design a data placement strategy based on ‐constraint and Lagrangian relaxation, to reduce the data access time and meanwhile limit the data placement cost to an ideal level. As a result, our proposed data placement strategy can effectively reduce data access time and control data placement costs. Simulation and comparative analysis results have demonstrated the validity of our proposed strategy. Zhijie Shen, Bowen Liu 0002, Wan-Chun Dou |
Concurr. Comput. Pract. Exp. | 1 |
| 2024 | 360 Layout Estimation via Orthogonal Planes Disentanglement and Multi-View Geometric Consistency PerceptionabstractExisting panoramic layout estimation solutions tend to recover room boundaries from a vertically compressed sequence, yielding imprecise results as the compression process often muddles the semantics between various planes. Besides, these data-driven approaches impose an urgent demand for massive data annotations, which are laborious and time-consuming. For the first problem, we propose an orthogonal plane disentanglement network (termed DOPNet) to distinguish ambiguous semantics. DOPNet consists of three modules that are integrated to deliver distortion-free, semantics-clean, and detail-sharp disentangled representations, which benefit the subsequent layout recovery. For the second problem, we present an unsupervised adaptation technique tailored for horizon-depth and ratio representations. Concretely, we introduce an optimization strategy for decision-level layout analysis and a 1D cost volume construction method for feature-level multi-view aggregation, both of which are designed to fully exploit the geometric consistency across multiple perspectives. The optimizer provides a reliable set of pseudo-labels for network training, while the 1D cost volume enriches each view with comprehensive scene information derived from other perspectives. Extensive experiments demonstrate that our solution outperforms other SoTA models on both monocular layout estimation and multi-view layout estimation tasks. Zhijie Shen, Chunyu Lin, Junsong Zhang, Lang Nie, Kang Liao, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Disentangling Orthogonal Planes for Indoor Panoramic Room Layout Estimation with Cross-Scale Distortion AwarenessabstractBased on the Manhattan World assumption, most existing indoor layout estimation schemes focus on recovering layouts from vertically compressed 1D sequences. However, the compression procedure confuses the semantics of different planes, yielding inferior performance with ambiguous interpretability. To address this issue, we propose to disentangle this 1D representation by pre-segmenting orthogonal (vertical and horizontal) planes from a complex scene, explicitly capturing the geometric cues for indoor layout estimation. Considering the symmetry between the floor boundary and ceiling boundary, we also design a soft-flipping fusion strategy to assist the pre-segmentation. Besides, we present a feature assembling mechanism to effectively integrate shallow and deep features with distortion distribution awareness. To compensate for the potential errors in pre-segmentation, we further leverage triple attention to reconstruct the disentangled sequences for better performance. Experiments on four popular benchmarks demonstrate our superiority over existing SoTA solutions, especially on the 3DIoU metric. The code is available at https://github.com/zhijieshen-bjtu/DOPNet. Zhijie Shen, Zishuo Zheng, Chunyu Lin, Lang Nie, Kang Liao, Shuai Zheng 0005, Yao Zhao 0001 |
CVPR | 1 |
| 2023 | S-OmniMVS: Incorporating Sphere Geometry into Omnidirectional Stereo MatchingabstractMulti-fisheye stereo matching is a promising task that employs the traditional multi-view stereo (MVS) pipeline with spherical sweeping to acquire omnidirectional depth. However, the existing omnidirectional MVS technologies neglect fisheye and omnidirectional distortions, yielding inferior performance. In this paper, we revisit omnidirectional MVS by incorporating three sphere geometry priors: spherical projection, spherical continuity, and spherical position. To deal with fisheye distortion, we propose a new distortion-adaptive fusion module to convert fisheye inputs into distortion-free spherical tangent representations by constructing a spherical projection space. Then these multi-scale features are adaptively aggregated with additional learnable offsets to enhance content perception. To handle omnidirectional distortion, we present a new spherical cost aggregation module with a comprehensive consideration of the spherical continuity and position. Concretely, we first design a rotation continuity compensation mechanism to ensure omnidirectional depth consistency of left-right boundaries without introducing extra computation. On the other hand, we encode the geometry-aware spherical position and push them into the cost aggregation to relieve panoramic distortion and perceive the 3D structure. Furthermore, to avoid the excessive concentration of depth hypothesis caused by inverse depth linear sampling, we develop a segmented sampling strategy that combines linear and exponential spaces to create S-OmniMVS, along with three sphere priors. Extensive experiments demonstrate the proposed method outperforms the state-of-the-art (SoTA) solutions by a large margin on various datasets both quantitatively and qualitatively. Zisong Chen, Chunyu Lin, Lang Nie, Zhijie Shen, Kang Liao, Yuanzhouhan Cao, Yao Zhao 0001 |
ACM Multimedia | 4 |
| 2023 | Complementary Bi-directional Feature Compression for Indoor 360° Semantic Segmentation with Self-distillationabstractSemantic segmentation on 360° images is a vital component of scene understanding due to the rich surrounding information. Recently, horizontal representation-based approaches outperform projection-based solutions, because the distortions can be effectively removed by compressing the spherical data in the vertical direction. However, these methods ignore the distortion distribution prior and are limited to unbalanced receptive fields, e.g., the receptive fields are sufficient in the vertical direction and insufficient in the horizontal direction. Differently, a vertical representation compressed in another direction can offer implicit distortion prior and enlarge horizontal receptive fields. In this paper, we combine the two different representations and propose a novel 360° semantic segmentation solution from a complementary perspective. Our network comprises three modules: a feature extraction module, a bi-directional compression module, and an ensemble decoding module. First, we extract multi-scale features from a panorama. Then, a bi-directional compression module is designed to compress features into two complementary low-dimensional representations, which provide content perception and distortion prior. Furthermore, to facilitate the fusion of bi-directional features, we design a unique self distillation strategy in the ensemble decoding module to enhance the interaction of different features and further improve the performance. Experimental results show that our approach outperforms the state-of-the-art solutions on quantitative evaluations while displaying the best performance on visual appearance. Zishuo Zheng, Chunyu Lin, Lang Nie, Kang Liao, Zhijie Shen, Yao Zhao 0001 |
WACV | 5 |
| 2022 | PanoFormer: Panorama Transformer for Indoor 360$^{\circ }$ Depth Estimation
Zhijie Shen, Chunyu Lin, Kang Liao, Lang Nie, Zishuo Zheng, Yao Zhao 0001 |
ECCV (1) | 1 |
| 2022 | An Improved Deliberation Network with Text Pre-training for Code-Switching Automatic Speech Recognition
Zhijie Shen, Wu Guo |
INTERSPEECH | 1 |
| 2022 | Neural Contourlet Network for Monocular 360° Depth EstimationabstractFor a monocular 360° image, depth estimation is a challenging because the distortion increases along the latitude. To perceive the distortion, existing methods devote to designing a deep and complex network architecture. In this paper, we provide a new perspective that constructs an interpretable and sparse representation for a 360° image. Considering the importance of the geometric structure in depth estimation, we utilize the contourlet transform to capture an explicit geometric cue in the spectral domain and integrate it with an implicit cue in the spatial domain. Specifically, we propose a neural contourlet network consisting of a convolutional neural network and a contourlet transform branch. In the encoder stage, we design a spatial–spectral fusion module to effectively fuse two types of cues. Contrary to the encoder, we employ the inverse contourlet transform with learned low-pass subbands and band-pass directional subbands to compose the depth in the decoder. Experiments on the three popular 360° panoramic image datasets demonstrate that the proposed approach outperforms the state-of-the-art schemes with faster convergence. Code is available athttps://github.com/zhijieshen-bjtu/Neural-Contourlet-Network-for-MODE. Zhijie Shen, Chunyu Lin, Lang Nie, Kang Liao, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Distortion-Tolerant Monocular Depth Estimation on Omnidirectional Images Using Dual-CubemapabstractEstimating the depth of omnidirectional images is more challenging than that of normal field-of-view (NFoV) images because the varying distortion can significantly twist an object’s shape. The existing methods suffer from troublesome distortion while estimating the depth of omnidirectional images, leading to inferior performance. To reduce the negative impact of the distortion influence, we propose a distortion-tolerant omnidirectional depth estimation algorithm using a dual-cubemap. It comprises two modules: Dual-Cubemap Depth Estimation (DCDE) module and Boundary Revision (BR) module. In DCDE module, we present a rotation-based dual-cubemap model to estimate the accurate NFoV depth, reducing the distortion at the cost of boundary discontinuity on omnidirectional depths. Then a boundary revision module is designed to smooth the discontinuous boundaries, which contributes to the precise and visually continuous omnidirectional depths. Extensive experiments demonstrate the superiority of our method over other state-of-the-art solutions. Zhijie Shen, Chunyu Lin, Lang Nie, Kang Liao, Yao Zhao 0001 |
ICME | 1 |
| 2014 | Spatial-Temporal Tag Mining for Automatic Geospatial Video AnnotationabstractVideos are increasingly geotagged and used in practical and powerful GIS applications. However, video search and management operations are typically supported by manual textual annotations, which are subjective and laborious. Therefore, research has been conducted to automate or semi-automate this process. Since a diverse vocabulary for video annotations is of paramount importance towards good search results, this article proposes to leverage crowdsourced data from social multimedia applications that host tags of diverse semantics to build a spatio-temporal tag repository, consequently acting as input to our auto-annotation approach. In particular, to build the tag store, we retrieve the necessary data from several social multimedia applications, mine both the spatial and temporal features of the tags, and then refine and index them accordingly. To better integrate the tag repository, we extend our previous approach by leveraging the temporal characteristics of videos as well. Moreover, we set up additional ranking criteria on the basis of tag similarity, popularity and location bias. Experimental results demonstrate that, by making use of such a tag repository, the generated tags have a wide range of semantics, and the resulting rankings are more consistent with human perception. Yifang Yin, Zhijie Shen, Roger Zimmermann |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | Orientation data correction with georeferenced mobile videosabstractSimilar to positioning data, camera orientation information has become a powerful contextual feature utilized by a number of GIS and social media applications. Such auxiliary information facilitates higher-level semantic analysis and management of video assets in such applications, e.g., video summarization and video indexing systems. However, it is problematic that raw sensor data collected from current mobile devices is often not accurate enough for subsequent geospatial analysis. To date, an effective orientation data correction system for mobile video content has been lacking. Here we present a content-based approach that improves the accuracy of noisy orientation sensor measurements generated by mobile devices in conjunction with video acquisition. Our preliminary experimental results demonstrate significant accuracy enhancements which benefit upstream sensor-aided GIS applications to access video content more precisely. Guanfeng Wang, Yifang Yin, Beomjoo Seo, Roger Zimmermann, Zhijie Shen |
SIGSPATIAL/GIS | 5 |
| 2013 | OSCOR: an orientation sensor data correction system for mobile generated contentsabstractIn addition to positioning data, other sensor information -- such as orientation data, have become a useful and powerful contextual feature. Such auxiliary information can facilitate higher-level semantic description inferences in many multimedia applications, e.g., video tagging and video summarization. However, sensor data collected from current mobile devices is often not accurate enough for upstream multimedia analysis. An effective orientation data correction system for mobile multimedia content has been an elusive goal so far. Here we present a system, termed Oscor, which aims to improve the accuracy of noisy orientation sensor measurements generated by mobile devices during image and video recording. We provide a user-friendly camera interface to facilitate the gathering of additional information, which enables the correction process on the server-side. Geographic field-of-view (FOV) visualizations based on the original and corrected sensor data help users understand the corrected contextual information and how the erroneous data possibly may affect further processes. Guanfeng Wang, Beomjoo Seo, Yifang Yin, Roger Zimmermann, Zhijie Shen |
ACM Multimedia | 5 |
| 2012 | Reducing cross-group traffic with a cooperative streaming architectureabstractCooperative approaches, such as P2P networks, have demonstrated their effectiveness in video delivery. However, with underlay structure considered, it is still possible to further improve traffic efficiency. In this paper, we discuss the problem of localizing the traffic traversal across peer groups, which are partitioned according to underlay characteristics. We first provide three concrete examples to demonstrate this common challenge, which we theoretically formulate afterwards. Finally, we propose a ring overlay approach, which performs excellently to solve the problem, while tolerating peer dynamics and supporting peer heterogeneity. Zhijie Shen, Roger Zimmermann |
ACM Multimedia | 1 |
| 2012 | Automatic music soundtrack generation for outdoor videos from contextual sensor informationabstractWe present a system to automatically generate soundtracks for user-generated outdoor videos (UGV) based on concurrently captured contextual sensor information with mobile apps for the ACM Multimedia 2012 Google challenge: Automatic Music Video Generation. Our method addresses the use case of making "a video much more attractive for sharing by adding a matching soundtrack to it." Our system correlates viewable scene information from sensors with geographic contextual tags from OpenStreetMap. The co-occurance of geo-tags and mood tags are investigated from a set of categories of the web site Foursquare.com and a mapping from geo-tags to mood tags is obtained. Finally, a music retrieval component returns music based on matching mood tags. The experimental results show that our system can successfully create soundtracks that are related to the mood and situation of UGVs and therefore enhance the enjoyment of viewers. Our system sends only sensor data to a cloud service and is therefore bandwidth efficient since video data does not need to be transmitted for analysis. Yi Yu 0001, Zhijie Shen, Roger Zimmermann |
ACM Multimedia | 2 |
| 2012 | ISP-friendly P2P live streaming: A roadmap to realizationabstractPeer-to-Peer (P2P) applications generate large amounts of Internet network traffic. The wide-reaching connectivity of P2P systems is creating resource inefficiencies for network providers. Recent studies have demonstrated that localizing cross-ISP (Internet service provider) traffic can mitigate this challenge. However, bandwidth sensitivity and display quality requirements complicate the ISP-friendly design for live streaming systems. To this date, although some prior techniques focusing on live streaming systems exist, the correlation between traffic localization and streaming quality guarantee has not been well explored. Additionally, the proposed solutions are often not easy to apply in practice. In our presented work, we demonstrate that the cross-ISP traffic of P2P live streaming systems can be significantly reduced with little impact on the streaming quality. First, we analytically investigate and quantify the tradeoff between traffic localization and streaming quality guarantee, determining the lower bound of the inter-AS (autonomous system) streaming rate below which streaming quality cannot be preserved. Based on the analysis, we further propose a practical ISP-friendly solution, termed IFPS , which requires only minor changes to the peer selection mechanism and can easily be integrated into both new and existing systems. Additionally, the significant opportunity for localizing traffic is underscored by our collected traces from PPLive, which also enabled us to derive realistic parameters to guide our simulations. The experimental results demonstrate that IFPS reduces cross-ISP traffic from 81% up to 98% while keeping streaming quality virtually unaffected. Zhijie Shen, Roger Zimmermann |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2011 | SRV-TaGS: An Automatic TAGging and Search System for Sensor-Rich Outdoor VideosabstractTagging facilitates video search in many social media and web applications. While manual tagging is time consuming, subjective and sometimes inaccurate, auto-tagging facilitated by content-based techniques is compute-intensive and challenging to apply across domains. We have developed a complementary system, named SRV-TAGS, to automatically generate tags for outdoor videos based on their geographic properties, to index the videos based on their generated tags and to provide textual search services. The system works with our geo-referenced video management web portal, enabling users to manage, search and watch videos. Zhijie Shen, Sakire Arslan Ay, Seon Ho Kim |
ACM Multimedia | 1 |
| 2011 | Automatic tag generation and ranking for sensor-rich outdoor videosabstractVideo tag annotations have become a useful and powerful feature to facilitate video search in many social media and web applications. The majority of tags assigned to videos are supplied by users - a task which is time consuming and may result in annotations that are subjective and lack precision. A number of studies have utilized content-based extraction techniques to automate tag generation. However, these methods are compute-intensive and challenging to apply across domains. Here, we describe a complementary approach for generating tags based on the geographic properties of videos. With today's sensor-equipped smartphones, the location and orientation of a camera can be continuously acquired in conjunction with the captured video stream. Our novel technique utilizes these sensor meta-data to automatically tag outdoor videos in a two step process. First, we model the viewable scenes of the video as geometric shapes by means of its accompanied sensor data and determine the geographic objects that are visible in the video by querying geo-information databases through the viewable scene descriptions. Subsequently we extract textual information about the visible objects to serve as tags. Second, we define six criteria to score the tag relevance and rank the obtained tags based on these scores. Then we associate the tags with the video and the accurately delimited segments of the video. To evaluate the proposed technique we implemented a prototype tag generator and conducted a user study. The results demonstrate significant benefits of our method in terms of automation and tag utility. Zhijie Shen, Sakire Arslan Ay, Seon Ho Kim, Roger Zimmermann |
ACM Multimedia | 1 |
| 2011 | LAN-awareness: improved P2P live streamingabstractThe popularity of P2P streaming systems has rapidly created extensive, far-reaching Internet traffic. Recent studies have demonstrated that localizing cross-ISP (Internet service provider) traffic can mitigate this challenge. Another trend shows that households own an increasing number of devices, which are sharing a LAN of 2 or more peers. To this date, however, no study has investigated the potential of localizing traffic within LANs. In our presented work, we propose the concept of LAN-awareness and introduce its threefold benefits: 1) reducing Internet streaming traffic, 2) lowering stream server workload, and 3) improving streaming quality. First we conduct a large-scale measurement on PPLive, confirming that a considerable number of peers (up to 21%) are connected to the LANs having 2 or more peers. Recognizing the opportunity of localizing traffic within LANs, we discuss the principles to construct a LAN-aware overlay and propose a heuristic. The results of our trace-driven simulations confirm the benefits outlined above. Zhijie Shen, Roger Zimmermann |
NOSSDAV | 1 |
| 2011 | Peer-to-Peer Media Streaming: Insights and New DevelopmentsabstractInternet media content delivery started to emerge roughly a decade ago, and it has subsequently had a major impact on network traffic and usage. Although traditional client-server systems were used initially for delivering media content, researchers and practitioners soon realized that peer-to-peer (P2P) systems, due to their self-scaling properties, had the potential to improve scalability compared with traditional client-server architectures. Consequently, various P2P media streaming systems have been deployed successfully, and corresponding theoretical investigations have been performed on such systems. The rapid developments in this field raise the need for up-to-date literature surveys to summarize them. In recent years, numerous technological discoveries have been achieved. The focus of this report is to survey and discuss these new findings, which include new technological developments, as well as new understandings of these developments and of the existing P2P streaming techniques, through both novel modeling methodologies and measurement-based studies. Zhijie Shen, Jun Luo 0001, Roger Zimmermann, Athanasios V. Vasilakos |
Proc. IEEE | 1 |
| 2009 | ISP-friendly peer selection in P2P networksabstractPeer-to-peer (P2P) multicast is a scalable solution adopted by many video streaming systems. However, a prevalence of P2P applications has caused heavy traffic on the Internet. While P2P streaming users enjoy high-quality online video, the Internet Service Providers (ISPs) may experience large expenditure because of cross-ISP traffic. To reduce this traffic, it has been suggested that ISPs cooperate with P2P systems by exposing information about the networking layer to guide the topology construction. In this paper, we propose a novel peer selection algorithm that leverages the ISPs' service to form a local clustering topology through gossip communication. Additionally, this algorithm is adaptive to prevent peers from being trapped in a local cluster where they cannot obtain an adequate streaming rate. Zhijie Shen, Roger Zimmermann |
ACM Multimedia | 1 |
| 2008 | Meta-Mesh Span Restoration and Increased Lightpath Routing in Sparse Network TopologiesabstractPrior work has shown that network capacity efficiency decreases significantly as a network's topology becomes sparse. Meta-mesh restoration was proposed in prior work as a means of increasing capacity efficiency in such networks, and showed significant improvements relative to span restoration. In the present work herein, we now study the effects of converting an existing (or planned) span-restorable network to one employing meta-mesh restoration. We develop and analyze two new ILP design models, extensions of the original meta-mesh restoration design models, to maximize the average per-demand increase in working lightpaths the network can serve and fully protect if converted to meta-mesh. John Doucette 0001, Zhijie Shen |
ICC | 2 |