Jiajie Tan

dblp:185/4058 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Vector Biomedical Dense Retrieval with Knowledge-Enhanced Entity-Type Clustering
abstract
Single-vector dense retrieval models, which are foundational to modern Retrieval-Augmented Generation (RAG) systems, struggle to represent the multifaceted semantics of complex documents, particularly in specialized fields like biomedicine. This semantic bottleneck limits their ability to provide comprehensive context for generation tasks. To address this, we propose ELK-Multi, a novel multi-vector retrieval framework that constructs fine-grained document representations through knowledge-enhanced entity-type clustering. By leveraging a knowledge-aware encoder, ELK-Multi first identifies and groups entities by their type, generating a distinct vector for each semantic cluster. These targeted representations are then combined with a global document vector using principled aggregation strategies to balance fine-grained detail with holistic context. Extensive experiments on the TREC-COVID and NFCorpus datasets validate our approach, where ELK-Multi establishes new state-of-the-art results in NDCG and Recall. This is complemented by a detailed efficiency analysis demonstrating that our model achieves this performance while remaining within the efficient dual-encoder paradigm, alongside a qualitative analysis with case studies and visualizations.
Jiajie Tan, Xiaorou Zheng, Shoubin Dong
ACM Trans. Knowl. Discov. Data2
2025 Knowledge enhanced edge-driven graph neural ranking for biomedical information retrieval
Xiaofeng Liu 0014, Jiajie Tan, Shoubin Dong
Expert Syst. Appl.2
2023 Semi-supervised Learning with Network Embedding on Ambient RF Signals for Geofencing Services
abstract
In applications such as elderly care, dementia anti-wandering and pandemic control, it is important to ensure that people are within a predefined area for their safety and well-being. We propose GEM, a practical, semi-supervised Geofencing system with network EMbedding, which is based only on ambient radio frequency (RF) signals. GEM models measured RF signal records as a weighted bipartite graph. With access points on one side and signal records on the other, it is able to precisely capture the relationships between signal records. GEM then learns node embeddings from the graph via a novel bipartite network embedding algorithm called BiSAGE, based on a Bipartite graph neural network with a novel bi-level SAmple and aggreGatE mechanism and non-uniform neighborhood sampling. Using the learned embeddings, GEM finally builds a one-class classification model via an enhanced histogram-based algorithm for in-out detection, i.e., to detect whether the user is inside the area or not. This model also keeps on improving with newly collected signal records. We demonstrate through extensive experiments in diverse environments that GEM shows state-of-the-art performance with up to 34% improvement in F-score. BiSAGE in GEM leads to a 54% improvement in F-score, as compared to the one without BiSAGE.
Weipeng Zhuo, Ka Ho Chiu, Jierun Chen, Jiajie Tan, Edmund Sumpena, Shueng-Han Gary Chan, Sangtae Ha, Chul-Ho Lee
ICDE4
2023 Self-Supervised Association of Wi-Fi Probe Requests Under MAC Address Randomization
abstract
Wi-Fi-on devices such as smartphones search for network availability by periodically broadcasting probe requests which encapsulate MAC addresses as device identifiers. To protect identity privacy, modern devices embed random MAC addresses in probe frames, the so-called MAC address randomization. Such randomization disrupts the frame association, inadvertently frustrating identity-oblivious statistical analytic efforts such as people counting and trajectory inference. To address that, we proposeCappuccino, a novel privacy-preserving approach that captures the association of probe requests under MAC address randomization. Cappuccino first estimates pairwise frame correlation and then associates frames over time. For frame correlation, it employs a self-supervised estimator that jointly considers multiple modalities, i.e., information elements, sequence number, and received signal strength. For multiple frame association, Cappuccino formulates frames as a minimum-cost flow optimization. To the best of our knowledge, this is the first piece of work that leverages self-supervised learning to estimate frame correlation based on multiple modalities and formulates the probe request association problem as the network flow optimization. We have conducted extensive experiments in a leading and crowded shopping mall for more than three months. Cappuccino achieves remarkable performance in terms of V-measure scores ($>0.85$).
Tianlang He, Jiajie Tan, Shueng-Han Gary Chan
IEEE Trans. Mob. Comput.2
2023 Implicit Multimodal Crowdsourcing for Joint RF and Geomagnetic Fingerprinting
abstract
In fingerprint-based indoor localization, fusing radio frequency (RF) and geomagnetic signals has been shown to achieve promising results. To efficiently collect fingerprints, implicit crowdsourcing can be used, where signals sampled by pedestrians are automatically labeled with their locations on a map. Previous work on crowdsourced fingerprinting is often based on a single signal, which is susceptible to signal bias and labeling error. We study, for the first time, implicit multimodal crowdsourcing for joint RF and geomagnetic fingerprinting. The scheme, termed UbiFin, exploits the spatial correlation among RF, geomagnetic, and motion signals to mitigate the impact of sensor noise, leading to highly accurate and robust fingerprinting without the need for any explicit manual intervention. Using clustering and dynamic programming, UbiFin correlates spatially different signals and filters effectively mislabeled signals. We conduct extensive experiments on our campus and a large multi-story shopping mall. Efficient and simple to implement, UbiFin outperforms other state-of-the-art crowdsourcing schemes to construct RF and geomagnetic fingerprints in terms of accuracy and robustness (cutting fingerprint error by 40 percent in general).
Jiajie Tan, Ka-Ho Chow 0001, Shueng-Han Gary Chan
IEEE Trans. Mob. Comput.1
2022 Tackling Multipath and Biased Training Data for IMU-Assisted BLE Proximity Detection
abstract
Proximity detection is to determine whether an IoT receiver is within a certain distance from a signal transmitter. Due to its low cost and high popularity, Bluetooth low energy (BLE) has been used to detect proximity based on the received signal strength indicator (RSSI). To address the fact that RSSI can be markedly influenced by device carriage states, previous works have incorporated RSSI with inertial measurement unit (IMU) using deep learning. However, they have not sufficiently accounted for the impact of multipath. Furthermore, due to the special setup, the IMU data collected in the training process may be biased, which hampers the system’s robustness and generalizability. This issue has not been studied before.We propose PRID, an IMU-assisted BLE proximity detection approach robust against RSSI fluctuation and IMU data bias. PRID histogramizes RSSI to extract multipath features and uses carriage state regularization to mitigate overfitting due to IMU data bias. We further propose PRID-lite based on a binarized neural network to substantially cut memory requirements for resource-constrained devices. We have conducted extensive experiments under different multipath environments, data bias levels, and a crowdsourced dataset. Our results show that PRID significantly reduces false detection cases compared with the existing arts (by over 50%). PRID-lite further reduces over 90% PRID model size and extends 60% battery life, with a minor compromise in accuracy (7%).
Tianlang He, Jiajie Tan, Weipeng Zhuo, Maximilian Printz, Shueng-Han Gary Chan
INFOCOM2
2022 A Syntax-enhanced model based on category keywords for biomedical relation extraction
Xiaofeng Liu 0014, Jiajie Tan, Jianye Fan, Kaiwen Tan 0001, Jinlong Hu 0002, Shoubin Dong
J. Biomed. Informatics2
2022 Pedometer-free Geomagnetic Fingerprinting with Casual Walking Speed
abstract
The geomagnetic field has been wildly advocated as an effective signal for fingerprint-based indoor localization due to its omnipresence and local distinctive features. Prior survey-based approaches to collect magnetic fingerprints often required surveyors to walk at constant speeds or rely on a meticulously calibrated pedometer (step counter) or manual training. This is inconvenient, error-prone, and not highly deployable in practice. To overcome that, we propose Maficon, a novel and efficient pedometer-free approach for geo ma gnetic fi ngerprint database con struction. In Maficon, a surveyor simply walks at casual (arbitrary) speed along the survey path to collect geomagnetic signals. By correlating the features of geomagnetic signals and accelerometer readings (user motions), Maficon adopts a self-learning approach and formulates a quadratic programming to accurately estimate the walking speed in each signal segment and label these segments with their physical locations. To the best of our knowledge, Maficon is the first piece of work on pedometer-free magnetic fingerprinting with casual walking speed. Extensive experiments show that Maficon significantly reduces walking speed estimation error (by more than 20%) and hence fingerprint error (by 35% in general) as compared with traditional and state-of-the-art schemes.
Jiajie Tan, Shueng-Han Gary Chan
ACM Trans. Sens. Networks2
2021 Efficient Association of Wi-Fi Probe Requests under MAC Address Randomization
abstract
Wi-Fi-enabled devices such as smartphones periodically search for available networks by broadcasting probe requests which encapsulate MAC addresses as the device identifiers. To protect privacy (user identity and location), modern devices embed random MAC addresses in their probe frames, the so-called MAC address randomization. Such randomization greatly hampers statistical analysis such as people counting and trajectory inference. To mitigate its impact while respecting privacy, we propose Espresso, a simple, novel and efficient approach which establishes probe request association under MAC address randomization. Espresso models the frame association as a flow network, with frames as nodes and frame correlation as edge cost. To estimate the correlation between any two frames, it considers the multimodality of request frames, including information elements, sequence numbers and received signal strength. It then associates frames with minimum-cost flow optimization. To the best of our knowledge, this is the first piece of work that formulates the probe request association problem as network flow optimization using frame correlation. We have implemented Espresso and conducted extensive experiments in a leading shopping mall. Our results show that Espresso outperforms the state-of-the-art schemes in terms of discrimination accuracy (> 80%) and V-measure scores (> 0.85).
Jiajie Tan, Shueng-Han Gary Chan
INFOCOM1
2019 Efficient Locality Classification for Indoor Fingerprint-Based Systems
abstract
Locality classification is an important component to enable location-based services. It entails two sequential queries: 1) whether a target is within the site or not, i.e., inside/outside region decision, and 2) if so, which area in the region the target is located, i.e., area classification. Locality classification is hence more coarse-grained and efficient as compared with pinpointing the exact target location in the region. The classification problem is challenging, because fingerprints may not exist outside the region for training. Furthermore, the target may sample an incomplete RSSI vector due to, say, random signal noise, momentary occlusion, or scanning duration. The algorithm also has to be computationally efficient. We propose INOA, a scalable and practical locality classification overcoming the above challenges. INOA may serve as a plug-in before any fingerprint-based localization, and can be incrementally extended to cover new areas or regions for large-scale deployment. Its preprocessor cherry-picks only those discriminating access points, which greatly enhances computational efficiency and accuracy. By formulating a “one-class” classifier using ensemble learning, INOA accurately decides whether the target is within the region or not. Extensive experimental trials in different sites validate the high efficiency and accuracy of INOA, without the need of full RSSI vectors collected at the target.
Ka-Ho Chow 0001, Suining He, Jiajie Tan, Shueng-Han Gary Chan
IEEE Trans. Mob. Comput.3
2019 Efficient Indoor Localization Based on Geomagnetism
abstract
Geomagnetism is promising for indoor localization due to its omnipresence, high stability, and availability of magnetometers in smartphones. Previous works often fuse it with pedometer via particles, which are computationally intensive and make strong user behavior assumptions. To overcome that, we propose Magil, an approach leveraging geo mag netism for i ndoor l ocalization. To our best knowledge, this is the first piece of work using geomagnetism for smartphone localization without the need of pedometer or user walking model. Magil is applicable to any open or complex indoor environment. In the offline phase, Magil collects and stores geomagnetic fingerprints while surveyors walk indoors. In the online phase, it employs a fast algorithm to match the geomagnetic variation of the target with the stored fingerprints. Given closely matched segments, Magil constructs user trajectory efficiently with a modified shortest path formulation by selecting and connecting these matched segments. To further improve accuracy and deployability, we propose MagFi, which extends Mag il by fusing it with Wi- Fi . MagFi further collects opportunistic Wi-Fi RSSI for fingerprint construction. We have implemented both Magil and MagFi and conducted extensive experiments in our campus. Results show that our schemes outperform state-of-the-art schemes by a wide margin (often cutting localization error by 30%).
Ziliang Mo, Jiajie Tan, Suining He, Shueng-Han Gary Chan
ACM Trans. Sens. Networks3
2016 Towards area classification for large-scale fingerprint-based system
abstract
In spacious and multi-area buildings, fingerprint-based localization often suffers from expensive location search. Besides, context knowledge like inside/outside-region and floor area is important for complete location service. To address above issues, beyond the algorithms finding the exact location point, we study accurate and efficient indoor area classification for large-scale fingerprint-based system. We first study leveraging the one-class classification to conduct inside/outside-region detection given only the inside fingerprints. Then we discuss different area determination algorithms, and compare their detection accuracy and deployment efficiency. To further enhance accuracy, we also discuss rejecting unclassifiable signals and calibrating heterogeneous devices. We have implemented different algorithms on Android platforms. Experimental trials (totally over 30,000 fingerprints and 15,000 test data) at an international airport, a business building, a premium shopping mall and a university campus have evaluated practicability and deployability of different classification schemes. Our studies can also serve as design guidelines for area classification.
Suining He, Jiajie Tan, Shueng-Han Gary Chan
UbiComp2