JTop algorithms for top-k join queries
Loading...
Date
Journal Title
Journal ISSN
Volume Title
Publisher
University of Waterloo
Abstract
Top-k join queries are very important, because there are many applications in which users need to join multiple inputs and are interested i the top-k join results based on some scoring function that combines some attribute values of each input. One of the most efficient algorithms for top-k join queries is the Rank-Join algorithm. In this report, we first study this algorithm and show that there are many cases where its threshold is lazy, i.e. decreases very slowly, and the algorithm needs to go too far in the lists. Then, we propose a family of efficient algorithms for processing top-k join queries. The main idea is to take advantage of the specific information on join attribute values as well as characteristics of the query and of the underlying system. Our contributions are as follows. First, we propose a general model for the problem of top-k join queries which is useful for databases as well as many other areas of computing. Second, we propose JTop, an efficient top-k join algorithm for systems where random accesses are not expensive. JTop takes advantage of both random and sorted accesses as well as information on the join condition. In contrast to Rank-Join, JTop's threshold is not lazy if at least one of the scoring attributes has a reasonable progressive impact on the scoring function. Third, we propose two new algorithms, LR_JTop and NR_JTop, for systems where random accesses are expensive or not supported, respectively. Forth, we propose a new algorithm called BP_JTop which is designed for systems with position-based indexing, i.e. when accessing a data item, the index gives its position. For each of our algorithms, we prove that over any database, it stops before or at the same position at which Rank-Join stops. We also show that there is a class of databases over which our algorithm stops at a position that is O(n) times lower than that of Rank-Join, where n is the number of data items. We also conducted an extensive experimental study to evaluate the performance of our algorithms under different data distributions. The performance evaluation shows that our algorithms obtain high performance gains against the Rank-Join algorithm.