跳到主要导航 跳到搜索 跳到主要内容

ACRES: efficient query answering on large compressed sequences

  • Bin Wang
  • , Xiaochun Yang*
  • , Guoren Wang
  • *此作品的通讯作者
  • Northeastern University China

科研成果: 期刊稿件文章同行评审

摘要

With the advances in next generation sequencing, the amount of genomic sequence data being produced continues to grow at an exponential rate. It is estimated that the entire genome of each individual human, each containing about 3 billion letters, could be made available in the next a few years. An increasingly pressing issue in genomics and medicine is how to efficiently store and query these massive amounts of sequence data. Recently a lossless compression technique has been proposed to drastically reduce the storage space of genomic sequences, taking advantage of the fact that any two genomes from the same species are highly similar and therefore only their differences need to be encoded. In this paper we study how to efficiently answer queries on the compressed sequences without first decompressing them. We study three important types of queries, including retrieving a subsequence, finding subsequences matching a given pattern, and finding subsequences similar to a pattern. We propose an index structure, filtering techniques, and efficient algorithms for answering these queries. We further demonstrate the utility of these algorithms using a real dataset.

源语言英语
页(从-至)1349-1376
页数28
期刊World Wide Web
21
5
DOI
出版状态已出版 - 1 9月 2018
已对外发布

联合国可持续发展目标

此成果有助于实现下列可持续发展目标:

  1. 可持续发展目标 3 - 良好健康与福祉
    可持续发展目标 3 良好健康与福祉

学术指纹

探究 'ACRES: efficient query answering on large compressed sequences' 的科研主题。它们共同构成独一无二的学术指纹。

引用此