QuickEd: high-performance exact sequence alignment based on bound-and-align

Carregant...
Miniatura
El pots comprar en digital a:
El pots comprar en paper a:

Projectes de recerca

Unitats organitzatives

Número de la revista

Títol de la revista

ISSN de la revista

Títol del volum

Col·laborador

Editor

Tribunal avaluador

Realitzat a/amb

Tipus de document

Article

Data publicació

Editor

Oxford University Press

Condicions d'accés

Accés obert

Llicència

Creative Commons
Aquesta obra està protegida pels drets de propietat intel·lectual i industrial corresponents. Llevat que s'hi indiqui el contrari, els seus continguts estan subjectes a la llicència de Creative Commons: Reconeixement 4.0 Internacional

Assignatures relacionades

Assignatures relacionades

Publicacions relacionades

Datasets relacionats

Datasets relacionats

Projecte CCD

Abstract

Motivation: Pairwise sequence alignment is a core component of multiple sequencing-data analysis tools. Recent advancements in sequencing technologies have enabled the generation of longer sequences at a much lower price. Thus, long-read sequencing technologies have become increasingly popular in sequencing-based studies. However, classical sequence analysis algorithms face significant scalability challenges when aligning long sequences. As a result, several heuristic methods have been developed to improve performance at the expense of accuracy, as they often fail to produce the optimal alignment. Results: This paper introduces QuickEd, a sequence alignment algorithm based on a bound-and-align strategy. First, QuickEd effectively bounds the maximum alignment-score using efficient heuristic strategies. Then, QuickEd utilizes this bound to reduce the computations required to produce the optimal alignment. Compared to Oðn2Þ complexity of traditional dynamic programming algorithms, QuickEd’s bound-and-align strategy achieves OðnbsÞ complexity, where n is the sequence length and bs is an estimated upper bound of the alignment-score between the sequences. As a result, QuickEd is consistently faster than other state-of-the-art implementations, such as Edlib and BiWFA, achieving performance speedups of 4:2-5:9× and 3:8-4:4×, respectively, aligning long and noisy datasets. In addition, QuickEd maintains a stable memory footprint below 35 MB while aligning sequences up to 1 Mbp. Availability and implementation: QuickEd code and documentation are publicly available at https://github.com/maxdoblas/QuickEd.

Descripció

Persones/entitats

Document relacionat

Versió de

Citació

Doblas, M. [et al.]. QuickEd: high-performance exact sequence alignment based on bound-and-align. "Bioinformatics (Oxford)", 4 Març 2025, vol. 41, núm. 3, btaf112.

Ajut

Forma part

Dipòsit legal

ISBN

ISSN

1367-4811

Altres identificadors

Referències