Parallelizing dense and banded linear algebra libraries using SMPSs

Badia Sala, Rosa Maria; Herrero Zaragoza, José Ramón; Labarta Mancho, Jesús José; Pérez Cáncer, Josep Maria; Quintana Ortí, Enrique Salvador; Quintana Ortí, Gregorio

doi:10.1002/cpe.1463

dc.contributor.author	Badia Sala, Rosa Maria
dc.contributor.author	Herrero Zaragoza, José Ramón
dc.contributor.author	Labarta Mancho, Jesús José
dc.contributor.author	Pérez Cáncer, Josep Maria
dc.contributor.author	Quintana Ortí, Enrique Salvador
dc.contributor.author	Quintana Ortí, Gregorio
dc.contributor.other	Universitat Politècnica de Catalunya. Departament d'Arquitectura de Computadors
dc.date.accessioned	2010-03-08T12:11:15Z
dc.date.available	2010-03-08T12:11:15Z
dc.date.created	2009-12-25
dc.date.issued	2009-12-25
dc.identifier.citation	Badia, R. [et al.]. Parallelizing dense and banded linear algebra libraries using SMPSs. "Concurrency and computation: practice and experience", 25 Desembre 2009, vol. 21, núm. 18, p. 2438-2456.
dc.identifier.issn	1532-0626
dc.identifier.uri	http://hdl.handle.net/2117/6565
dc.description.abstract	The promise of future many-core processors, with hundreds of threads running concurrently, has led the developers of linear algebra libraries to rethink their design in order to extract more parallelism, further exploit data locality, attain better load balance, and pay careful attention to the critical path of computation. In this paper we describe how existing serial libraries such as (C)LAPACK and FLAME can be easily parallelized using the SMPSs tools, consisting of a few OpenMP-like pragmas and a runtime system. In the LAPACK case, this usually requires the development of blocked algorithms for simple BLAS-level operations, which expose concurrency at a finer grain. For better performance, our experimental results indicate that column-major order, as employed by this library, needs to be abandoned in benefit of a block data layout. This will require a deeper rewrite of LAPACK or, alternatively, a dynamic conversion of the storage pattern at run-time. The parallelization of FLAME routines using SMPSs is simpler as this library includes blocked algorithms (or algorithms-by-blocks in the FLAME argot) for most operations and storage-by-blocks (or block data layout) is already in place.
dc.format.extent	19 p.
dc.language.iso	eng
dc.subject	Àrees temàtiques de la UPC::Informàtica::Arquitectura de computadors::Arquitectures paral·leles
dc.subject.lcsh	Embedded computer systems
dc.subject.other	Linear algebra libraries
dc.subject.other	Programmability
dc.subject.other	High performance
dc.subject.other	Dynamic scheduling
dc.subject.other	Multi-core processors
dc.title	Parallelizing dense and banded linear algebra libraries using SMPSs
dc.type	Article
dc.subject.lemac	Ordinadors immersos, Sistemes d'
dc.contributor.group	Universitat Politècnica de Catalunya. CAP - Grup de Computació d'Altes Prestacions
dc.identifier.doi	10.1002/cpe.1463
dc.description.peerreviewed	Peer Reviewed
dc.relation.publisherversion	https://onlinelibrary.wiley.com/doi/abs/10.1002/cpe.1463
dc.rights.access	Restricted access - publisher's policy
local.identifier.drac	1613909
dc.description.version	Postprint (published version)
local.citation.author	Badia, R.; Herrero, J.; Labarta, J.; Pérez, J.; Quintana-Ortí, E.; Quintana-Ortí, G.
local.citation.publicationName	Concurrency and computation: practice and experience
local.citation.volume	21
local.citation.number	18
local.citation.startingPage	2438
local.citation.endingPage	2456

Fitxers d'aquest items

Nom:: Badia.pdf
Mida:: 1,060Mb
Format:: PDF

Visualitza/Obre

Aquest ítem apareix a les col·leccions següents

Articles de revista [1.049]
Articles de revista [382]

Mostra el registre d'ítem simple

UPCommons. Portal del coneixement obert de la UPC

Parallelizing dense and banded linear algebra libraries using SMPSs

Fitxers d'aquest items

Aquest ítem apareix a les col·leccions següents

Explora