Rebalancing the core front-end through HPC code analysis

Milic, Ugljesa; Carpenter, Paul Matthew; Rico, Alejandro; Ramirez, Alex

doi:10.1109/IISWC.2016.7581273

Visualitza/Obre

Rebalancing the Core Front-End through HPC Code.pdf (542,3Kb)

Veure estadístiques d'ús d'UPCommons

Estadístiques de LA Referencia / Recolecta

Cita com:

Mostra el registre d'ítem complet

Milic, Ugljesa

Carpenter, Paul Matthew

Rico, Alejandro

Ramirez, Alex

Tipus de documentText en actes de congrés

Data publicació2016-10-10

EditorIEEE

Condicions d'accésAccés obert

Tots els drets reservats. Aquesta obra està protegida pels drets de propietat intel·lectual i industrial corresponents. Sense perjudici de les exempcions legals existents, queda prohibida la seva reproducció, distribució, comunicació pública o transformació sense l'autorització del titular dels drets

ProjecteCOMPUTACION DE ALTAS PRESTACIONES VII (MINECO-TIN2015-65316-P)

Abstract

There is a need to increase performance under the same power and area envelope to achieve Exascale technology in high performance computing (HPC). The today's chip multiprocessor (CMP) design is tailored by traditional desktop and server workloads, different from parallel applications commonly run in HPC. In this work, we focus on the HPC code characteristics and processor front-end which factors around 30% of core power and area on the emerging lean-core type of processors used in HPC. Separating serial from parallel code sections inside applications, we characterize three HPC benchmark suites and compare them to a traditional set of desktop integer workloads. HPC applications have biased and mostly backward taken branches, small dynamic instruction footprints, and long basic blocks. Our findings suggest smaller branch predictors (BP) with the additional loop BP, smaller branch target buffers (BTB), and smaller L1 instruction caches (I-cache) with wider lines. Still, the aforementioned downsizing applies only to the cores meant to run parallel code. The difference between serial and parallel code sections in HPC applications points to an asymmetric CMP design, with one baseline core for sequential and many HPCtailored cores designed for parallel code. Predictions using Sniper simulator and McPAT show that an HPC-tailored lean core saves 16% of the core area and 7% of power compared to a baseline core, without performance loss. Using the area savings to add an extra core, an asymmetric CMP with one baseline and eight tailored cores has the same area budget as a symmetric CMP composed out of eight baseline cores demanding 4% more power and providing 12% shorter execution time on average.

CitacióMilic, Ugljesa [et al.]. Rebalancing the core front-end through HPC code analysis. A: IEEE International Symposium on Workload Characterization (IISWC), 25-27 Sept. 2016. "Workload Characterization (IISWC), 2016 IEEE International Symposium on". IEEE, 2016, p. 128-137.

URIhttp://hdl.handle.net/2117/100362

DOI10.1109/IISWC.2016.7581273

ISBN978-1-5090-3896-1

Versió de l'editorhttp://ieeexplore.ieee.org/document/7581273/

Col·leccions

Computer Sciences - Ponències/Comunicacions de congressos [574]

Veure estadístiques d'ús d'UPCommons

Mostra el registre d'ítem complet

Fitxers	Descripció	Mida	Format	Visualitza
Rebalancing the ... t-End through HPC Code.pdf		542,3Kb	PDF	Visualitza/Obre

UPCommons. Portal del coneixement obert de la UPC

Rebalancing the core front-end through HPC code analysis

Visualitza/Obre

Explora