Poster (After Call) FZJ-2026-04710

http://join2-wiki.gsi.de/foswiki/pub/Main/Artwork/join2_logo100x88.png
Best Practices for Training LLMs on HPC Systems

 ;  ;  ;  ;

2026

SciCoCo.nrw – Science, Compute, Connect, SciCoCo.nrw, AachenAachen, Germany, 21 Sep 2026 - 23 Sep 20262026-09-212026-09-23

Abstract: The training of large language models (LLMs) requires substantial computational resources, complex software stacks, and carefully designed workflows to achieve scalability and efficiency. This poster presents best practices and insights gained from the training of open, multilingual LLMs optimized for European languages. We detail the use of high-performance computing (HPC) systems, primarily JUWELS Booster at the Jülich Supercomputing Centre, for training a 7-billion-parameter transformer model. We include details about system architecture, training infrastructure, software choices, profiling and benchmarking tools, as well as engineering and operational challenges. Measured throughput data is provided that sheds light on 3D parallelism and the impact of features such as flash attention during training.


Contributing Institute(s):
  1. Jülich Supercomputing Center (JSC)
Research Program(s):
  1. 5112 - Cross-Domain Algorithms, Tools, Methods Labs (ATMLs) and Research Groups (POF4-511) (POF4-511)
  2. ATML-X-DEV - ATML Accelerating Devices (ATML-X-DEV) (ATML-X-DEV)
  3. NOVAS - Novel System Architectures Design (NOVAS) (NOVAS)

Appears in the scientific report 2026
Click to display QR Code for this record

The record appears in these collections:
Dokumenttypen > Präsentationen > Poster
Workflowsammlungen > Öffentliche Einträge
Workflowsammlungen > In Bearbeitung
Institutssammlungen > JSC
Online First

 Datensatz erzeugt am 2026-10-01, letzte Änderung am 2026-10-08


Restricted:
Volltext herunterladen PDF
Dieses Dokument bewerten:

Rate this document:
1
2
3
 
(Bisher nicht rezensiert)