Preprint FZJ-2026-04462

http://join2-wiki.gsi.de/foswiki/pub/Main/Artwork/join2_logo100x88.png
NUMA balancing hampering performance of spiking network simulations

 ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;

2026
arXiv

arXiv () [10.48550/ARXIV.2607.22275]

This record in other databases:  

Please use a persistent id in citations: doi:

Abstract: Computing centers today mostly operate conventional CPU- and GPU-based systems, where the direct way of decreasing energy consumption is a reduction in the applications' runtime. Neuromorphic computing promises an alternative architecture with improved energy efficiency for artificial intelligence. In this endeavor, code for the simulation of large-scale spiking networks on conventional supercomputers is the reference. We show that turning off automatic NUMA balancing may reduce energy consumption by 30%. This dwarfs other attempts of increasing the energy efficiency of a computing center with respect to cost effectiveness. The memory access pattern of spiking network simulation code dynamically interacts with automatic NUMA balancing. This does not affect the correctness of simulation results and thus goes unnoticed in day-to-day neuroscience research. In performance analysis, however, time measurements fluctuate obstructing attempts to optimize simulation technology. A new time- and compute-node resolved performance display exposes the fine-grained temporal variability of distributed spiking network simulations. The analysis uncovers that automatic NUMA balancing is of disadvantage and affects the jemalloc library for thread-aware memory allocation in a transient manner. The method also allows developers to detect perturbations of the HPC system and target specific improvements to simulation technology. As a consequence, we have equipped our supercomputers with an option to turn on or off automatic NUMA balancing on a per-job basis on the user level. This gives researchers the opportunity to find the best setting for the application at hand. There are indications in the literature that the effect has been observed before, yet it does not seem common knowledge in scientific computing. It remains to be investigated how widespread the phenomenon is among scientific codes.

Keyword(s): Distributed, Parallel, and Cluster Computing (cs.DC) ; Neurons and Cognition (q-bio.NC) ; FOS: Computer and information sciences ; FOS: Biological sciences


Contributing Institute(s):
  1. Computational and Systems Neuroscience (IAS-6)
  2. Neuromorphic Software Eco System (PGI-15)
  3. Jülich Supercomputing Center (JSC)
Research Program(s):
  1. 5231 - Neuroscientific Foundations (POF4-523) (POF4-523)
  2. 5232 - Computational Principles (POF4-523) (POF4-523)
  3. BMFTR 03ZU2106CB - NeuroSys: Algorithm-Hardware Co-Design (Projekt C) - B (03ZU2106CB) (03ZU2106CB)
  4. BMBF 03ZU1106CA - NeuroSys: Algorithm-Hardware Co-Design (Projekt C) - A (03ZU1106CA) (03ZU1106CA)
  5. EBRAINS 2.0 - EBRAINS 2.0: A Research Infrastructure to Advance Neuroscience and Brain Health (101147319) (101147319)
  6. JL SMHB - Joint Lab Supercomputing and Modeling for the Human Brain (JL SMHB-2021-2027) (JL SMHB-2021-2027)
  7. GRK 2416 - GRK 2416: MultiSenses-MultiScales: Neue Ansätze zur Aufklärung neuronaler multisensorischer Integration (368482240) (368482240)
  8. DFG project G:(GEPRIS)545776403 - FOR 5880: Ganzheitliche Energie- und Leistungsmodellierung für nachhaltiges Rechnen (Mod4Comp) (545776403) (545776403)
  9. HiRSE - Helmholtz Platform for Research Software Engineering (HiRSE-20250220) (HiRSE-20250220)

Appears in the scientific report 2026
Click to display QR Code for this record

The record appears in these collections:
Institute Collections > IAS > IAS-6
Institute Collections > PGI > PGI-15
Document types > Reports > Preprints
Workflow collections > Public records
Institute Collections > JSC
Publications database

 Record created 2026-09-17, last modified 2026-09-23


External link:
Download fulltext
Fulltext
Rate this document:

Rate this document:
1
2
3
 
(Not yet reviewed)