Parallel execution with mache.parallel
mache.parallel provides a machine-aware interface for launching parallel
workloads based on each machine’s config file.
Typical downstream workflow
Downstream software (for example, Polaris software) can:
Load machine config with
MachineInfo.Build a parallel-system object with
get_parallel_system().Query available resources (
cores,nodes,gpus,memory, andmpi_allowed).Build a machine-correct launcher command with
get_parallel_command().Use the command for either generated job scripts or direct subprocess calls.
Example: build a launcher command
from mache import MachineInfo
from mache.parallel import get_parallel_system
machine_info = MachineInfo()
parallel_system = get_parallel_system(machine_info.config)
args = ["python", "-m", "your_package.run_task", "--case", "smoke"]
command = parallel_system.get_parallel_command(
args=args,
ntasks=4,
cpus_per_task=2,
gpus_per_task=0,
)
print(" ".join(command))
On a batch allocation, this returns an srun/mpiexec command using the
machine’s configured launcher and resource flags. On login nodes for slurm
or pbs systems, get_parallel_system() falls back to login, where MPI is
intentionally disabled.
For slurm systems, get_parallel_system() also falls back to login when
SLURM_JOB_ID is set but the allocation it names has already ended. salloc
does not kill its shell when the allocation is revoked, so that shell keeps the
variable and goes on claiming nodes it no longer has. mache asks the scheduler
and warns before falling back. If the scheduler cannot be reached at all, it
raises instead, rather than quietly demoting a real allocation to a login node.
The scheduler is only asked when nothing local can answer. A process running on
one of the allocation’s own nodes – a batch script and anything it launches –
is proof that the allocation is live, because Slurm kills a job’s processes
before it releases its nodes. mache recognizes that case from SLURMD_NODENAME
and SLURM_JOB_NODELIST and issues no query at all. This matters to callers
that run many short-lived processes inside one allocation, since a query per
process is a rate that sites ask jobs to stay well under.
GPU-per-task flags
When gpus_per_task > 0 is passed to get_parallel_command():
slurmsystems add--gpus-per-task <N>by default. This can be overridden withgpus_per_task_flagin the machine’s[parallel]config.pbssystems require a machine-specificgpus_per_task_flagto be set in config before a GPU-per-task argument is added.
Hyperthreading
mache.parallel does not currently have a dedicated
hyperthreading = true/false switch. Instead, hyperthreading behavior is
controlled through the machine’s [parallel] config and the resource values
passed to get_parallel_command(). The most important config knobs are
cores_per_node, max_mpi_tasks_per_node, cpu_bind, and any
launcher-specific arguments included in parallel_executable.
The default convention in mache’s machine configs is to describe CPU resources in terms of physical cores, not hardware threads. For E3SM itself, and for most downstream software, this means:
cores_per_nodeshould usually be the number of physical CPU cores per nodemax_mpi_tasks_per_nodeshould usually reflect the intended non-hyperthreaded MPI rank count per nodecpu_bind = coresis often a good default when the launcher and machine topology support it, but some systems such as Frontier prefercpu_bind = threadscpus_per_taskshould usually be sized assuming physical cores
This is why several shipped machine configs explicitly document
cores_per_node as the count “without hyperthreading”.
If a downstream application wants to take advantage of hyperthreading, it should opt in by overriding the relevant parallel config values for that use case. In practice, that usually means switching from physical-core counts to hardware-thread counts and adjusting binding accordingly. For example, on a machine with 64 physical cores and 2 hardware threads per core:
[parallel]
cores_per_node = 128
max_mpi_tasks_per_node = 128
cpu_bind = threads
Then, calls to get_parallel_command() should use cpus_per_task and
ntasks values that match that threaded layout.
The important point is that hyperthreading is opt-in. Mache’s default machine configs should generally preserve the physical-core layout that is appropriate for E3SM and most downstream tools, while still allowing downstream users to provide a config override when they intentionally want thread-level placement.
Using this in generated job scripts
A common pattern is to generate scheduler directives separately, then use
mache.parallel only for launch lines. For example:
Use
MachineInfo.get_account_defaults()to populate account/partition/QOS.Use
MachineInfo.get_queue_specs(),MachineInfo.get_partition_specs()orMachineInfo.get_qos_specs()for optional scheduler-target policy metadata (min_nodes,max_nodes,max_wallclock,max_wallclock_bins) when available.Render scheduler headers (
#SBATCHor#PBS) in your template logic.Use
get_parallel_command()to build the executable line.
This keeps scheduler policy in your tool while reusing machine-specific launch
behavior from mache.
Slurm distribution options
For slurm systems, mache supports two ways to control srun -m:
distribution = <value>passes a raw Slurm distribution string directly as-m <value>, for exampleblock:cyclicorblock:blockplacement = <value>preserves mache’s legacy behavior and expands to-m <value>=<max_mpi_tasks_per_node>, for exampleplane=56
If both are present, distribution takes precedence. Prefer distribution
for machines whose documented Slurm usage relies on explicit values like
block:cyclic rather than the older plane=<tasks> form.
Note
The placement config option here is the Slurm task distribution and is
unrelated to the placement argument to get_parallel_command() described in
Placing concurrent launches within one allocation. Passing that argument drops this config
option from the command.
Selecting scheduler options by node count
mache.parallel also provides helpers for selecting queue/partition/QOS from
machine metadata:
ParallelSystem.get_scheduler_target(config, target_type, nodes)selects one ofqueue,partition, orqos.ParallelSystem.resolve_submission(config, nodes, target_type, min_nodes_allowed=None, requested=None, desired_wall_time=None)returns aSubmissionResolutionwith fieldstarget,requested_nodes,effective_nodes,adjustment(exact,decrease, orincrease),honored, andreason.SlurmSystem.resolve_slurm_options(config, nodes, min_nodes_allowed=None, partition=None, qos=None, constraint=None, desired_wall_time=None, scheduler_target=None)returns aSlurmOptionsobject with fieldspartition,qos,constraint,gpus_per_node,max_wallclock,effective_nodes,wall_time,honored, andreason.PbsSystem.resolve_pbs_options(config, nodes, min_nodes_allowed=None, queue=None, constraint=None, desired_wall_time=None, scheduler_target=None)returns aPbsOptionsobject with fieldsqueue,constraint,gpus_per_node,max_wallclock,filesystems,effective_nodes,wall_time,honored, andreason.
For invalid gaps between scheduler ranges, node count is adjusted to the
nearest valid value, preferring lower adjustments when feasible. If
min_nodes_allowed disallows lower adjustments, resolution moves to the next
valid higher range. If no feasible target exists, these functions raise
ValueError.
Note
SlurmSystem.get_slurm_options() and PbsSystem.get_pbs_options() return the
same values as tuples. They are deprecated as of v3.11.0 in favor of
resolve_slurm_options() and resolve_pbs_options(), which can be extended
with new fields without breaking positional unpacking.
Requesting a specific queue, partition or QOS
Callers that want a particular scheduler target – for example a test suite
that should run in the debug QOS – can ask for one directly instead of
rewriting the machine’s config:
from mache import MachineInfo
from mache.parallel.slurm import SlurmSystem
config = MachineInfo(machine="pm-cpu").config
options = SlurmSystem.resolve_slurm_options(
config=config,
nodes=4,
qos="debug",
desired_wall_time="02:00:00",
)
if not options.honored:
print(f"Falling back to the {options.qos} qos: {options.reason}")
A requested target is a preference, not an assertion. mache honors it when the
machine’s metadata allows it and otherwise resolves the default target,
setting honored = False and putting a printable explanation in reason. A
request is not honored when:
the target is not in the machine’s
[parallel]queues/partitions/qoslist,clamping the node count to the target’s
min_nodes/max_nodeswould fall belowmin_nodes_allowed, ordesired_wall_timeis longer than the target’smax_wallclock.
A constraint can be requested the same way. Unlike a queue, partition or QOS,
it has no node-count or wall-clock metadata and no [constraint.*] section, so
it is validated only against the machine’s [parallel] constraints list: a
constraint that is not on that list falls back to the machine’s default with a
reason, exactly as the other targets do, and a machine that defines no
constraints ignores the request entirely.
Clamping the node count on its own does not prevent a target from being
honored. The clamp is reported through effective_nodes and adjustment, and
min_nodes_allowed is the guard for a clamp the caller cannot live with.
requested values of None, an empty string, and placeholders of the form
<<<default>>> all mean “no target was requested”, so config-driven callers
can pass their raw config value through without guarding against unset
placeholders.
A request is also ignored, rather than denied, when the machine defines no
targets of that type at all. Machines hang a concept like “debug” off
different axes – a partition on Chrysalis, a QOS on Frontier and Perlmutter,
a queue on Aurora – so a caller that asks on more than one axis should not be
told its request was refused on the axes the machine does not use. There was
no choice to deny, so honored stays True and reason stays None. A
target missing from a list the machine does define is still a denied
request.
Requesting a target without naming its axis
Asking on every axis is still awkward for a caller whose intent is simply
“use this machine’s debug target”. scheduler_target says it once and lets
mache work out which axis this machine uses:
options = SlurmSystem.resolve_slurm_options(
config=config,
nodes=2,
scheduler_target="debug",
desired_wall_time="00:20:00",
)
This selects the debug partition on Chrysalis, the debug QOS on Frontier
and Perlmutter (leaving the default batch partition in place on Frontier),
and the debug queue on Aurora, with no spurious reason on any of them.
Slurm machines are searched partitions-first and then QOS; PBS machines
schedule by queue, so only queues are searched. partition, qos and queue
take precedence on the axis they name, so a caller can set a broad
scheduler_target and still pin one axis explicitly.
Once an axis is chosen the target is resolved exactly as if it had been
requested there, so it is still subject to that target’s node and wall-clock
metadata: scheduler_target="debug" with a three-hour wall time on Frontier
falls back to the normal QOS and says why. A name that appears on no axis at
all is a genuine failed request, and the reason lists what each axis does
offer.
When desired_wall_time is given, the returned wall_time is that value
capped at the selected target’s max_wallclock. The same capping is available
on its own:
from mache.parallel.system import cap_wall_time
cap_wall_time("04:00:00", "00:30:00") # "00:30:00"
Note that mache resolves the partition and the QOS independently and has no concept of one being valid only with the other. A caller that requests both is responsible for asking for a combination its machine accepts.
Wall-clock limits that depend on job size
On some machines, the maximum wall time depends on how many nodes a job asks
for. Frontier’s batch partition allows 2 hours for 1-91 nodes, 6 hours for
92-183 nodes, and 12 hours above that. These machines describe their policy
with max_wallclock_bins rather than a single max_wallclock (see
Adding a new config file), and mache selects the bin that matches the
resolved node count:
options = SlurmSystem.resolve_slurm_options(config=config, nodes=8)
options.max_wallclock # "02:00:00" on Frontier
options = SlurmSystem.resolve_slurm_options(config=config, nodes=200)
options.max_wallclock # "12:00:00" on Frontier
When both a partition and a QOS set a limit, the more restrictive of the two
is reported, and that is also the limit wall_time is capped at.
Placing concurrent launches within one allocation
get_parallel_command() normally asks for resources in the abstract – this
many tasks, this many CPUs each – and lets the machine decide where the work
runs. That is enough while a tool runs one piece of work at a time. As soon as
it wants to run two inside the same allocation, the two launches are given
overlapping resources or, more often, the second waits until the first has
finished.
Note
placement, memory_cap and memory_per_node are new in v3.12.0. A tool
that depends on them should require at least that version rather than test
for the capability: on an older mache, get_parallel_command() takes no
placement at all, and a launch that expected to be confined to part of a
node would instead be free to use the whole allocation – which looks like a
working run right up until two of them collide.
An optional placement says where a launch should run:
from mache import MachineInfo
from mache.parallel import ResourcePlacement, get_parallel_system
parallel_system = get_parallel_system(MachineInfo().config)
placement = ResourcePlacement(
nodes=["nid001373"],
cores=list(range(8, 16)),
)
command = parallel_system.get_parallel_command(
args=["./run_step.py"],
ntasks=1,
cpus_per_task=8,
placement=placement,
)
A placement carries three things: the nodes the launch may use, the cores it may use on each of them, and how many GPUs it needs in total. mache renders them into whatever the machine’s launcher needs, so callers do not have to know which flags a given site takes.
A call without a placement produces exactly the command it produced before this feature existed.
Checking what a machine supports
Not every machine can confine a launch. Check before running things concurrently, rather than discovering the answer as a hang or as silent oversubscription:
from mache.parallel import PlacementSupport
support = parallel_system.placement_support
if support is PlacementSupport.NONE:
print("this machine cannot place launches; run steps one at a time")
elif support is PlacementSupport.CPU_BINDING:
print("placement is by CPU binding, which the work itself could ignore")
The three values are:
PlacementSupport.SCHEDULER– the batch system reserves what each launch asks for, so a launch cannot exceed what it was given. This is Slurm 20.11 and newer.PlacementSupport.CPU_BINDING– the launcher binds each task to specific cores, which keeps concurrent launches apart but reserves nothing. Work that rebinds itself is not prevented from doing so. This is Slurm before 20.11, PBS with PALS, andsingle_node.PlacementSupport.NONE– there is no mechanism here. Passing a placement raisesValueErrorrather than producing a command that would be accepted and then silently do nothing.
This is determined at run time from the launcher actually present, not from the machine’s config, because a site can be upgraded without its mache config changing.
GPUs are a total, not a count per task
ResourcePlacement.gpus is the number of GPUs for the whole launch. This is
deliberately unlike cpus_per_task: asking for a number of GPUs per task
was measured not to confine a launch on either of the GPU machines mache
supports, while a per-launch total does.
gpus defaults to 0, and 0 is rendered as an explicit request for no GPUs.
On Slurm this matters more than it sounds: a launch that says nothing about
GPUs is read as claiming every one on the node, so the next launch waits.
Callers whose work uses no GPUs – most of them – get correct behavior
without having to know GPUs were ever a consideration.
On PALS the same request is belt and braces rather than the mechanism that makes concurrency work, since nothing there reserves a GPU in the first place. See Assigning GPUs on PBS with PALS.
Which cores are honored
cores is an explicit set rather than a count, because the usable cores on a
node may not be contiguous and may not start at zero – Aurora reserves core 0
and cores 49-52 – and because a count cannot say which cores.
How much of that set is honored depends on the mechanism:
where the scheduler reserves resources, only the size of the set is used; Slurm is asked for that many cores and picks which ones itself, and an explicit core list is rejected outright alongside
-cwhere placement is by CPU binding, the set is used exactly as given, in order, split into one contiguous chunk of
cpus_per_taskcores per task
Either way, mache raises ValueError if the set is too small for
ntasks x cpus_per_task.
Assigning GPUs on PBS with PALS
PALS has no scheduler to hand out GPUs, so isolation there is by the vendor’s
visible-device variable – ZE_AFFINITY_MASK on Aurora,
CUDA_VISIBLE_DEVICES on Polaris. mache renders it into the mpiexec command
and needs to be told which devices to name:
placement = ResourcePlacement(
nodes=["x4401c1s0b0n0"],
cores=list(range(1, 9)),
gpus=1,
gpu_ids=[2],
)
gpu_ids are indices from 0 to gpus_per_node - 1; mache maps them to
whatever form the machine’s variable takes, including Aurora’s device.tile
addressing. Only the caller knows about every launch running at that moment,
so only the caller can assign disjoint GPUs – mache renders what it is given
and never guesses. A placement with gpus > 0 and no gpu_ids raises on
PALS, and len(gpu_ids) must equal gpus.
gpu_ids is ignored where the scheduler assigns GPUs itself, which is every
Slurm machine.
A placement with no GPUs sets the variable to an empty value. Note that this
is weaker than the equivalent on Slurm: PALS reserves nothing, so a launch
that stays quiet about GPUs does not block the next one, and how much an
empty value actually hides has not been measured. An empty
CUDA_VISIBLE_DEVICES means “no devices”, but an empty ZE_AFFINITY_MASK
may instead mean “no mask”, which is every tile. Do not rely on it to keep a
GPU launch and a CPU launch off the same device – give the GPU launch
explicit gpu_ids instead.
Two config options support this, both already set on the machines that need
them: gpu_visible_devices_var names the variable, and the ordered
gpu_bind = list:... binding list, where a machine has one, says how its
devices are named.
What a placement overrides
A placement is the authority on which resources a launch gets, so it supersedes the machine’s config options that describe spreading a launch over a whole node:
distributionand the legacyplacementconfig option are dropped, since the placement has already said which nodes and how many cores the launch getscpu_bind,gpu_bindandmem_bindare dropped when they name specific cores or devices, as Aurora’s docpu_bindis also dropped wherever the placement renders its own bindinggpu_bindis dropped when the placement asks for no GPUs, and also when it isnone, which asks for no binding at all
A binding policy such as cpu_bind = cores or gpu_bind = closest is kept
where it does not conflict, since it still applies within whatever the launch
was given.
Note
gpu_bind = none is dropped because keeping it appears to cost a placement
its GPUs. Of four concurrent placed launches on Perlmutter GPU asking for one
GPU each, one was given a GPU and the other three got none, ran anyway and
exited 0. Frontier, whose gpu_bind is closest, gave all four disjoint
GPUs from a nearly identical command. Dropping none takes nothing away,
since Slurm does not bind tasks to GPUs without the option either.
Note
Verifying GPU placement from inside a launch needs the scheduler’s global GPU
identifiers, such as SLURM_STEP_GPUS. CUDA_VISIBLE_DEVICES is renumbered
per launch, so four launches on four different GPUs all report device 0.
How much memory a machine has
A caller deciding how many pieces of work fit inside one allocation needs to
know how much memory it has to divide up. mache reports it beside the core
and GPU counts:
from mache import MachineInfo
from mache.parallel import get_parallel_system
parallel_system = get_parallel_system(MachineInfo().config)
print(parallel_system.memory_per_node) # MB on one node
print(parallel_system.memory) # MB across the whole allocation
memory is memory_per_node times the node count, exactly as cores is
cores_per_node times the node count. Both are in MB, which is the unit
Slurm’s memory options default to.
The figure is the memory a job may actually use – what the site reports as
available, rounded down – not the hardware capacity of a node. On a Slurm
machine that is the MEMORY column of sinfo, which is already net of what
the operating system and the site’s own services hold back. The two differ by
several percent, and the whole point of the number is that a caller can pack
up to it.
Warning
Most shipped values are still estimates, rounded down from the node memory each site documents, because the figure that matters can only be read off the machine itself. A config whose value has not been measured there says so in a comment above the option. Estimates err low on purpose: packing less work than a node could hold is wasteful, while packing more is a job killed for exhausting the node.
A machine whose config does not set memory_per_node reports None for both,
rather than 0, since no machine has no memory. Every machine mache ships a
config for sets it; a site-specific or user config may not. The login-node
system always reports None: memory_per_node describes a compute node, and
a login node neither has that much nor hands out what it does have.
Note
memory_per_node describes the machine, and a ResourcePlacement
deliberately does not carry memory: a placement says where a launch runs,
and how much memory it may use is a separate statement, made with a separate
argument – see Capping the memory a launch may use. Deciding how much each
piece of work may take remains the caller’s, because only the caller knows
what else it is running.
Capping the memory a launch may use
get_parallel_command() takes an optional memory_cap, in MB, which is the
most memory the launch may use on each node it runs on:
command = parallel_system.get_parallel_command(
args=['./run.py'],
ntasks=4,
cpus_per_task=2,
placement=placement,
memory_cap=16000,
)
It is absent by default. Without it, nothing about memory is rendered and the command is exactly what it would have been – which is also what every existing caller gets.
The unit and the per-node denomination match memory_per_node, since that is
the figure a caller divides up between the launches it runs at once.
Warning
memory_cap is a cap, not a reservation. Nothing measured suggests it
sets memory aside for the launch, so staying under a cap does not protect a
launch from a concurrent one that ignores its own. What it does do, where the
batch system acts on it, is kill the launch for exceeding it: a step allowed
1024 MB and told to allocate 4 GB is killed at 960 MB on Perlmutter GPU and
on Frontier.
Where a cap is worth anything
Not every machine will hold a launch to a cap, so mache renders one only
where it will be acted on, and reports which kind of machine this is:
from mache.parallel import MemoryCapSupport
if parallel_system.memory_cap_support is MemoryCapSupport.NONE:
print('memory caps are not enforced here')
MemoryCapSupport.ENFORCED– the batch system holds the launch to its cap and kills it for exceeding it. Slurm 20.11 and newer.MemoryCapSupport.NONE– nothing here will hold a launch to a cap, somacherenders none. Either the launcher has no memory option at all, as PALS on Aurora does not, or it accepts one and does not act on it, as Chrysalis’s Slurm does with--mem.
Passing a cap to a machine that reports NONE is not an error and does not
raise: the work still runs correctly, it is simply unprotected. This differs
from passing a placement to a machine that cannot honor one, which does
raise, because that means concurrent launches collide rather than merely go
unguarded. Check memory_cap_support when it matters that the cap is real.
Note
mache renders nothing on a machine that will not act on a cap, rather than
rendering an option the machine accepts and ignores. A figure on the command
line that nothing keeps reads to anyone who sees the command as a safety net
that is not there.
What a placement does not do about memory
A placement does not cap memory on its own. A placed single-core launch and an unplaced control, neither mentioning memory, both allocated twice what a single core’s proportional share of the node would be, and neither was touched – measured on Perlmutter CPU and Frontier. Asking Slurm for exactly the cores a launch needs does not hand it a slice of the node’s memory to go with them.
The reassuring half of the same result is that saying nothing about memory does not repeat the trap that saying nothing about GPUs sets. An unstated memory requirement is not read as a claim on the node’s memory: four concurrent launches that say nothing about it start together and run.