Getting Started#
seamm_exec runs flowcharts (run_flowchart, run_from_jobserver) and runs
the external codes that steps need. A step gets the executor from its flowchart and
asks it to run a program; the executor handles the program’s configuration from
<root>/<program>.ini (a conda environment, environment modules, a Docker image,
or a plain executable), where the program runs, and which files come back.
Running a code: executor.run()#
The original interface, used by most steps:
result = self.flowchart.executor.run(
config, # the program's section of <root>/<program>.ini
cmd=["{code}", "input.dat", ">", "output.txt"],
directory=self.directory,
files={"input.dat": text}, # written before the run
return_files=["output.txt", "*.log"],
shell=True,
)
output = result["output.txt"]["data"]
in_situ chooses where the code runs: None (the default) runs in a scratch
directory under a scheduler such as SLURM ($TMPDIR, so node-local storage) and
in place otherwise; True always runs in place, so the output can be watched as
it is written; False always uses scratch. Only the return_files come back to
the step directory.
Running codes as tasks#
Since 2026.10.2 a step can describe each calculation as a Task
and run any number of them through a TaskSet:
from seamm_exec import Resources, Task, TaskSet
tasks = TaskSet(self, archive=True)
for key, text in inputs.items():
tasks.add(
Task(
key=key, # stable across reruns, e.g. a fragment key
program="orca",
cmd=["{code}", "orca.inp", ">", "orca.out"],
files={"orca.inp": text},
return_files=["orca.out", "orca.engrad"],
resources=Resources(ntasks=4), # mem_per_cpu in bytes if needed
success_text={"orca.out": "ORCA TERMINATED NORMALLY"},
estimated_seconds=30,
shell=True,
)
)
for result in tasks.run():
if result.ok:
parse(result.files["orca.out"])
else:
print(f"{result.key} failed: {result.reason}")
What this gives a step:
Concurrency. The tasks run through a pool sized to the machine or to the SLURM allocation, each in its own process group; a task’s
Resourcessay how many cores and how much memory it needs, andntasks=Nonemeans all of it.tasks.capacity()reports the pool’s size, for a step that sizes its input (ORCA’s%pal) before writing it.Restart. The set keeps
tasks/manifest.jsonin the step directory and aDONEmarker in each task’s directory,tasks/<key>/. Rerunning the job in the same directory restores finished tasks from their markers and never recomputes them; a task whose inputs changed is recomputed. A task that failed is not retried within a run, but is tried again on a rerun, up to three attempts in all (the count resets when the inputs change). A process that a crashed run left behind is killed on the rerun, after checking it is the same process.Success beyond the exit code.
success_textnames text that must appear in a result file; ORCA exits 0 on an error termination, for example.Archiving. With
archive=Truefinished task directories are packed into one tar per bundle, so a step with thousands of tasks leaves a handful of files.
A step with a single calculation uses the same machinery through
run_task(), which also keeps the output in the step directory, as
executor.run() did:
result = run_task(task, self, directory=self.directory)
A step that handles <program>.ini itself passes the program’s config
(and env) on the task, and the pool uses them as they are. A task without
config is configured where it runs, from that machine’s
<root>/<program>.ini and the program’s resolver (an entry point in
org.molssi.seamm.exec.resolvers; ORCA’s adds its full path and the OpenMPI
paths). In the command, {code} is the program and {code_dir} its
directory, when it has one.
Many structures, one model chemistry#
A step that needs the energy (and gradients, and stress) of many structures
with the flowchart’s model chemistry uses an Evaluator
rather than tasks or MDI directly:
from seamm_exec import Evaluator
with Evaluator(self, properties=("energy", "gradients")) as evaluator:
for configuration in configurations:
evaluator.submit(configuration, key=f"c{configuration.id}")
for result in evaluator.results():
if result.ok:
store(result.key, result.energy, result.gradients) # kJ/mol, kJ/mol/Å
else:
print(f"{result.key} failed: {result.reason}")
The evaluator chooses the path, not the step or the user. On a queue target a
program that can make tasks runs as tasks there; otherwise a warm MDI engine
evaluates the structures one after another, unless the program prefers tasks
(ORCA, whose engine starts ORCA for each structure anyway). Both paths give the
same numbers. A structure the program cannot run as a task (a periodic system
for ORCA or MOPAC) goes to its MDI engine; if there is none here it fails alone,
with the reason. The batch path keeps the task layer’s restart, so a rerun
reuses finished structures. options on submit carries what a fragment
needs: atom_indices, ghost_atoms, charge and multiplicity.
resources on the Evaluator (a Resources: ranks, memory
per rank) sets the size of each calculation on the batch path; the provider’s
get_task receives it. Without it the provider chooses.
A program offers the batch path through three classmethods beside
get_model_chemistry_options: get_task, analyze_task and, optionally,
can_run_task; see seamm_exec.evaluator.
The sign of the stress. The stress comes back as the program gives it, in
GPa, and programs differ in its sign. A provider that returns a stress must
therefore declare its convention in its get_model_chemistry_options entry:
options["stress_convention"] is "pressure" (positive when the system
pushes outward: VASP’s in kB line, MDI’s <STRESS) or "stress"
(sigma = -P, as ASE and xnn use). The Evaluator passes the stress through
unchanged; a consumer converts it with that declaration and refuses a level
that lacks it. The declaration covers both paths, batch and MDI, so a provider
whose two paths would give opposite signs must make them agree. Pinning one convention for every provider later would be a
documented change of this contract.
Where things run#
Tasks go to the job’s target, a section of the JobServer’s
<root>/<jobserver-name>.ini (see seamm_scheduler.config). The JobServer
writes the job’s section into the job directory as target.json; for a run by
hand, SEAMM_TARGET=<section> (with SEAMM_TARGETS=<ini file> if it is not
<root>/<hostname>.ini) does the same. Without a target, or with
tasks = pool, tasks run on this machine or inside the current allocation, as
before.
With tasks = queue the TaskSet submits them, in bundles, as batch jobs:
[arc]
type = local
tasks = queue
scheduler = slurm
transport = ssh
host = tinkercliffs
remote_root = /projects/seamm/psaxe/tasks ; task directories, staged
remote_python = /projects/seamm/SEAMM/venv/bin/python
account = seamm
partition = normal_q
qos = tc_normal_short
export = NONE
bundle_tasks = 8 ; or bundle_walltime = 04:00:00
max_queued_tasks = 800
Each bundle is one job that runs
python -m seamm_exec.task_worker bundle.jsonon the cluster, which runs the bundle’s tasks through a pool sized to the allocation. The program is configured there, from that machine’s<root>/<program>.iniand the program’s resolver (an entry point inorg.molssi.seamm.exec.resolvers), so the code need not be installed where the evaluator runs.Without a shared filesystem the task directories are copied to
remote_rootwithrsyncbefore the job and back after it; withshared_filesystem = yes(or the local transport) nothing is copied.A rerun of the step polls bundles still in the queue rather than submitting them again, and a bundle that ran out of time leaves its finished tasks done.
Tasks estimated to take less than
inline_belowseconds (default 60) run on the evaluator’s machine when their program is installed there.A task that carries its own
configwas configured for the evaluator’s machine, so on an ssh target it runs there instead, with a warning. Steps that still configure their program themselves do this; ORCA (from orca_step 2026.10.3.1) and the Evaluator’s tasks name only their program.
Bundle files are in <step>/tasks/_bundles/<bundle>.<n>/: bundle.json,
run.sh and the scheduler’s log.