2026-10-02 – Parallel execution: the task layer, schedulers and parallel loops#

Status: design complete and all six open questions settled with Paul on 2026-10-02. Phases 0 and 1 are done and released (NOTES_phase0.rst, NOTES_phase1.rst). Phase 2 (seamm_scheduler, the seamm_slurm shim, the SchedulerBackend, targets and the resolver hook) is implemented and validated live from this Mac to TinkerCliffs and MolSSI10 and on TinkerCliffs alone, committed locally, not released; see NOTES_phase2.rst. The canonical copy of this design is ~/Work/SEAMM/Parallel_execution_design.rst at the workspace root; this is the campaign copy, to be kept in sync while the design changes and to gain NOTES* files as the work proceeds. It continues the JobServer SLURM campaign (seamm_jobserver campaigns/2026-08-05) and the multi-queue routing campaign (campaigns/2026-08-10), and it answers the execution question left open in MBE_correction_step_design.rst.

Purpose#

SEAMM jobs run one step at a time, in one process, and every external code runs inline in that process. Two common patterns need more than that:

  1. Many independent calculations inside one step. The many-body correction (MBE) step needs about 2,800 fragment calculations per periodic frame; the N-fragment counterpoise step, the Energy step over a trajectory, and the dimer-builder labelling all have the same shape.

  2. Loops whose iterations are independent. A Loop over structures, table rows or parameters whose body may be one MOPAC step or a multi-step LAMMPS workflow, with results collected into tables or into the structure database.

This document designs one mechanism that serves both, works on a disconnected laptop, on a server plus cluster pair (ChemAI + ARC), and on a cluster alone (ARC), and supports SLURM, PBS and other queueing systems as well as none.

The design follows the model Paul used at Materials Design (MedeA): the JobServer manages jobs; a flowchart evaluator does the lightweight work of interpreting the flowchart, preparing inputs and analyzing outputs; and heavy codes run as tasks handed to whatever runs programs on the compute resource, either a very simple TaskServer or the site’s queueing system, which keeps the state.

Requirements#

Deployment shapes#

All three must work with the same flowcharts and the same plug-in code:

Shape

Where things run

Laptop, offline

JobServer, evaluator and tasks all on the laptop. No queue, possibly no network.

Server + cluster

JobServer and evaluator on the server (ChemAI). Tasks on the server’s own SLURM, or on a remote cluster (ARC) over ssh. Datastore and job directories stay on the server.

Cluster alone

No persistent services allowed on the login nodes. The evaluator is a one-core queue job; it submits tasks to the same queue from the compute node.

Other requirements#

  • Queueing systems: SLURM and PBS from the start in the interface, SLURM first in code; LSF, Grid Engine and others later with no change to plug-ins. No queueing system must also work.

  • Restart: never recompute a finished calculation. A frame of 2,800 fragments or a loop of 1,000 iterations must resume after a crash, a walltime limit, or a JobServer upgrade.

  • Site limits: respect per-user queued-job caps (TinkerCliffs: 1,000, each array element counting), bundle small tasks into shared allocations, and keep the file count down (the MBE prototype hit a 10.5 M-inode quota overnight).

  • Live monitoring: running a code in place so its output can be watched (MOPAC, LAMMPS trajectories) must remain possible; it is a per-task choice, not a global one.

  • Visibility: the Dashboard should be able to show a job’s tasks and their states.

  • Codes need not be installed where the evaluator runs. The evaluator may be on ChemAI while ORCA is only on ARC.

  • Nesting: a loop body may itself contain a step that fans out, or another parallel loop.

What exists today#

Facts established from the code on 2026-10-02:

  • seamm_exec.Base.run(config, cmd, directory, input_data, files, env, return_files, shell, in_situ, ce) already has the task shape: input files as a dict, a command template, return-file globs, a temporary working directory honouring $TMPDIR when not in situ, and copy-back of only the requested files. It is blocking, runs one command, and is given the whole allocation.

  • computational_environment() reports the full SLURM allocation or the whole machine. Nothing shares an allocation among concurrent tasks. ORCA sets %pal to all of NTASKS.

  • Each code step reads its own <root>/<code>.ini (conda environment, modules, executable) in the evaluator process and resolves the command itself.

  • seamm_slurm has SlurmBackend (submit, poll_many, cancel), LocalSlurm and SshSlurm transports, status.classify(), script.build_script() and RsyncStager. It is already decoupled from the JobServer, which was intended to allow a per-step executor later.

  • seamm_jobserver submits a whole flowchart as one sbatch when a queue section is of type slurm, or spawns run_from_jobserver as a local subprocess. Queue sections live in <root>/<jobserver-name>.ini and a job’s parameters["queue"] selects one. Per-queue concurrency caps and resubmission exist.

  • Flowcharts do not resume. The JobServer’s resubmit-on-loss logic assumes a flowchart restarts from its first incomplete step. No such code exists in seamm, exec_flowchart or loop_step; a resubmitted job reruns from the top.

  • loop_step runs iterations sequentially in one process. Iterations share the global variables dict, the in-memory pandas tables and their current index, the SystemDB’s current system pointer, and the references database. Nothing checks independence and finished iterations are not skipped.

  • Tables exist only in memory, as {"type": "pandas", "table": DataFrame, ...} variables, until a Table step saves them. store_results writes cells with table.at[index, column].

  • The structure database is SQLite in the job directory (seamm.db, WAL mode), owned by the one evaluator process.

  • MDI engines are sequential by construction: one warm engine, one structure at a time. The ORCA engine runs a subprocess per evaluation in its own temporary directory, so for 27-second fragments MDI gains nothing over batch inputs.

  • There is no parent/child job concept in the datastore, dashboard or JobServer.

Architecture#

Roles#

Dashboard / Tk client / MCP
        |  submit job (flowchart, files, project, target)
        v
JobServer  -- manages jobs; spawns one evaluator per job (subprocess, as today)
        |
        v
Flowchart evaluator (run_from_jobserver) -- interprets the flowchart, runs lightweight steps,
        |     prepares inputs, analyzes outputs, writes the job database and checkpoint
        |  tasks (program, files, resources, return globs)
        v
Task layer (seamm_exec) -- manifest, bundling, pruning, archiving, reattachment
        |
        +--> LocalPool      : in-process, partitions this machine or this allocation
        +--> TaskServer     : tiny service on a compute machine without a queue
        +--> Scheduler      : SLURM | PBS | ... via local commands or ssh
                 (staging: shared filesystem | rsync over ssh | TaskServer transfer)

The JobServer does not become a task broker. It keeps doing what it does: pick up submitted jobs, start an evaluator, cap concurrency per target, reattach to or resubmit lost evaluators, and record state. The task layer runs inside the evaluator. The queueing system or the TaskServer keeps the task state; the evaluator keeps only the ids it needs to reattach.

How the three deployment shapes use this#

  • Laptop: the LocalPool back end. Tasks run concurrently, sized to the machine. A TaskServer on the laptop is optional and only buys independence from JobServer restarts.

  • Server + cluster: the evaluator runs on the server under the JobServer. Its target section says scheduler = slurm, transport = local for the server’s own queue, or transport = ssh, host = arc for the cluster, with rsync staging of each task’s directory. Only the heavy tasks leave the server.

  • Cluster alone: the JobServer (on the server, or none) submits the evaluator itself as a one-core job, which is today’s whole-flowchart path. The evaluator’s own target section then says scheduler = slurm, transport = local and tasks are submitted from the compute node. The prototype’s feeder job proved that sbatch works from TinkerCliffs compute nodes. Because the evaluator can outlive its walltime, this shape needs flowchart-level checkpointing (below).

The TaskServer#

A deliberately small service for a machine without a queueing system (a workstation, a second Mac, a cloud VM). Its protocol is the task API over HTTP, and nothing else:

  • POST /tasks: program name, files, command template, resources, return globs, in-situ flag. Returns an id.

  • GET /tasks/{id}: state (queued, running, finished, failed, cancelled), return code, listing.

  • GET /tasks/{id}/files/{name}: a returned file.

  • DELETE /tasks/{id}: cancel.

  • GET /programs: the programs it knows how to run, from its own <code>.ini files.

It owns the code configuration for its machine and a small pool for concurrency. It persists its queue to a local SQLite file so it survives its own restart. Because it changes rarely, the JobServer and plug-ins can be upgraded without losing running tasks.

Transport and trust: the TaskServer binds to loopback (127.0.0.1) only. A remote evaluator reaches it through an ssh port-forward using the same passwordless ssh the scheduler transport already needs, so trust is the user’s ssh key and there are no new credentials, tokens or certificates to manage. The url of a tasks = taskserver target is therefore the local end of that tunnel, and the back end opens the tunnel itself (ssh -N -L) when the target has a host.

Resource vocabulary#

Tasks state resources in scheduler-neutral terms; each back end translates:

ntasks, cpus_per_task, mem_per_cpu, ngpus, walltime, partition, account, qos, nodes. A code step fills these from its own knowledge (ORCA: ntasks = 4, mem_per_cpu = 2 GB) or from the step’s parameters, and may read site defaults from the target section. Within a task the code sees its own allocation through the existing computational_environment() because the back end runs it under the scheduler, or sets the equivalent variables for the pool.

The task API (seamm_exec)#

Objects#

@dataclass
class Task:
    key: str                    # unique within the step, stable across restarts (e.g. fragment key)
    program: str                # "orca", "mopac", "vasp", "run_flowchart", ...
    cmd: list[str]              # template; {code}, {code_dir}, {NTASKS}, ... filled by the back end
    files: dict[str, str | bytes]
    return_files: list[str]     # globs; may include "@subdir+pattern" as today
    resources: Resources        # ntasks=None means the whole of the back end's capacity
    env: dict[str, str] = {}
    in_situ: bool | None = None # True = run and leave output in the task directory for watching
    shell: bool = False
    input_data: str | None = None            # stdin, as Base.run() takes
    estimated_seconds: float | None = None   # plug-in's cost estimate; drives the inline rule
    target: str | None = None   # None = the job's target (reserved; no per-step override today)
    directory: Path | None = None  # None = <step dir>/tasks/<key>/ (fan-out); a single-calculation
                                   # step passes its own directory so its output stays put
    config: dict | None = None  # local override of the <program>.ini section; remote back ends ignore it
    fingerprint: str | None = None # restart identity; None hashes cmd + files (never env)
    success_text: dict[str, str | list[str]] | None = None
                                # each file must contain its text(s), for codes such as ORCA
                                # that exit 0 after an error termination

@dataclass
class TaskResult:
    key: str
    state: str                  # finished | failed | cancelled | lost
    returncode: int | None
    stdout: str; stderr: str
    directory: Path | None      # Task.directory, or <step dir>/tasks/<key>/
    files: dict[str, bytes | str]
    attempts: int; history: list[dict]   # over all runs, each with its reason
    restored: bool              # from an earlier run's DONE, not recomputed
    reason: str | None          # why it failed: return code, success check, lost, ...

class TaskBackend(Protocol):
    name: str
    def submit(self, tasks: list[Task], directories: list[Path],
               on_start=None) -> list[str]: ...  # backend ids; on_start(task, info) per start
    def status(self, ids: list[str]) -> dict[str, str]: ...
    def cancel(self, ids: list[str]) -> None: ...
    def fetch(self, task: Task, backend_id: str) -> TaskResult: ...
    # optional: wait(ids, timeout), reattach(records) -> {key: state}, capacity(), has_program(task)
    # a back end that bundles (the SchedulerBackend, phase 2) also has:
    #   bundles = True        -> one submit() per bundle, with bundle=<name>, markers=[tasks/<key>]
    #   room() -> int | None  -> how many more bundles the queue takes now (max_queued_tasks)
    #   adopt(task, directory, marker, record) -> id | None   -> poll it after a restart
    #   reason(id) -> str     -> why a task was lost (job state, the tail of its log)
    #   accepts_config: bool  -> False on an ssh target: Task.config stays on this machine

class TaskSet:
    """What a step uses: submit many, wait, iterate results as they finish."""
    def __init__(self, node, target=None, *, archive=False, bundle_tasks=None,
                 max_attempts=3, max_lost_retries=2, inline_below=60.0, ...): ...
    def add(self, task: Task) -> None: ...
    def capacity(self) -> dict: ...                 # {"cores", "memory", "ngpus"}, to size tasks first
    def run(self) -> Iterator[TaskResult]: ...      # submits what is not done, polls, yields results
    def summary(self) -> dict: ...

def run_task(task, node=None, **kwargs) -> TaskResult: ...   # a one-task TaskSet, for single calculations

Base.run() stays for backward compatibility and becomes TaskSet with one task and a synchronous LocalPool with one slot and no manifest; existing steps keep working unchanged and their directories gain no new files. <step dir>/tasks/<key>/ is the default task directory only for fan-out: a step running one calculation gives Task.directory as its own directory, and only its bookkeeping (tasks/manifest.json, tasks/<key>/DONE) is new.

The manifest and restart#

Each step that uses tasks gets <step dir>/tasks/manifest.json recording, per key: the backend, its id, the state, timestamps and the attempt count, plus <step dir>/tasks/<key>/DONE on completion. On TaskSet.run():

  1. keys with DONE whose fingerprint matches are yielded from their stored result and never resubmitted (a changed fingerprint is rerun with a warning);

  2. keys with a live backend id are reattached by polling, not resubmitted. The LocalPool cannot adopt another process’s children, so it kills a leftover process group it finds and reruns the task;

  3. everything else is submitted. Within one run a failed task (nonzero return code, or a failed success_text check) is never retried and a lost one is retried up to twice; across runs a failed or lost task is eligible again until it has had max_attempts (default 3) attempts in all, after which it is reported failed (“attempts exhausted”) with its attempt history. New inputs (a changed fingerprint) reset the count; deleting the task’s entry in the manifest does too.

This is the same trust-the-record pattern the JobServer uses for jobs, and it gives fragment- and iteration-level restart for free.

Target and the inline rule for tiny tasks#

Every task goes to the job’s target, chosen at submission (parameters["queue"]). There is no per-step target override; a flowchart that needs two targets is split into two jobs. Task.target is reserved so an override can be added later without changing plug-ins.

The one exception is the inline rule, which catches the really bad cases such as submitting a 100 ms MOPAC run to SLURM:

  • the plug-in supplies a cost estimate, not a routing decision: Task.estimated_seconds computed from what the step knows (MOPAC: atom count and method; ORCA: atoms, basis and method class). A plug-in with no idea leaves it unset;

  • each target has a threshold, inline_below (default 60 s); tasks under it run in the evaluator’s own LocalPool instead of going to the queue, provided the program is installed where the evaluator runs (the pool checks for the program’s ini section). Otherwise the task goes to the target as usual, where bundling still groups thousands of tiny tasks into one allocation.

Bundling#

Small tasks are grouped into bundles that share one allocation. A bundle is itself a scheduler job that runs a generic worker script (shipped with seamm_exec, pure Python, no SEAMM import) which executes the bundle’s tasks in order, writes each task’s DONE, and skips tasks already done. Bundling is in the task layer, not implemented through native job arrays, because array support and limits differ between SLURM and PBS and between sites. The target section sets bundle sizes (bundle_tasks, bundle_walltime) or the step computes them from an estimate per task, as the MBE prototype did (240 ORCA runs per 4-core job, 8 VASP fragments per 8-core job).

Pruning and archiving#

  • return_files is the contract: nothing else comes back from a scratch run. In-situ tasks delete files not matched by return_files when they finish (Base already does this).

  • A step may declare archive = True on a TaskSet; finished task directories are then packed into tasks/<bundle>.tar as their bundle completes, leaving the manifest, the archives and the step’s own results. The MBE step requires this (about 14 files per frame instead of 10,000).

  • Code steps should list the files worth keeping explicitly (for VASP: INCAR, KPOINTS, POSCAR, OUTCAR, OSZICAR, vasprun.xml).

Back ends#

LocalPool

Runs tasks as subprocesses with a slot budget computed from the machine or the current allocation (cores and memory). It reads the local <code>.ini files exactly as the steps do today. Honour in_situ. Set OMP_NUM_THREADS and the MPI binding policy per task as orca_step does now, so concurrent tasks don’t pile onto the same cores.

TaskServerClient

Speaks the TaskServer protocol. Staging is part of the protocol (files in the POST, files back by GET).

SchedulerBackend

Parameterized by a scheduler module (below) and a transport (local commands or ssh), plus a stager (none on a shared filesystem, rsync over ssh otherwise). Writes one script per task or per bundle, submits it, polls many ids in one command, cancels, and fetches return_files after staging back. A target may set shared_filesystem = yes to skip staging when the evaluator and the target cluster see the same storage (TinkerCliffs, Falcon and Owl share /projects).

Cross-cluster from a queued evaluator. An evaluator that is itself a queue job may target another cluster over ssh. It is allowed, not precluded, but it is a last resort that the user sets up and tests: outbound ssh from compute nodes is site-dependent, and staging from a node whose allocation can end is fragile. The documentation says to test it from an interactive job first. The normal way to fan out across clusters is from the server shape, where the evaluator runs on ChemAI.

Code configuration moves to the back end#

Steps currently resolve the executable, conda environment or modules themselves from <root>/<code>.ini. In the new model a step names the program and the back end resolves it on the machine where it runs: the LocalPool from the local ini files, the TaskServer from its own, the scheduler back end from the target section’s setup text and the remote ini files. The command template keeps the existing {code}/{NTASKS} placeholders. This is the one refactor every code step shares, and it is done once in seamm_exec; a step’s change is limited to replacing its executor.run(...) call with a Task.

As built (phase 2). A bundle runs on the compute node as <python> -m seamm_exec.task_worker bundle.json, which pushes its tasks through a LocalPool sized to the allocation, so {code}, {NTASKS}, conda/modules, $TMPDIR scratch and return_files behave exactly as in the evaluator. The program is resolved there: the [local] section of <root>/<program>.ini on that machine, then the program’s resolver, an entry point in the group org.molssi.seamm.exec.resolvers named after the program, hook(config, cmd, env, ce, root) -> (config, cmd, env), which adjusts the configuration, command and environment for that machine and the task’s share of it. <python> is the evaluator’s own on the local transport and the target’s remote_python on ssh. Task.config is never sent to an ssh target, because it was resolved on the evaluator’s machine and names its paths: a task that carries one is kept in the evaluator’s pool, with a warning, until its step names only the program (ORCA and MOPAC today; their resolvers come with get_task in phase 3). On the local transport Task.config is sent and used as the LocalPool uses it.

Scheduler abstraction#

Generalize seamm_slurm into a new package, seamm_scheduler, with one module per queueing system and a shared interface. seamm_slurm remains as a thin compatibility shim re-exporting from seamm_scheduler so seamm_jobserver keeps working until it is updated to import the new package.

class Scheduler(Protocol):
    name: str                                   # "slurm", "pbs", ...
    def directives(self, resources: Resources, extra: dict) -> dict: ...  # this scheduler's directives
    def directive_lines(self, directives: dict) -> list[str]: ...  # "#SBATCH ..." lines
    def submit_cmd(self, script_path) -> list[str]: ...     # ["sbatch", "--parsable", ...]
    def parse_submit(self, stdout) -> str: ...              # job id
    def status_cmd(self, ids) -> list[str]: ...             # squeue/sacct or qstat
    def parse_status(self, stdout, ids) -> dict[str, str]: ...  # id -> queued|running|finished|failed|lost
    def cancel_cmd(self, ids) -> list[str]: ...
    env_names: dict[str, str]                   # {"ntasks": "SLURM_NTASKS", ...} for computational_environment
    # as built (phase 2):
    def poll(self, run, ids) -> dict[str, JobStatus]: ...  # composes the above; SLURM: squeue, then
                                                # sacct for the rest; --json probed, text fallback
    poll_failed: bool                           # the queue could not be asked: "missing" is not "gone"
    def count_cmd(self) -> list[str] | None: ...  # the user's own jobs (squeue --me -r), for room()
    def log_directives(self, directory) -> dict: ...  # where the job's own output goes

What is SLURM-specific in seamm_slurm today is exactly this set: the directive syntax, the submit and status commands, the state vocabulary and the --json versus text parsing. The transports, the stager, the script builder and the JobServer-facing backend are already generic and move up unchanged. The JobServer’s whole-flowchart submission and the task layer use the same scheduler modules, so a new queueing system is one module plus tests.

computational_environment() gains the same abstraction: it asks the scheduler module for the names of its environment variables instead of hard-coding SLURM_*.

Configuration#

The existing <root>/<jobserver-name>.ini sections become targets that describe both where a flowchart evaluator may run and where its tasks run. New keys are additive; current files keep working.

[DEFAULT]
default = local

[local]
type = local                 ; evaluator runs as a local subprocess
tasks = pool                 ; tasks: pool | taskserver | queue
max_concurrent_jobs = 4

[chemai]
type = local                 ; evaluator on this machine
tasks = queue                ; tasks go to this machine's SLURM
scheduler = slurm
transport = local
partition = normal
bundle_tasks = 50

[arc]
type = local                 ; evaluator still on ChemAI
tasks = queue                ; tasks on ARC
scheduler = slurm
transport = ssh
host = tinkercliffs
remote_root = /projects/seamm/psaxe/tasks
account = seamm
partition = normal_q
max_queued_tasks = 800
setup = module load ORCA/6.1.1
remote_python = /projects/seamm/SEAMM/venv/bin/python   ; runs the tasks there (phase 2)
export = NONE                ; a login environment for jobs submitted over ssh

[arc-all]
type = slurm                 ; evaluator itself is a one-core job on ARC (today's path)
transport = ssh
host = tinkercliffs
tasks = queue                ; and it submits tasks locally from the compute node
scheduler = slurm

[workstation]
type = local
tasks = taskserver
url = https://workstation.local:5500

Task keys, as built (phase 2). All optional; a section without tasks = means what it always did. tasks (pool, queue, taskserver), scheduler (slurm default, pbs), shared_filesystem (default yes for the local task transport, no for ssh), bundle_tasks, bundle_walltime (SLURM time syntax; also a bundle’s --time when its tasks give none), max_queued_tasks, inline_below (default 60 s), poll_interval (default 30 s), remote_root (where task directories are staged), remote_python (ssh: a Python with seamm_exec on the cluster), remote_seamm_root (its <program>.ini files; default the venv’s root) and url (TaskServer). The section’s directive keys (partition, account, qos, export, …) are the bundles’ site defaults; each bundle’s resources override them, and setup runs before the worker. For type = slurm (the evaluator is itself a batch job) tasks always use the local transport and a shared filesystem, since the evaluator is inside the cluster. On [arc] above, remote_python is required and export = NONE is what makes module load work in jobs submitted over ssh.

How the evaluator finds its target. The JobServer writes the job’s section as <job dir>/target.json when it starts the job (before staging), so an evaluator that is itself a batch job on a cluster that cannot read the JobServer’s ini file still finds it. The evaluator takes, in order: an explicit target, target.json, $SEAMM_TARGET (a section of <root>/<hostname>.ini or of $SEAMM_TARGETS, for runs by hand), else its own LocalPool. target.json is a copy of the section, setup text included, in a directory the Dashboard shows, so a section must never hold secrets (none does today).

A job’s parameters["queue"] already selects a section; its meaning becomes “this job’s target”. The Tk submit dialog’s queue picker and the GET /api/queues route need no conceptual change. A step may override the target for its own tasks through a parameter when that is genuinely needed (a cheap preprocessing step staying local while the heavy step goes to the cluster), but the default is the job’s target.

Model Chemistry batch contract#

MDI remains the right interface for a warm engine on one machine. Farming out needs the prepare/analyze split. Add to the provider interface, next to get_model_chemistry_options() and get_mdi_engine_command():

# classmethods of the program's plug-in class, beside get_model_chemistry_options
@classmethod
def get_task(cls, configuration, model_chemistry, *, key,
             properties=("energy", "gradients"), options=None, resources=None) -> Task: ...
@classmethod
def analyze_task(cls, result: TaskResult, model_chemistry, configuration, *,
                 properties=("energy", "gradients"), options=None) -> dict: ...
    # {"energy": kJ/mol, "gradients": (n, 3) kJ/mol/Å, "stress": GPa as the program gives it}
@classmethod
def can_run_task(cls, configuration, model_chemistry, *, options=None) -> bool: ...  # optional

options carries what fragments need: atom_indices, ghost_atoms, charge, multiplicity (enough for seamm_bsse’s N-fragment counterpoise jobs). A structure for which can_run_task is False (a periodic system for ORCA or MOPAC, an open-shell one for MOPAC) goes to the program’s MDI engine, if there is one here, even when the rest run as tasks; otherwise it fails alone. analyze_task raises a clear error, never returns partial numbers, when a result lacks a requested property (the stress is required only for a periodic structure). For a given model chemistry the batch path and the MDI path return the same numbers (for MOPAC the “energy” is the heat of formation, as its MDI engine reports it; Paul, 2026-10-03), and a test of ORCA and MOPAC water protects that to a stated tolerance.

Consumers (Energy, MBE, N-fragment counterpoise, dimer builder, the finite-difference Hessian of normal-mode sampling) use a facade in seamm_exec with the same calls as seamm_mdi.MDIEngine but asynchronous: submit(configuration) -> key and results() -> iterator. It lives in seamm_exec, not seamm, because seamm_exec already depends on seamm (a facade in seamm would need lazy imports to dodge the circle) and how things run is seamm_exec’s job; seamm_mdi is imported only on the MDI path, so a machine without pymdi still runs the batch path. The facade, not the user or the step, chooses the path: the batch path when the job’s target sends tasks to a queue (or a TaskServer) and the provider has get_task; MDI when the tasks stay local and the provider has an MDI engine (so MLFF, MOPAC and xTB stay at milliseconds per structure), unless the provider declares prefers_batch in its options (ORCA does: its MDI engine runs a subprocess per evaluation, so the pool’s concurrency beats a sequential warm engine for many structures); otherwise the batch path in the local pool. A path= argument exists for tests only. Energy, MBE and the counterpoise code inherit the rule with no per-step logic, and the Energy step keeps its MDI path. The Hessian for normal-mode sampling: on a queue target the target wins – the finite difference as tasks on the cluster, with no local engine started even to ask; otherwise the analytic Hessian over MDI when the program has one (its analytic_hessian option, else the engine’s <HESSIAN), else the finite difference as tasks if the evaluator chose them, or over the warm engine. Order of implementation: ORCA (its MDI engine already contains the analyze half), MOPAC (tests the whole chain on a laptop in seconds), then the VASP registered-fragment mode the MBE step needs. Options a consumer must be able to pass through: ghost atoms (counterpoise), point charges, an initial guess file, and the registered-box parameters for VASP.

Tables in the job database#

Tables move from in-memory pandas DataFrames to tables in the job’s SQLite file (seamm.db), beside the structures and properties. Reasons:

  • the job’s whole state is one file that is already on disk and already staged by rsync, so checkpointing needs no table serialization;

  • merging parallel iterations is a SQL insert keyed by iteration index, through the same path as properties;

  • the Dashboard, the web UI and the MCP server can read a job’s tables directly, and “save or lose it” goes away.

Revised 2026-10-03 (phase 4, see NOTES_phase4.rst in the campaign directory). An explicit, small SEAMM Table API (append_row(s), set_cell/get_cell, add_column, rows(where=...), current row, export, and a read-only to_dataframe() for printing and plotting), not a facade that imitates a DataFrame: the four steps that touch tables directly rely on pandas idioms (concat and replace, column assignment, .at enlargement) that a facade would have to chase. The SQLite backend is built on molsystem’s _Table, which supplies the connection and the WAL settings; user tables have prefixed SQL names, a _tables registry holds each table’s declared column types, index column and current row, and a _table_changes journal records every change for the parallel Loop’s merge. Rows have an internal id; users see the index column or the position. The Table step’s Save/Save as keep exporting CSV, Excel and JSON. Rule, unchanged: one writer per database file, which is why children never write into the parent’s file (see the Loop).

Why not a multi-writer DBMS. PostgreSQL would allow concurrent writers but is a server to install, secure and reach from every compute node, and would mean porting molsystem off the sqlite3 module. DuckDB is single-writer like SQLite. rqlite/dqlite funnel writes through one leader behind a daemon. The deciding constraint is not the engine but the network: a task or iteration on an ARC compute node has no route back to ChemAI and its allocation can end mid-write, so any shared-write design needs a network path the cluster shapes do not reliably provide. Pull-merge needs none: the child writes its own file, the file is staged back with the task, and the single writer merges it. A job also stays a self-contained directory that the dashboard reads, rsync moves and Zenodo archives. No DBMS change is on the roadmap.

Checkpointing#

Task level (first)#

The manifest and DONE markers above. Cheap, and it makes almost all of the compute safe to interrupt. This alone meets the MBE requirement that no finished fragment is recomputed.

Flowchart level (second)#

Needed for the cluster-alone shape, where the evaluator outlives its walltime, and to make the JobServer’s resubmit-on-loss assumption true. Scheme:

  • After each node completes, the evaluator writes checkpoint.json in the job directory: the id of the next node, the JSON-serializable variables, the current system and configuration ids, and the completed node ids. Tables and properties are already in the database. Non-serializable variables are listed by name and recreated by the node that made them, or the checkpoint refuses and says why.

  • The Loop node records its iteration index and the per-iteration completion in the same file, so a loop resumes at the first unfinished iteration.

  • On restart, nodes with a completion record are not run; their side effects are restored from the checkpoint. A node that was mid-flight re-enters run() and its TaskSet reattaches through the manifest.

  • Nodes must therefore be idempotent on re-entry up to their tasks: they regenerate inputs deterministically (the task key makes this checkable) and must not append to tables before their tasks finish. A short audit of the code steps is part of this phase.

  • The JobServer’s “trust job_data.json, else resubmit up to max_resubmits” path becomes correct without change.

The parallel Loop#

A parallel option on the Loop step (default off) runs each iteration as a task whose program is the flowchart evaluator:

  • Child flowchart: the loop body as a flowchart, with a start node that loads a snapshot: the variables at loop entry plus the loop variable and _loop_index, and a child seamm.db holding only the iteration’s selected configuration(s) and copies of the tables the body reads (the default). A Loop option, give iterations the whole database, copies the parent’s seamm.db for bodies that read other structures, such as pairing the current structure against others.

  • Contract (documented, user-declared, not verified): the body communicates back only through table rows, properties on its configuration(s), and files in its iteration directory. Variables set inside an iteration are not visible after the loop. Table rows written by iterations are appended in iteration order when the parent merges. This is the same contract a job array imposes.

  • Merge: the parent collects each child’s seamm.db and inserts its new table rows and properties into its own database, keyed by iteration index; files stay in iter_N/.

  • Placement: one choice per target, from the Loop’s parameters: run the body’s codes inline in the iteration’s allocation (default; right for a uniform body such as one MOPAC or one small ORCA step, with per-iteration resources set in the Loop) or as separate tasks submitted by the child (right for a body mixing a long LAMMPS run with short analyses). Iteration evaluators are small, so the task layer bundles several per allocation on a cluster.

  • Shapes: on the laptop iteration evaluators run concurrently in the pool; on ChemAI + ARC they run on ChemAI and only their heavy tasks cross to ARC; on ARC alone they are one-core queue jobs.

  • Nesting works because an iteration containing another parallel loop is a task spawning tasks.

  • Restart: each child has its own manifest and checkpoint; the parent’s manifest tracks iterations.

  • Errors: the existing errors parameter (continue / exit / raise) applies per iteration; failed iterations are reported in the loop’s results and never merged.

How the MBE step maps onto this#

  • Enumerate fragments; each fragment is a Task from the provider’s get_task(), keyed by its canonical fragment key, with resources from the level (VASP 8 ranks, 2 GB per rank; ORCA 4 ranks) and return_files restricted to the parse set.

  • Two TaskSets per frame (periodic and molecular levels), archive = True, bundle sizes from the per-task estimate. The cell calculation is a single task.

  • The step waits on both sets, then runs the increment algebra from seamm_mbe; a frame with a failed fragment is reported and never written as a label.

  • Restart comes from the manifest; a resubmitted job finishes the frame.

  • Execution model (b) of the MBE document (“fan-out”) is realized without new JobServer capability.

Dashboard visibility#

Add a tasks summary to the job’s job_data.json (counts by state per step) that the evaluator updates as its TaskSets progress, and a GET /api/jobs/{id}/tasks route in seamm_webui that reads the manifests. No datastore migration is needed. A later option is for the JobServer to register child iteration jobs as real datastore jobs with a parent_id in the parameters JSON, which would show them in the job list; not needed for the first release.

Rollout and compatibility#

This is a sweeping change, so every phase is additive, old behaviour is the default, and each phase is released on its own with a one-line rollback. Production never takes a change it did not opt into.

Compatibility promise#

Through every phase:

  • Old flowcharts run unchanged. The flowchart format does not change. The parallel Loop is a new parameter, off by default. No new migration follows the 3.0 one.

  • Old ini files keep their meaning. A JobServer target section without the new keys behaves as today; <code>.ini files are still read where the code runs.

  • Old job directories stay readable. New files (tasks/manifest.json, checkpoint.json, the _tables registry, and <job dir>/target.json, written by the JobServer only for a target with the task keys) are added beside the existing ones; nothing existing is renamed or removed.

  • Unconverted plug-ins behave byte for byte as today. Base.run() is reimplemented on the task layer with a one-slot LocalPool; a plug-in that has not been converted cannot tell the difference.

  • Tables in the database cut over in one release (revised 2026-10-03; originally a switch). Its rollback is the phase 0 venv rollback, seamm-manager environment rollback, and its soak and comparison run in ~/SEAMM_DEV before release. A job started with the new version and then rolled back is simply rerun; old job directories are unaffected either way.

Three rules#

  1. Old code paths stay; new ones are opt-in; defaults are unchanged. See the promise above. Each new capability is reachable only through a new key, parameter or method.

  2. Running jobs never see a mixed environment. A job imports from the venv it started in, and SEAMM imports lazily in places, so a venv updated in place while jobs run can mix versions. Rule: never update a venv in place while jobs run. With uv the venv is cheap: seamm-manager builds the new one beside the old, switches a symlink, and deletes the old venv only when its jobs have finished. Restarting the JobServer is already safe for local jobs (it reattaches by pid); phase 5 makes it safe for everything.

  3. Validate in a separate installation against production output. ~/SEAMM_DEV (own venv, JobServer, datastore) sits beside production. The gate for every phase is the same: run the Testing flowcharts and a sample of real campaign flowcharts in both installations and diff the results files and tables. Server order: paul.local, then ChemAI only on an explicit ask between runs, then ARC (no services; a venv swap). Release order follows the shared-library rule: seamm_exec and seamm_scheduler first, pinned as minimum versions in each converted plug-in, so a PyPI install can never pair new plug-ins with an old library.

Risk by phase#

Phase

Risk

Containment

0 preparation

none

Tooling and harness only (below).

1 task layer

medium

Touches the path every code runs through; the Base.run() shim contains it. Convert ORCA and MOPAC first, then other codes one at a time on their own merits.

2 scheduler pkg

low

seamm_slurm becomes a shim; the JobServer does not change. Validate from ~/SEAMM_DEV to MolSSI10 and TinkerCliffs.

3 batch + MBE

low

New provider methods, a new step; Energy keeps MDI as its default.

4 tables in DB

high

Table step, store_results and every table-reading step, converted to an explicit Table API. The gate runs the whole local and Dropbox flowchart corpus through the comparison harness, with a soak period in ~/SEAMM_DEV before the release; rollback is the venv rollback.

5 checkpointing

low

New files only; resume runs only when the JobServer asks for it on a new job. The node idempotence audit is the real work.

6 parallel Loop

none

Opt-in parameter.

7 TaskServer/PBS

none

New and opt-in.

Phases#

  1. Preparation. The venv swap-not-overwrite rule in seamm-manager (build beside, switch a symlink, delete when idle); a comparison harness that runs a flowchart in two installations and diffs results files and tables; and the compatibility promise above recorded in the campaign doc.

  2. Task layer in ``seamm_exec`` with LocalPool, the manifest, bundling, pruning and archiving, and Base.run() reimplemented on it. Convert orca_step and mopac_step to Task. Test on the laptop with concurrent MOPAC and ORCA runs; verify restart by killing mid-run.

  3. Scheduler abstraction and the SLURM back end. Refactor seamm_slurm into scheduler modules used by both the JobServer and the task layer; local and ssh transports with rsync staging; bundled worker script. Validate ORCA fragments from this Mac to TinkerCliffs and from ChemAI to its own SLURM. Start the PBS module as a compile-time proof of the interface, tested against a mocked qstat.

  4. Model Chemistry batch contract for ORCA and MOPAC, the asynchronous facade, and the Energy step using it. Then the MBE step (its Phase 3) on this.

  5. Tables in the database behind an explicit Table API; store_results and the Table, Loop, Properties and Geometry Analysis steps converted; one-release cutover.

  6. Flowchart-level checkpointing in exec_flowchart, Node.run() and the Loop; node idempotence audit of the code steps; JobServer resubmission validated for real on ARC-alone.

  7. Parallel Loop with the snapshot, merge and placement options; validate on the laptop (pool), on ChemAI + ARC, and on ARC alone.

  8. TaskServer and its client; PBS validated on a real PBS site when one is available; Dashboard task view. PBS done 2026-10-03: MolSSI10 (test-only) replaced SLURM with OpenPBS 23.06.06; the PBS back end was fixed against it (-W depend, cd into the job directory, -V, qselect -x) and its output recorded as test fixtures, as was SLURM 20.11’s before removal; JobServer sections gained type = queue + scheduler = pbs|slurm; a flowchart ran as a PBS job, and with its calculations as PBS tasks, from SEAMM_DEV over ssh; after the review fixes, job 4008 (the validation record) ran its bundles with the queue’s select memory and brought pbs.out back. See seamm_scheduler docs/developer_guide/campaigns/2026-10-03/. (OpenPBS’s last release is 2023; it stands in for PBS Professional sites.)

Phase 0 protects every later phase. Phases 1 to 3 unblock the MBE step with low production exposure; 4 and 5 are prerequisites of 6; 4 is the one that needs a soak period before it is released.

Decisions (2026-10-02)#

Settled with Paul, in the order the open questions were discussed:

  • The JobServer stays a job manager; the task layer lives in the evaluator (MedeA model).

  • The queueing system or the TaskServer keeps task state; the evaluator keeps ids in a manifest.

  • Steps request a program by name plus scheduler-neutral resources; the back end owns code configuration.

  • Bundling is generic (worker script), not native job arrays.

  • Tables move into the job database behind an explicit Table API (revised 2026-10-03 from a DataFrame facade); SQLite stays, one writer per file, results from children return by staging and merge. No multi-writer DBMS.

  • Batch contract on the Model Chemistry providers; MDI kept for local warm engines, chosen by the facade.

  • Task-level restart first, flowchart-level checkpointing second.

  • Parallel Loop iterations are evaluator tasks with a declared independence contract.

  • Q1 package: new seamm_scheduler; seamm_slurm becomes a compatibility shim.

  • Q2 cross-cluster from a queued evaluator: allowed as a last resort, the user’s responsibility to set up and test; shared_filesystem flag skips staging on sites like ARC.

  • Q3 per-step target: no override; the job’s target for everything, except the plug-in cost-estimate inline rule for tiny tasks. Task.target reserved for later.

  • Q4 TaskServer: loopback only, reached through an ssh port-forward.

  • Q5 iteration snapshot: only the selected configurations by default, with a whole-database option.

  • Q6 Energy step: keeps MDI; the facade picks MDI versus batch.

Remaining questions#

None blocking. Items to settle during implementation:

  1. The exact estimated_seconds heuristics per code, and whether the threshold should also consider the queue’s current wait.

  2. The merge policy when two iterations write the same table index (error, or last wins with a warning).

  3. Whether the Dashboard shows child iterations as rows (parent_id in parameters) in the first release or only the per-step task counts.

Notes#