2026-10-02 -- Parallel execution: the task layer, schedulers and parallel loops =============================================================================== Status: design complete and all six open questions settled with Paul on 2026-10-02. Phases 0 and 1 are done and released (``NOTES_phase0.rst``, ``NOTES_phase1.rst``). Phase 2 (``seamm_scheduler``, the ``seamm_slurm`` shim, the ``SchedulerBackend``, targets and the resolver hook) is implemented and validated live from this Mac to TinkerCliffs and MolSSI10 and on TinkerCliffs alone, committed locally, not released; see ``NOTES_phase2.rst``. The canonical copy of this design is ``~/Work/SEAMM/Parallel_execution_design.rst`` at the workspace root; this is the campaign copy, to be kept in sync while the design changes and to gain ``NOTES*`` files as the work proceeds. It continues the JobServer SLURM campaign (``seamm_jobserver`` ``campaigns/2026-08-05``) and the multi-queue routing campaign (``campaigns/2026-08-10``), and it answers the execution question left open in ``MBE_correction_step_design.rst``. Purpose ======= SEAMM jobs run one step at a time, in one process, and every external code runs inline in that process. Two common patterns need more than that: 1. **Many independent calculations inside one step.** The many-body correction (MBE) step needs about 2,800 fragment calculations per periodic frame; the N-fragment counterpoise step, the Energy step over a trajectory, and the dimer-builder labelling all have the same shape. 2. **Loops whose iterations are independent.** A Loop over structures, table rows or parameters whose body may be one MOPAC step or a multi-step LAMMPS workflow, with results collected into tables or into the structure database. This document designs one mechanism that serves both, works on a disconnected laptop, on a server plus cluster pair (ChemAI + ARC), and on a cluster alone (ARC), and supports SLURM, PBS and other queueing systems as well as none. The design follows the model Paul used at Materials Design (MedeA): the **JobServer** manages jobs; a **flowchart evaluator** does the lightweight work of interpreting the flowchart, preparing inputs and analyzing outputs; and heavy codes run as **tasks** handed to whatever runs programs on the compute resource, either a very simple **TaskServer** or the site's queueing system, which keeps the state. Requirements ============ Deployment shapes ----------------- All three must work with the same flowcharts and the same plug-in code: =================== ==================================================================================== Shape Where things run =================== ==================================================================================== Laptop, offline JobServer, evaluator and tasks all on the laptop. No queue, possibly no network. Server + cluster JobServer and evaluator on the server (ChemAI). Tasks on the server's own SLURM, or on a remote cluster (ARC) over ssh. Datastore and job directories stay on the server. Cluster alone No persistent services allowed on the login nodes. The evaluator is a one-core queue job; it submits tasks to the same queue from the compute node. =================== ==================================================================================== Other requirements ------------------ - **Queueing systems:** SLURM and PBS from the start in the interface, SLURM first in code; LSF, Grid Engine and others later with no change to plug-ins. No queueing system must also work. - **Restart:** never recompute a finished calculation. A frame of 2,800 fragments or a loop of 1,000 iterations must resume after a crash, a walltime limit, or a JobServer upgrade. - **Site limits:** respect per-user queued-job caps (TinkerCliffs: 1,000, each array element counting), bundle small tasks into shared allocations, and keep the file count down (the MBE prototype hit a 10.5 M-inode quota overnight). - **Live monitoring:** running a code in place so its output can be watched (MOPAC, LAMMPS trajectories) must remain possible; it is a per-task choice, not a global one. - **Visibility:** the Dashboard should be able to show a job's tasks and their states. - **Codes need not be installed where the evaluator runs.** The evaluator may be on ChemAI while ORCA is only on ARC. - **Nesting:** a loop body may itself contain a step that fans out, or another parallel loop. What exists today ================= Facts established from the code on 2026-10-02: - ``seamm_exec.Base.run(config, cmd, directory, input_data, files, env, return_files, shell, in_situ, ce)`` already has the task shape: input files as a dict, a command template, return-file globs, a temporary working directory honouring ``$TMPDIR`` when not in situ, and copy-back of only the requested files. It is blocking, runs one command, and is given the whole allocation. - ``computational_environment()`` reports the full SLURM allocation or the whole machine. Nothing shares an allocation among concurrent tasks. ORCA sets ``%pal`` to all of ``NTASKS``. - Each code step reads its own ``/.ini`` (conda environment, modules, executable) in the evaluator process and resolves the command itself. - ``seamm_slurm`` has ``SlurmBackend`` (submit, poll_many, cancel), ``LocalSlurm`` and ``SshSlurm`` transports, ``status.classify()``, ``script.build_script()`` and ``RsyncStager``. It is already decoupled from the JobServer, which was intended to allow a per-step executor later. - ``seamm_jobserver`` submits a whole flowchart as one sbatch when a queue section is of type ``slurm``, or spawns ``run_from_jobserver`` as a local subprocess. Queue sections live in ``/.ini`` and a job's ``parameters["queue"]`` selects one. Per-queue concurrency caps and resubmission exist. - **Flowcharts do not resume.** The JobServer's resubmit-on-loss logic assumes a flowchart restarts from its first incomplete step. No such code exists in ``seamm``, ``exec_flowchart`` or ``loop_step``; a resubmitted job reruns from the top. - ``loop_step`` runs iterations sequentially in one process. Iterations share the global variables dict, the in-memory pandas tables and their ``current index``, the SystemDB's current system pointer, and the references database. Nothing checks independence and finished iterations are not skipped. - Tables exist only in memory, as ``{"type": "pandas", "table": DataFrame, ...}`` variables, until a Table step saves them. ``store_results`` writes cells with ``table.at[index, column]``. - The structure database is SQLite in the job directory (``seamm.db``, WAL mode), owned by the one evaluator process. - MDI engines are sequential by construction: one warm engine, one structure at a time. The ORCA engine runs a subprocess per evaluation in its own temporary directory, so for 27-second fragments MDI gains nothing over batch inputs. - There is no parent/child job concept in the datastore, dashboard or JobServer. Architecture ============ Roles ----- .. code-block:: text Dashboard / Tk client / MCP | submit job (flowchart, files, project, target) v JobServer -- manages jobs; spawns one evaluator per job (subprocess, as today) | v Flowchart evaluator (run_from_jobserver) -- interprets the flowchart, runs lightweight steps, | prepares inputs, analyzes outputs, writes the job database and checkpoint | tasks (program, files, resources, return globs) v Task layer (seamm_exec) -- manifest, bundling, pruning, archiving, reattachment | +--> LocalPool : in-process, partitions this machine or this allocation +--> TaskServer : tiny service on a compute machine without a queue +--> Scheduler : SLURM | PBS | ... via local commands or ssh (staging: shared filesystem | rsync over ssh | TaskServer transfer) The JobServer does not become a task broker. It keeps doing what it does: pick up submitted jobs, start an evaluator, cap concurrency per target, reattach to or resubmit lost evaluators, and record state. The task layer runs inside the evaluator. The queueing system or the TaskServer keeps the task state; the evaluator keeps only the ids it needs to reattach. How the three deployment shapes use this ---------------------------------------- - **Laptop:** the ``LocalPool`` back end. Tasks run concurrently, sized to the machine. A TaskServer on the laptop is optional and only buys independence from JobServer restarts. - **Server + cluster:** the evaluator runs on the server under the JobServer. Its target section says ``scheduler = slurm, transport = local`` for the server's own queue, or ``transport = ssh, host = arc`` for the cluster, with rsync staging of each task's directory. Only the heavy tasks leave the server. - **Cluster alone:** the JobServer (on the server, or none) submits the evaluator itself as a one-core job, which is today's whole-flowchart path. The evaluator's own target section then says ``scheduler = slurm, transport = local`` and tasks are submitted from the compute node. The prototype's feeder job proved that sbatch works from TinkerCliffs compute nodes. Because the evaluator can outlive its walltime, this shape needs flowchart-level checkpointing (below). The TaskServer -------------- A deliberately small service for a machine without a queueing system (a workstation, a second Mac, a cloud VM). Its protocol is the task API over HTTP, and nothing else: - ``POST /tasks``: program name, files, command template, resources, return globs, in-situ flag. Returns an id. - ``GET /tasks/{id}``: state (queued, running, finished, failed, cancelled), return code, listing. - ``GET /tasks/{id}/files/{name}``: a returned file. - ``DELETE /tasks/{id}``: cancel. - ``GET /programs``: the programs it knows how to run, from its own ``.ini`` files. It owns the code configuration for its machine and a small pool for concurrency. It persists its queue to a local SQLite file so it survives its own restart. Because it changes rarely, the JobServer and plug-ins can be upgraded without losing running tasks. **Transport and trust:** the TaskServer binds to loopback (``127.0.0.1``) only. A remote evaluator reaches it through an ssh port-forward using the same passwordless ssh the scheduler transport already needs, so trust is the user's ssh key and there are no new credentials, tokens or certificates to manage. The ``url`` of a ``tasks = taskserver`` target is therefore the local end of that tunnel, and the back end opens the tunnel itself (``ssh -N -L``) when the target has a ``host``. Resource vocabulary ------------------- Tasks state resources in scheduler-neutral terms; each back end translates: ``ntasks``, ``cpus_per_task``, ``mem_per_cpu``, ``ngpus``, ``walltime``, ``partition``, ``account``, ``qos``, ``nodes``. A code step fills these from its own knowledge (ORCA: ``ntasks = 4``, ``mem_per_cpu = 2 GB``) or from the step's parameters, and may read site defaults from the target section. Within a task the code sees its own allocation through the existing ``computational_environment()`` because the back end runs it under the scheduler, or sets the equivalent variables for the pool. The task API (``seamm_exec``) ============================= Objects ------- .. code-block:: python @dataclass class Task: key: str # unique within the step, stable across restarts (e.g. fragment key) program: str # "orca", "mopac", "vasp", "run_flowchart", ... cmd: list[str] # template; {code}, {code_dir}, {NTASKS}, ... filled by the back end files: dict[str, str | bytes] return_files: list[str] # globs; may include "@subdir+pattern" as today resources: Resources # ntasks=None means the whole of the back end's capacity env: dict[str, str] = {} in_situ: bool | None = None # True = run and leave output in the task directory for watching shell: bool = False input_data: str | None = None # stdin, as Base.run() takes estimated_seconds: float | None = None # plug-in's cost estimate; drives the inline rule target: str | None = None # None = the job's target (reserved; no per-step override today) directory: Path | None = None # None = /tasks// (fan-out); a single-calculation # step passes its own directory so its output stays put config: dict | None = None # local override of the .ini section; remote back ends ignore it fingerprint: str | None = None # restart identity; None hashes cmd + files (never env) success_text: dict[str, str | list[str]] | None = None # each file must contain its text(s), for codes such as ORCA # that exit 0 after an error termination @dataclass class TaskResult: key: str state: str # finished | failed | cancelled | lost returncode: int | None stdout: str; stderr: str directory: Path | None # Task.directory, or /tasks// files: dict[str, bytes | str] attempts: int; history: list[dict] # over all runs, each with its reason restored: bool # from an earlier run's DONE, not recomputed reason: str | None # why it failed: return code, success check, lost, ... class TaskBackend(Protocol): name: str def submit(self, tasks: list[Task], directories: list[Path], on_start=None) -> list[str]: ... # backend ids; on_start(task, info) per start def status(self, ids: list[str]) -> dict[str, str]: ... def cancel(self, ids: list[str]) -> None: ... def fetch(self, task: Task, backend_id: str) -> TaskResult: ... # optional: wait(ids, timeout), reattach(records) -> {key: state}, capacity(), has_program(task) # a back end that bundles (the SchedulerBackend, phase 2) also has: # bundles = True -> one submit() per bundle, with bundle=, markers=[tasks/] # room() -> int | None -> how many more bundles the queue takes now (max_queued_tasks) # adopt(task, directory, marker, record) -> id | None -> poll it after a restart # reason(id) -> str -> why a task was lost (job state, the tail of its log) # accepts_config: bool -> False on an ssh target: Task.config stays on this machine class TaskSet: """What a step uses: submit many, wait, iterate results as they finish.""" def __init__(self, node, target=None, *, archive=False, bundle_tasks=None, max_attempts=3, max_lost_retries=2, inline_below=60.0, ...): ... def add(self, task: Task) -> None: ... def capacity(self) -> dict: ... # {"cores", "memory", "ngpus"}, to size tasks first def run(self) -> Iterator[TaskResult]: ... # submits what is not done, polls, yields results def summary(self) -> dict: ... def run_task(task, node=None, **kwargs) -> TaskResult: ... # a one-task TaskSet, for single calculations ``Base.run()`` stays for backward compatibility and becomes ``TaskSet`` with one task and a synchronous ``LocalPool`` with one slot and no manifest; existing steps keep working unchanged and their directories gain no new files. ``/tasks//`` is the default task directory only for fan-out: a step running one calculation gives ``Task.directory`` as its own directory, and only its bookkeeping (``tasks/manifest.json``, ``tasks//DONE``) is new. The manifest and restart ------------------------ Each step that uses tasks gets ``/tasks/manifest.json`` recording, per key: the backend, its id, the state, timestamps and the attempt count, plus ``/tasks//DONE`` on completion. On ``TaskSet.run()``: 1. keys with ``DONE`` whose fingerprint matches are yielded from their stored result and never resubmitted (a changed fingerprint is rerun with a warning); 2. keys with a live backend id are reattached by polling, not resubmitted. The ``LocalPool`` cannot adopt another process's children, so it kills a leftover process group it finds and reruns the task; 3. everything else is submitted. Within one run a failed task (nonzero return code, or a failed ``success_text`` check) is never retried and a lost one is retried up to twice; across runs a failed or lost task is eligible again until it has had ``max_attempts`` (default 3) attempts in all, after which it is reported failed ("attempts exhausted") with its attempt history. New inputs (a changed fingerprint) reset the count; deleting the task's entry in the manifest does too. This is the same trust-the-record pattern the JobServer uses for jobs, and it gives fragment- and iteration-level restart for free. Target and the inline rule for tiny tasks ----------------------------------------- Every task goes to the **job's target**, chosen at submission (``parameters["queue"]``). There is no per-step target override; a flowchart that needs two targets is split into two jobs. ``Task.target`` is reserved so an override can be added later without changing plug-ins. The one exception is the **inline rule**, which catches the really bad cases such as submitting a 100 ms MOPAC run to SLURM: - the plug-in supplies a *cost estimate*, not a routing decision: ``Task.estimated_seconds`` computed from what the step knows (MOPAC: atom count and method; ORCA: atoms, basis and method class). A plug-in with no idea leaves it unset; - each target has a threshold, ``inline_below`` (default 60 s); tasks under it run in the evaluator's own ``LocalPool`` instead of going to the queue, **provided the program is installed where the evaluator runs** (the pool checks for the program's ini section). Otherwise the task goes to the target as usual, where bundling still groups thousands of tiny tasks into one allocation. Bundling -------- Small tasks are grouped into bundles that share one allocation. A bundle is itself a scheduler job that runs a generic worker script (shipped with ``seamm_exec``, pure Python, no SEAMM import) which executes the bundle's tasks in order, writes each task's ``DONE``, and skips tasks already done. Bundling is in the task layer, not implemented through native job arrays, because array support and limits differ between SLURM and PBS and between sites. The target section sets bundle sizes (``bundle_tasks``, ``bundle_walltime``) or the step computes them from an estimate per task, as the MBE prototype did (240 ORCA runs per 4-core job, 8 VASP fragments per 8-core job). Pruning and archiving --------------------- - ``return_files`` is the contract: nothing else comes back from a scratch run. In-situ tasks delete files not matched by ``return_files`` when they finish (``Base`` already does this). - A step may declare ``archive = True`` on a ``TaskSet``; finished task directories are then packed into ``tasks/.tar`` as their bundle completes, leaving the manifest, the archives and the step's own results. The MBE step requires this (about 14 files per frame instead of 10,000). - Code steps should list the files worth keeping explicitly (for VASP: INCAR, KPOINTS, POSCAR, OUTCAR, OSZICAR, vasprun.xml). Back ends --------- ``LocalPool`` Runs tasks as subprocesses with a slot budget computed from the machine or the current allocation (cores and memory). It reads the local ``.ini`` files exactly as the steps do today. Honour ``in_situ``. Set ``OMP_NUM_THREADS`` and the MPI binding policy per task as ``orca_step`` does now, so concurrent tasks don't pile onto the same cores. ``TaskServerClient`` Speaks the TaskServer protocol. Staging is part of the protocol (files in the POST, files back by GET). ``SchedulerBackend`` Parameterized by a scheduler module (below) and a transport (local commands or ssh), plus a stager (none on a shared filesystem, rsync over ssh otherwise). Writes one script per task or per bundle, submits it, polls many ids in one command, cancels, and fetches ``return_files`` after staging back. A target may set ``shared_filesystem = yes`` to skip staging when the evaluator and the target cluster see the same storage (TinkerCliffs, Falcon and Owl share ``/projects``). **Cross-cluster from a queued evaluator.** An evaluator that is itself a queue job may target another cluster over ssh. It is allowed, not precluded, but it is a last resort that the user sets up and tests: outbound ssh from compute nodes is site-dependent, and staging from a node whose allocation can end is fragile. The documentation says to test it from an interactive job first. The normal way to fan out across clusters is from the server shape, where the evaluator runs on ChemAI. Code configuration moves to the back end ---------------------------------------- Steps currently resolve the executable, conda environment or modules themselves from ``/.ini``. In the new model a step names the **program** and the back end resolves it on the machine where it runs: the ``LocalPool`` from the local ini files, the TaskServer from its own, the scheduler back end from the target section's ``setup`` text and the remote ini files. The command template keeps the existing ``{code}``/``{NTASKS}`` placeholders. This is the one refactor every code step shares, and it is done once in ``seamm_exec``; a step's change is limited to replacing its ``executor.run(...)`` call with a ``Task``. **As built (phase 2).** A bundle runs on the compute node as `` -m seamm_exec.task_worker bundle.json``, which pushes its tasks through a ``LocalPool`` sized to the allocation, so ``{code}``, ``{NTASKS}``, conda/modules, ``$TMPDIR`` scratch and ``return_files`` behave exactly as in the evaluator. The program is resolved there: the ``[local]`` section of ``/.ini`` on that machine, then the program's **resolver**, an entry point in the group ``org.molssi.seamm.exec.resolvers`` named after the program, ``hook(config, cmd, env, ce, root) -> (config, cmd, env)``, which adjusts the configuration, command and environment for that machine and the task's share of it. ```` is the evaluator's own on the local transport and the target's ``remote_python`` on ssh. ``Task.config`` is never sent to an ssh target, because it was resolved on the evaluator's machine and names its paths: a task that carries one is kept in the evaluator's pool, with a warning, until its step names only the program (ORCA and MOPAC today; their resolvers come with ``get_task`` in phase 3). On the local transport ``Task.config`` is sent and used as the ``LocalPool`` uses it. Scheduler abstraction ===================== Generalize ``seamm_slurm`` into a new package, ``seamm_scheduler``, with one module per queueing system and a shared interface. ``seamm_slurm`` remains as a thin compatibility shim re-exporting from ``seamm_scheduler`` so ``seamm_jobserver`` keeps working until it is updated to import the new package. .. code-block:: python class Scheduler(Protocol): name: str # "slurm", "pbs", ... def directives(self, resources: Resources, extra: dict) -> dict: ... # this scheduler's directives def directive_lines(self, directives: dict) -> list[str]: ... # "#SBATCH ..." lines def submit_cmd(self, script_path) -> list[str]: ... # ["sbatch", "--parsable", ...] def parse_submit(self, stdout) -> str: ... # job id def status_cmd(self, ids) -> list[str]: ... # squeue/sacct or qstat def parse_status(self, stdout, ids) -> dict[str, str]: ... # id -> queued|running|finished|failed|lost def cancel_cmd(self, ids) -> list[str]: ... env_names: dict[str, str] # {"ntasks": "SLURM_NTASKS", ...} for computational_environment # as built (phase 2): def poll(self, run, ids) -> dict[str, JobStatus]: ... # composes the above; SLURM: squeue, then # sacct for the rest; --json probed, text fallback poll_failed: bool # the queue could not be asked: "missing" is not "gone" def count_cmd(self) -> list[str] | None: ... # the user's own jobs (squeue --me -r), for room() def log_directives(self, directory) -> dict: ... # where the job's own output goes What is SLURM-specific in ``seamm_slurm`` today is exactly this set: the directive syntax, the submit and status commands, the state vocabulary and the ``--json`` versus text parsing. The transports, the stager, the script builder and the JobServer-facing backend are already generic and move up unchanged. The JobServer's whole-flowchart submission and the task layer use the same scheduler modules, so a new queueing system is one module plus tests. ``computational_environment()`` gains the same abstraction: it asks the scheduler module for the names of its environment variables instead of hard-coding ``SLURM_*``. Configuration ============= The existing ``/.ini`` sections become **targets** that describe both where a flowchart evaluator may run and where its tasks run. New keys are additive; current files keep working. .. code-block:: ini [DEFAULT] default = local [local] type = local ; evaluator runs as a local subprocess tasks = pool ; tasks: pool | taskserver | queue max_concurrent_jobs = 4 [chemai] type = local ; evaluator on this machine tasks = queue ; tasks go to this machine's SLURM scheduler = slurm transport = local partition = normal bundle_tasks = 50 [arc] type = local ; evaluator still on ChemAI tasks = queue ; tasks on ARC scheduler = slurm transport = ssh host = tinkercliffs remote_root = /projects/seamm/psaxe/tasks account = seamm partition = normal_q max_queued_tasks = 800 setup = module load ORCA/6.1.1 remote_python = /projects/seamm/SEAMM/venv/bin/python ; runs the tasks there (phase 2) export = NONE ; a login environment for jobs submitted over ssh [arc-all] type = slurm ; evaluator itself is a one-core job on ARC (today's path) transport = ssh host = tinkercliffs tasks = queue ; and it submits tasks locally from the compute node scheduler = slurm [workstation] type = local tasks = taskserver url = https://workstation.local:5500 **Task keys, as built (phase 2).** All optional; a section without ``tasks =`` means what it always did. ``tasks`` (``pool``, ``queue``, ``taskserver``), ``scheduler`` (``slurm`` default, ``pbs``), ``shared_filesystem`` (default yes for the local task transport, no for ssh), ``bundle_tasks``, ``bundle_walltime`` (SLURM time syntax; also a bundle's ``--time`` when its tasks give none), ``max_queued_tasks``, ``inline_below`` (default 60 s), ``poll_interval`` (default 30 s), ``remote_root`` (where task directories are staged), ``remote_python`` (ssh: a Python with ``seamm_exec`` on the cluster), ``remote_seamm_root`` (its ``.ini`` files; default the venv's root) and ``url`` (TaskServer). The section's directive keys (``partition``, ``account``, ``qos``, ``export``, ...) are the bundles' site defaults; each bundle's resources override them, and ``setup`` runs before the worker. For ``type = slurm`` (the evaluator is itself a batch job) tasks always use the local transport and a shared filesystem, since the evaluator is inside the cluster. On ``[arc]`` above, ``remote_python`` is required and ``export = NONE`` is what makes ``module load`` work in jobs submitted over ssh. **How the evaluator finds its target.** The JobServer writes the job's section as ``/target.json`` when it starts the job (before staging), so an evaluator that is itself a batch job on a cluster that cannot read the JobServer's ini file still finds it. The evaluator takes, in order: an explicit target, ``target.json``, ``$SEAMM_TARGET`` (a section of ``/.ini`` or of ``$SEAMM_TARGETS``, for runs by hand), else its own ``LocalPool``. ``target.json`` is a copy of the section, ``setup`` text included, in a directory the Dashboard shows, so a section must never hold secrets (none does today). A job's ``parameters["queue"]`` already selects a section; its meaning becomes "this job's target". The Tk submit dialog's queue picker and the ``GET /api/queues`` route need no conceptual change. A step may override the target for its own tasks through a parameter when that is genuinely needed (a cheap preprocessing step staying local while the heavy step goes to the cluster), but the default is the job's target. Model Chemistry batch contract ============================== MDI remains the right interface for a warm engine on one machine. Farming out needs the prepare/analyze split. Add to the provider interface, next to ``get_model_chemistry_options()`` and ``get_mdi_engine_command()``: .. code-block:: python # classmethods of the program's plug-in class, beside get_model_chemistry_options @classmethod def get_task(cls, configuration, model_chemistry, *, key, properties=("energy", "gradients"), options=None, resources=None) -> Task: ... @classmethod def analyze_task(cls, result: TaskResult, model_chemistry, configuration, *, properties=("energy", "gradients"), options=None) -> dict: ... # {"energy": kJ/mol, "gradients": (n, 3) kJ/mol/Å, "stress": GPa as the program gives it} @classmethod def can_run_task(cls, configuration, model_chemistry, *, options=None) -> bool: ... # optional ``options`` carries what fragments need: ``atom_indices``, ``ghost_atoms``, ``charge``, ``multiplicity`` (enough for ``seamm_bsse``'s N-fragment counterpoise jobs). A structure for which ``can_run_task`` is False (a periodic system for ORCA or MOPAC, an open-shell one for MOPAC) goes to the program's MDI engine, if there is one here, even when the rest run as tasks; otherwise it fails alone. ``analyze_task`` raises a clear error, never returns partial numbers, when a result lacks a requested property (the stress is required only for a periodic structure). For a given model chemistry the batch path and the MDI path return the same numbers (for MOPAC the "energy" is the heat of formation, as its MDI engine reports it; Paul, 2026-10-03), and a test of ORCA and MOPAC water protects that to a stated tolerance. Consumers (Energy, MBE, N-fragment counterpoise, dimer builder, the finite-difference Hessian of normal-mode sampling) use a facade in ``seamm_exec`` with the same calls as ``seamm_mdi.MDIEngine`` but asynchronous: ``submit(configuration) -> key`` and ``results() -> iterator``. It lives in ``seamm_exec``, not ``seamm``, because ``seamm_exec`` already depends on ``seamm`` (a facade in ``seamm`` would need lazy imports to dodge the circle) and how things run is ``seamm_exec``'s job; ``seamm_mdi`` is imported only on the MDI path, so a machine without pymdi still runs the batch path. **The facade, not the user or the step, chooses the path:** the batch path when the job's target sends tasks to a queue (or a TaskServer) and the provider has ``get_task``; MDI when the tasks stay local and the provider has an MDI engine (so MLFF, MOPAC and xTB stay at milliseconds per structure), unless the provider declares ``prefers_batch`` in its options (ORCA does: its MDI engine runs a subprocess per evaluation, so the pool's concurrency beats a sequential warm engine for many structures); otherwise the batch path in the local pool. A ``path=`` argument exists for tests only. Energy, MBE and the counterpoise code inherit the rule with no per-step logic, and the Energy step keeps its MDI path. **The Hessian for normal-mode sampling:** on a queue target the target wins -- the finite difference as tasks on the cluster, with no local engine started even to ask; otherwise the analytic Hessian over MDI when the program has one (its ``analytic_hessian`` option, else the engine's ``.ini`` files are still read where the code runs. - **Old job directories stay readable.** New files (``tasks/manifest.json``, ``checkpoint.json``, the ``_tables`` registry, and ``/target.json``, written by the JobServer only for a target with the task keys) are added beside the existing ones; nothing existing is renamed or removed. - **Unconverted plug-ins behave byte for byte as today.** ``Base.run()`` is reimplemented on the task layer with a one-slot ``LocalPool``; a plug-in that has not been converted cannot tell the difference. - **Tables in the database cut over in one release** (revised 2026-10-03; originally a switch). Its rollback is the phase 0 venv rollback, ``seamm-manager environment rollback``, and its soak and comparison run in ``~/SEAMM_DEV`` before release. A job started with the new version and then rolled back is simply rerun; old job directories are unaffected either way. Three rules ----------- 1. **Old code paths stay; new ones are opt-in; defaults are unchanged.** See the promise above. Each new capability is reachable only through a new key, parameter or method. 2. **Running jobs never see a mixed environment.** A job imports from the venv it started in, and SEAMM imports lazily in places, so a venv updated in place while jobs run can mix versions. Rule: never update a venv in place while jobs run. With uv the venv is cheap: ``seamm-manager`` builds the new one beside the old, switches a symlink, and deletes the old venv only when its jobs have finished. Restarting the JobServer is already safe for local jobs (it reattaches by pid); phase 5 makes it safe for everything. 3. **Validate in a separate installation against production output.** ``~/SEAMM_DEV`` (own venv, JobServer, datastore) sits beside production. The gate for every phase is the same: run the Testing flowcharts and a sample of real campaign flowcharts in both installations and diff the results files and tables. Server order: paul.local, then ChemAI only on an explicit ask between runs, then ARC (no services; a venv swap). Release order follows the shared-library rule: ``seamm_exec`` and ``seamm_scheduler`` first, pinned as minimum versions in each converted plug-in, so a PyPI install can never pair new plug-ins with an old library. Risk by phase ------------- ================ ======== ================================================================================ Phase Risk Containment ================ ======== ================================================================================ 0 preparation none Tooling and harness only (below). 1 task layer medium Touches the path every code runs through; the ``Base.run()`` shim contains it. Convert ORCA and MOPAC first, then other codes one at a time on their own merits. 2 scheduler pkg low ``seamm_slurm`` becomes a shim; the JobServer does not change. Validate from ``~/SEAMM_DEV`` to MolSSI10 and TinkerCliffs. 3 batch + MBE low New provider methods, a new step; Energy keeps MDI as its default. 4 tables in DB high Table step, ``store_results`` and every table-reading step, converted to an explicit Table API. The gate runs the whole local and Dropbox flowchart corpus through the comparison harness, with a soak period in ``~/SEAMM_DEV`` before the release; rollback is the venv rollback. 5 checkpointing low New files only; resume runs only when the JobServer asks for it on a new job. The node idempotence audit is the real work. 6 parallel Loop none Opt-in parameter. 7 TaskServer/PBS none New and opt-in. ================ ======== ================================================================================ Phases ====== 0. **Preparation.** The venv swap-not-overwrite rule in ``seamm-manager`` (build beside, switch a symlink, delete when idle); a comparison harness that runs a flowchart in two installations and diffs results files and tables; and the compatibility promise above recorded in the campaign doc. 1. **Task layer in ``seamm_exec``** with ``LocalPool``, the manifest, bundling, pruning and archiving, and ``Base.run()`` reimplemented on it. Convert ``orca_step`` and ``mopac_step`` to ``Task``. Test on the laptop with concurrent MOPAC and ORCA runs; verify restart by killing mid-run. 2. **Scheduler abstraction and the SLURM back end.** Refactor ``seamm_slurm`` into scheduler modules used by both the JobServer and the task layer; local and ssh transports with rsync staging; bundled worker script. Validate ORCA fragments from this Mac to TinkerCliffs and from ChemAI to its own SLURM. Start the PBS module as a compile-time proof of the interface, tested against a mocked ``qstat``. 3. **Model Chemistry batch contract** for ORCA and MOPAC, the asynchronous facade, and the Energy step using it. Then the MBE step (its Phase 3) on this. 4. **Tables in the database** behind an explicit Table API; ``store_results`` and the Table, Loop, Properties and Geometry Analysis steps converted; one-release cutover. 5. **Flowchart-level checkpointing** in ``exec_flowchart``, ``Node.run()`` and the Loop; node idempotence audit of the code steps; JobServer resubmission validated for real on ARC-alone. 6. **Parallel Loop** with the snapshot, merge and placement options; validate on the laptop (pool), on ChemAI + ARC, and on ARC alone. 7. **TaskServer** and its client; **PBS** validated on a real PBS site when one is available; Dashboard task view. *PBS done 2026-10-03:* MolSSI10 (test-only) replaced SLURM with OpenPBS 23.06.06; the PBS back end was fixed against it (``-W depend``, ``cd`` into the job directory, ``-V``, ``qselect -x``) and its output recorded as test fixtures, as was SLURM 20.11's before removal; JobServer sections gained ``type = queue`` + ``scheduler = pbs|slurm``; a flowchart ran as a PBS job, and with its calculations as PBS tasks, from SEAMM_DEV over ssh; after the review fixes, job 4008 (the validation record) ran its bundles with the queue's ``select`` memory and brought ``pbs.out`` back. See seamm_scheduler ``docs/developer_guide/campaigns/2026-10-03/``. (OpenPBS's last release is 2023; it stands in for PBS Professional sites.) Phase 0 protects every later phase. Phases 1 to 3 unblock the MBE step with low production exposure; 4 and 5 are prerequisites of 6; 4 is the one that needs a soak period before it is released. Decisions (2026-10-02) ====================== Settled with Paul, in the order the open questions were discussed: - The JobServer stays a job manager; the task layer lives in the evaluator (MedeA model). - The queueing system or the TaskServer keeps task state; the evaluator keeps ids in a manifest. - Steps request a program by name plus scheduler-neutral resources; the back end owns code configuration. - Bundling is generic (worker script), not native job arrays. - Tables move into the job database behind an explicit Table API (revised 2026-10-03 from a DataFrame facade); SQLite stays, one writer per file, results from children return by staging and merge. No multi-writer DBMS. - Batch contract on the Model Chemistry providers; MDI kept for local warm engines, chosen by the facade. - Task-level restart first, flowchart-level checkpointing second. - Parallel Loop iterations are evaluator tasks with a declared independence contract. - **Q1 package:** new ``seamm_scheduler``; ``seamm_slurm`` becomes a compatibility shim. - **Q2 cross-cluster from a queued evaluator:** allowed as a last resort, the user's responsibility to set up and test; ``shared_filesystem`` flag skips staging on sites like ARC. - **Q3 per-step target:** no override; the job's target for everything, except the plug-in cost-estimate inline rule for tiny tasks. ``Task.target`` reserved for later. - **Q4 TaskServer:** loopback only, reached through an ssh port-forward. - **Q5 iteration snapshot:** only the selected configurations by default, with a whole-database option. - **Q6 Energy step:** keeps MDI; the facade picks MDI versus batch. Remaining questions =================== None blocking. Items to settle during implementation: 1. The exact ``estimated_seconds`` heuristics per code, and whether the threshold should also consider the queue's current wait. 2. The merge policy when two iterations write the same table index (error, or last wins with a warning). 3. Whether the Dashboard shows child iterations as rows (``parent_id`` in ``parameters``) in the first release or only the per-step task counts. Notes ===== .. toctree:: :glob: :maxdepth: 1 NOTES*