Phase 2 notes#
2026-10-02 – the scheduler abstraction and the SLURM back end for tasks#
Done locally, on main of a new local repository, seamm_scheduler (no
GitHub repository yet), and on dev in seamm_slurm, seamm_exec and
seamm_jobserver. Committed locally; not pushed and not released.
What was built#
seamm_scheduler(new package, decision Q1)One module per queueing system behind a
Schedulerclass. Copied from the layout ofseamm_slurm: versioningit, devops workflows, docs, HISTORY, uv CI with notest_env.yaml.scheduler.py: the interface.directives(resources, extra),directive_lines,submit_cmd/parse_submit,status_cmd/parse_status,cancel_cmd,count_cmd(the user’s own queued jobs) andlog_directives.poll(run, ids)composes these. SLURM overrides it, because it needssqueueand thensacct.poll_failed(see “Outages” below) andenv_names.JobStatuswithtask_state,TASK_STATES(pending -> queued, completed -> finished, cancelled -> lost, …) andget_scheduler(name).
slurm.py: SLURM.The directive syntax and the state vocabulary.
The
--jsonprobe with text fallback, and SLURM 25.11’s nestedreturn_code.The historical
SlurmBackend/LocalSlurm/SshSlurmand their error classes.
pbs.py: PBS Professional / OpenPBS, tested only against mockedqsub/qstat/qdeloutput (no PBS site).Resources become one
selectstatement, andpartitionis the queue.qstat -x -f -F json, falling back to theqstat -xtable plusqstat -x -ffor exit statuses.Exit 271 (a
qdel) means cancelled.
backend.py:QueueBackend(scheduler, transport):submit,poll_many,cancel/cancel_many,count_jobs.local.py/ssh.py:LocalTransport/SshTransport.TASK_SSH_OPTIONSand a command timeout are opt-in, so the JobServer’s argv is unchanged.stage.py: as before, pluspush/pullof many relative paths in onersync --files-from=-, which openrsync on macOS supports. The ssh options and the timeout are opt-in here too.config.py:TargetSection(SlurmSectionis an alias) with the task keys (below),build_task_backend/build_task_stager,task_settings/from_settingsfortarget.json,load_target(load_slurm_configis an alias) andlist_sections.script.py:build_script(directives, payload, scheduler="slurm").
171 tests: the 112 ported from
seamm_slurmplus the interface, PBS, the task keys, staging, and SLURM 25.11 JSON captured on TinkerCliffs.seamm_slurm(the shim)Every module re-exports from
seamm_scheduler.local,sshandstagekeep animport subprocess, because callers patchseamm_slurm.<module>.subprocess.run. Its own 112 tests pass unchanged, as doseamm_jobserver’s.seamm_execscheduler_backend.py:SchedulerBackend, theTaskBackendfortasks = queue.Each
submitcall is one bundle: one batch job that runspython -m seamm_exec.task_worker bundle.jsonin SEAMM mode.Backend ids are
<job id>#<bundle>.<n>#<key>, so a restarted evaluator can rebuild everything from the manifest record (adopt).room()ismax_queued_tasksless the user’ssqueue --me -rcount, refreshed once per poll interval.A QOS submit-limit error, or a failure to reach the cluster, raises
QueueFull, so the bundle is held, not failed.
bundle_runner.py: the SEAMM-mode bundle.It runs the tasks through a
LocalPoolsized to the allocation, withresolve_programs=True. Programs come from the<program>.inifiles where the bundle runs, then the resolver hook.So
{code}/{NTASKS}, conda/modules,$TMPDIRscratch andreturn_filesbehave exactly as in the evaluator’s pool, and tasks that fit side by side in the allocation run concurrently.It writes
DONE(in the task layer’s own form, with thefileslist) orFAILED(with the reason) intasks/<key>/.task_worker.pyonly loads it by name for"mode": "seamm", so the pure path still imports nothing from SEAMM (phase 1’s test enforces this).
resolve.py: the per-program resolver hook (phase 1 requirement a).Entry-point group
org.molssi.seamm.exec.resolvers, named after the program.hook(config, cmd, env, ce, root) -> (config, cmd, env).read_configandregister(for tests).
targets.py:find_target(explicit, then<job>/target.json, then$SEAMM_TARGETwith$SEAMM_TARGETS, then none) andwrite_target.TaskSet:It picks its back end from the target, and takes
bundle_tasks,bundle_walltimeandinline_belowfrom it. On a queue with neither bundling setting, each task is its own bundle.Bundling is by count and by summed walltime (or estimate).
A back end with
bundles = Truegets onesubmitper bundle (withbundle=andmarkers=), while it hasroom(). The rest wait in_held.A failed
submitrestores the attempt count._reattachadopts queued/running tasks from a back end withadopt().Lost tasks of one pass go back together, so a bundle’s lost tasks return as one bundle.
A lost task’s reason comes from the back end (with the tail of the bundle’s log).
route()keeps a task that carriesTask.configon this machine when the back end does not accept it (an ssh target), with one warning._restorelists the returned files for a pure-workerDONEwithout afileslist (phase 1 requirement b; the SEAMM-mode worker writes the list itself, and the TaskSet rewritesDONEon collection).
computational_environment(): the job variables come fromseamm_scheduler(SLURM, then PBS), and a_pbs()reader was added.Base’s in-situ check uses the same variables.91 tests: 51 from phase 1, 40 new. The new ones run with a fake queue that executes each real batch script with bash, so the real worker, pool and markers are exercised.
seamm_jobserverIt imports
seamm_scheduler. A section withtasks =is written as<job dir>/target.jsonbefore the job starts (and so before staging); sections without it write nothing. 87 tests.
Target keys (all optional; a section without tasks = means what it did)#
|
|
|
|
|
default: yes for the local task transport, no for ssh |
|
tasks per bundle |
|
SLURM time syntax; also the bundle’s |
|
the user’s queued + running jobs, counted with |
|
seconds (default 60) |
|
ssh targets: a Python with |
|
its |
|
seconds (default 30) |
|
|
The other keys of a section (partition, account, qos, export,
constraint, …) are the bundles’ site defaults, and each bundle’s resources
override them. setup lines run before the worker. For type = slurm (the
evaluator is itself a batch job) tasks always use the local transport with a
shared filesystem, because the evaluator is already inside the cluster.
bundle_walltime is parsed with SLURM’s rules: a bare number is minutes.
config._parse_time used to read a bare number as seconds. This only matters
for a .limits bound written as a bare number, which no deployed file has.
Decisions made with the design session#
How the evaluator finds its target.
The JobServer writes
<job dir>/target.json, a new job-level file (added to the rollout promise in the design).It works for an evaluator that is itself a batch job on a cluster that cannot read the JobServer’s ini file, and it survives resubmission.
The order is: explicit,
target.json,$SEAMM_TARGET(for hand runs withrun_flowchart), then theLocalPool.
Tasks run on the compute node through seamm_exec’s own ``_run_task``.
The bundle runs with
remote_python(ssh) orsys.executable(local transport), so the semantics are identical to theLocalPool.Task.configis never sent to an ssh target: a task that carries it is kept in the evaluator’s pool with a warning, and a task without it is resolved where it runs. On the local transportTask.configis sent and used as the pool uses it.The pure worker path stays, for machines without
seamm_exec.
The protocol has ``poll(run, ids)``, with the SLURM override. Bundling and throttling live in the
TaskSet.
Consequence for ORCA.
orca_stepresolves its own configuration in the evaluator: it passesTask.config, the Mac’slibrary-pathexport prefix incmdand the OpenMPI binding inenv.Those paths mean nothing on another cluster, so on an ssh target the
TaskSetkeeps ORCA (and MOPAC, which also passesconfig) on the evaluator’s machine, exactly as today.Sending them out needs an
orcaresolver hook inorca_stepand a bare command. That comes withget_taskin phase 3.Verified with
Testing/test.flowand the job’s target set to TinkerCliffs: both tasks ran locally, the warning was logged, and nothing was submitted.
How a bundle’s partial completion is handled#
The worker skips any task whose
tasks/<key>/DONEexists. So running the samebundle.jsonagain, after a walltime limit or a node failure, runs only what is left (tested). When a bundle’s job ends, a task withoutDONEorFAILEDis lost, and its reason names the job’s state and the last line of its log.The
TaskSetresubmits the lost tasks of one pass together. They keep their bundle name, so they become_bundles/<bundle>.<n+1>/with abundle.jsonof just those tasks, and the finished ones are never rerun.Each lost resubmission counts an attempt (
max_lost_retries2 per run,max_attempts3 overall).On a shared filesystem, results are yielded as the worker writes each
DONE, while the bundle still runs. Without one, they come back when the bundle’s job ends and its directories are pulled.
Outages (the laptop case)#
Paul’s laptop sleeps, changes networks and needs a VPN at home, so the evaluator may lose the cluster for minutes to an hour.
A job counts as missing (three polls in a row) only when the queue actually answered.
Scheduler.poll_failedis set whensacct(orqstat) could not be asked, and then nothing counts.A failed poll or a failed pull is logged and retried at the next poll.
An ssh or rsync failure at submission holds the bundle (
QueueFull) instead of failing it.The task layer’s ssh uses
BatchMode,ConnectTimeout=30,ServerAliveInterval=15/CountMax=4(a connection dead after sleep is given up within about a minute) andClearAllForwardings. Thetinkercliffsalias has aLocalForward, which otherwise fails to bind on every command. Commands time out after 300 s and rsync after an hour.The JobServer’s ssh is unchanged.
Validation#
This Mac -> TinkerCliffs (ssh, rsync staging, SLURM 25.11).
Setup:
phase2_driver.py(here) with eight B3LYP/def2-SVP single points, four MPI ranks each, bundled four to a job.Target
[tc]inTesting/phase2/targets.ini:account = seamm,normal_q,tc_normal_short,export = NONE.A development venv in
/projects/seamm/psaxe/phase2/venv: the frozen production stack plus these checkouts./projects/seamm/SEAMM/venvwas not touched, and ORCA comes from the productionorca.ini(installation = modules).
The evaluator was killed with
kill -9while both bundles were queued. The second run logged “still with queue:tc … polling it rather than submitting again” for all eight tasks, staged back, and finished withattempts=1and no new jobs.ORCA ran with “4 parallel MPI-processes” in
/localscratch/<jobid>(node-local, from$TMPDIR).The script carried
--ntasks=4 --mem-per-cpu=1200M --time=00:40:00(the tasks’ walltimes summed).
This Mac -> MolSSI10 (ssh, rsync, SLURM 20.11, no ``–json``). Eight PM6 MOPAC single points (conda installation) in three bundles, all finished.
squeue --json/sacct --jsonare “unrecognized” there, so status came from the text path.Cluster alone (an evaluator on TinkerCliffs, ``type = slurm``, so local transport, shared filesystem, no staging). The driver on the login node gave the same eight ORCA energies as from the Mac, to the last digit. Results arrived while bundles were still running.
The SEAMM_DEV A/B comparison.
The sides:
A: the current freeze with the released
seamm-exec2026.10.2,orca-step2026.10.2.1,mopac-step2026.10.2 andseamm-util2026.10.2.B: A plus the four checkouts.
Both were built beside the current version, as
venvs/phase2-A/-B, without switching.
seamm-manager --root ~/SEAMM_DEV compareontest.flow,harness_water,builder_loopandbsse:identical energies, charges and tables;
the only differences were versions and citation wrapping, pids, timings, timestamps,
tasks/manifest ids and chargemol’s last-digit noise, the same as in phase 1.
test.flowwithSEAMM_TARGET=tcwas likewise clean (above).
Review (2026-10-02)#
An independent review of the four diffs reproduced five bugs and found nine
plausible ones. All are fixed except where noted. The scripts are in
/private/tmp/claude-502/review/; each fixed bug now has a test.
A pull-back that failed once was never retried, and the evaluator spun at 100% CPU (the job was terminal, so it was no longer polled). Jobs that ended but are not back are now worked out from every tracked job, the poll time is set on every pass, outages do not count as failures, and five real failures (e.g. purged scratch) make the tasks lost with that reason.
The JobServer could not start the design’s
[arc](type = local,transport = ssh):_build_cmdused the remoterun_from_jobserverfor a local evaluator. Onlytype = slurmwithtransport = sshnow means a remote evaluator.A restart with changed inputs adopted the old bundle and recorded its output under the new fingerprint. Adoption now requires the same fingerprint,
DONE/FAILEDare honoured only if their fingerprint matches (in the backend and in the worker), and the old job is cancelled when no adopted task still runs in it.Unparseable ``squeue``/``sacct`` output (a banner, a schema change) read as “every job is gone”. It now sets
poll_failed.PBS:
Fwithout an exit status (deleted while queued) never ended; it is now cancelled (lost). SLURM spellings in a section are dropped with a warning;qselectcounts and finds jobs.A submission whose outcome was unknown (the connection dropped after
sbatchran) was submitted again as a new bundle. Each bundle’s job name is unique (seamm-<bundle>.<n>-<nonce>), and the next attempt asks the queue for that name first.The manifest was written only after all bundles were submitted. It is now written after each.
finallymarked tasks cancelled even whenscancelfailed (an outage), so a restart resubmitted them. They are marked only when the cancel worked.On a shared filesystem a just-written
DONEhidden by attribute caching could make a task lost. A job gets one more poll after it ends, and markers are read from a freshos.listdir.Remote job directories could collide across installations sharing a
remote_root(job numbers are per datastore). The remote name now always carries a hash of the local path.Clusters without SLURM accounting would have had
poll_failedforever. “accounting storage is disabled” is recognised andsqueuetrusted.Bundles submitted from inside an evaluator job inherited its
SLURM_*variables. The local task transport dropsSLURM_*andPBS_*.Behaviour changes, kept and recorded in HISTORY:
Base’s in-situ default andcomputational_environment()now recognise PBS jobs (no SEAMM installation runs under PBS today); a bare.limitstime is minutes; a bad task key in a section raises when the file is read, as a badtypealways did.Minor:
xargs -0for the remote marker cleanup; a bundle held for more than 15 minutes logs a warning; a staged task outside the job directory fails before anything is submitted. Not changed:_list_returnedkeepsBase’s flat file names.
After the fixes, the live Mac -> TinkerCliffs and Mac -> MolSSI10 runs were repeated: both clean, with the same energies.
Found on the way#
``computational_environment()`` crashed in any SLURM job without ``–ntasks`` (no
SLURM_NTASKS:KeyError). Its hostlist expansion crashed ontc[053,059]and dropped the zero padding oftc[053-055]. Both are fixed. The first was found when a probe bundle withoutntasksdied before the worker started. Bundles now always requestntasks.The first cluster-alone run had no evaluator root (no flowchart), so the worker had no
orca.ini. The local transport now falls back toseamm_util.root.current_root().A conda-installed code needs
shell=True, sinceLocal.execmakes it aconda run ...string. Without it the whole string is taken as a program name. This is pre-existing and was hit by the driver, not by a plug-in.SLURM 25.11’s
squeue --jsonreports["TIMEOUT", "COMPLETING"]: the first flag is the state.arc_pick.py(the arc-slurm-submit skill) could not read TinkerCliffs’s state during a short network outage on this side; its choice was made by hand (sbatch --test-only).~/SEAMM_DEV/PaulsPersonal.local.iniis orphaned again: this Mac is nowPaulVT.local. Reported to the design session; left alone.
Deferred#
The
orca(andmopac) resolver hooks and a bare command, withget_task, in phase 3. Until then those steps’ tasks stay local on ssh targets.Partial progress without a shared filesystem: results come back only when a bundle ends. A periodic pull of the markers would give them earlier.
max_queued_taskscounts the user’s jobs on the cluster, but two evaluators can still race for the last slots. A rejectedsbatchis held and retried, so this is harmless.PBS counts and finds jobs with
qselect -u "$USER"(-Nfor a name), implemented but untested on a real site.Remote staging directories (
<remote_root>/<Job>-<hash>/) are never removed. MolSSI10’s home has no purge, so they accumulate there until removed by hand. Cleanup once aTaskSethas collected everything, or when the job is deleted, is left for later.TinkerCliffs’s production venv is not versioned yet. That is a later rollout step.
Cleanup#
Test directories:
~/Work/SEAMM/Testing/phase2/(Job_9000xx)tinkercliffs:/projects/seamm/psaxe/phase2/{tasks,Job_9000xx}molssi10:~/phase2/tasks
Development venvs:
tinkercliffs:/projects/seamm/psaxe/phase2/venvmolssi10:~/phase2/venv~/SEAMM_DEV/venvs/phase2-{A,B}, to be pruned once phase 2 is released.
Design session’s review (2026-10-02)#
A second, independent review. Each item is fixed, with a test.
The squeue JSON path never flagged a transient failure. With
rc=255(“ssh: timed out”),poll_failedstayed False. Any squeue error other than “invalid job id” now marks squeue failed, andpoll_failedis set when squeue failed and a job is still missing aftersacct(sacct may not have recorded a new job yet), with or without accounting.A restart could still duplicate a submission. The record said queued with no id if the evaluator died between
sbatchand the manifest write, or while a bundle was parked during an outage.Each bundle’s unique job name and directory are now written to its tasks’ records, flushed, before
sbatch(on_prepared).A record without an id is adopted by name.
findissqueue --me --name, thensacct --namefor a job that already finished.If neither knows the job, the bundle never reached the queue: lost, and submitted again.
A bundle that may have reached the queue stays recorded as queued.
Live: TinkerCliffs refuses a 30-day
sacctrange (“Too wide of a date range”), so it uses 14 days, then 2. Checked on TinkerCliffs and MolSSI10 with names of earlier bundles.
QOS-full churn. Each retry made a new
bundle.<n+1>, rewrote the inputs and restaged them. A held bundle is now prepared once (_Prepared: directory, job name, script, staged, maybe-submitted) and reused until it is queued.expand_hostlistfollows SLURM’s grammar: a suffix (tc[01-02]-ib) and several bracket groups (r[1-2]n[1-2], the product).Paths with whitespace are refused up front with a clear message (batch directives and rsync’s remote paths cannot carry them safely).
An earlier attempt’s output can no longer pass a success check. The
success_textfiles are removed from the task directory, locally and on the cluster, at submission, and the worker trusts a file on disk only if this run wrote it (its mtime).Abandoned jobs (inputs changed) are polled until they have ended,
COMPLETINGincluded (two minutes at most), before the new inputs are submitted into the same directories.Remote staging directories are not cleaned up: recorded under Deferred.
Also:
Task.configis described as kept local, not ignored.directives()returns a dict in the design’s protocol.A section must hold no secrets, since it is copied to
target.json; this is in the design, the config module and the JobServer guide.Bundles refuse a non-local executor instead of silently running locally.
SLURM_CONFsurvives the dropping of the evaluator’sSLURM_*.The task digest is computed once per backend entry, not on every poll.
There are tests for bare
.limitstimes (minutes). A local and queued mix in flight is covered bytest_inline_rule_with_a_real_scheduler_backend.
2026-10-02 – release PRs#
Paul created molssi-seamm/seamm_scheduler (public, empty); PyPI publishing
uses the organization secret the shared Release workflow already uses. Its
main is an initial commit (LICENSE, README, .gitignore) and dev the work,
joined by an ours merge so the PR applies cleanly.
PRs, to be merged and released in this order:
seamm_scheduler#1, 2026.10.2. CI green.seamm_exec#34, 2026.10.2.1. It pinsseamm-scheduler>=2026.10.2, so its CI is red until (1) is on PyPI (the phase 1 lesson).seamm_slurm#9, 2026.10.2 (the shim). Same pin; CI on uv (test_env.yamlremoved).seamm_jobserver#24, 2026.10.2. It needsseamm_scheduler>=2026.10.2; CI on uv (test_env.yamlremoved).
After the first release of seamm_scheduler, enable GitHub Pages for it
(gh api -X POST repos/molssi-seamm/seamm_scheduler/pages ...), as for every
new repository. If the merges happen after 2026-10-02, the versions in the
HISTORY entries should follow the release date.
2026-10-03 – released#
Merged in order and released with the 2026.10.2 dates (Paul’s call). The
checkouts are synced with make update (dev == main).
Package |
Version |
PR |
|---|---|---|
|
2026.10.2 |
#1 (new repository; GitHub Pages enabled) |
|
2026.10.2.1 |
#34 (requires |
|
2026.10.2 |
#9 (the shim) |
|
2026.10.2 |
#24 |
Nothing is rolled out to an installation yet; that is a separate step.
Lesson (again)#
A package that pins an unreleased library is red in CI, and cannot even be installed locally, until the library is on PyPI.
Phase 1 learned this when
orca_stepwas pushed beforeseamm_exec.Phase 2 met it on purpose. The PRs for
seamm_exec,seamm_slurmandseamm_jobservercarriedseamm-scheduler>=2026.10.2and said “red until #1 is released”.Locally,
make installuninstalledseamm_execand then could not resolve the pin, which left the development environment without it.pip install --no-deps .(anduv pip install --no-depson TinkerCliffs) was the way through.The rule stays the same: release the library first, merge the dependants after, and expect red CI only in between. Say so in each dependant PR.
Two more points:
Live runs keep finding what tests cannot. TinkerCliffs refuses a 30-day
sacctrange, andcomputational_environment()crashed in jobs without--ntasks. Each code path that talks to a real scheduler was exercised at least once against TinkerCliffs (SLURM 25.11) and MolSSI10 (20.11) before release.Two independent reviews (a subagent, then the design session) found 14 and 8 + misc issues. Most were state-machine holes around restarts and outages that the first round of tests did not reach. Converting every repro into a test was worth it.
Still deferred#
ORCA and MOPAC: their resolvers (
orca: full path, OpenMPIlibrary-path, binding;mopac) and a bare command, withget_task, in phase 3. Until then their tasks stay on the evaluator’s machine for ssh targets.Partial progress without a shared filesystem: results arrive when a bundle ends. A periodic pull of the markers would give them sooner.
Remote staging directories (
<remote_root>/<Job>-<hash>/) are never removed. MolSSI10’s home has no purge.PBS has been tested against recorded output only (
qsub/qstat/qselect). Validate it on a real site (phase 7).Several evaluators on one machine each think they own it (from phase 1), and
max_queued_taskscounts are racy between evaluators. Rejected submissions are held and retried.TinkerCliffs’s production venv is not versioned, and has none of these releases. Rolling out is a separate step.
``seamm_webui`` still imports
seamm_slurm(fine through the shim). Move it toseamm_schedulerwhen it is next released.
Cleanup (2026-10-03, with Paul’s agreement)#
Deleted:
tinkercliffs:/projects/seamm/psaxe/phase2(the development venv, sources, staged tasks and test jobs; 1.5 GB);molssi10:~/phase2(1 GB);~/SEAMM_DEV/venvs/phase2-{A,B}.
SEAMM_DEV’s current version is untouched. ~/Work/SEAMM/Testing/phase2
(targets.ini and the local test jobs) is kept as a record.