Phase 1 notes ============= 2026-10-02 -- the task layer in seamm_exec, and ORCA and MOPAC on it -------------------------------------------------------------------- Done on ``dev`` in ``seamm_exec``, ``orca_step`` and ``mopac_step``. Committed locally, not pushed and not released. What was built (``seamm_exec``) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ - ``tasks.py``: ``Resources``, ``Task``, ``TaskResult``, the ``TaskBackend`` protocol, ``Manifest`` and ``TaskSet``, plus ``run_task()`` for a step that runs one calculation. - The manifest is ``/tasks/manifest.json``. It is written lazily (dirty flag, at most once a second, and always at the start and end of a run), because a 2,800-task frame would otherwise rewrite a multi-megabyte file thousands of times. - Whether a task finished is recorded by its ``tasks//DONE``, which is written at once. - ``local_pool.py``: ``LocalPool``. - Its capacity is ``computational_environment()``: cores, available memory and GPUs. - Tasks start in submission order as soon as their cores and memory are free. A task larger than the pool is clamped to it and runs alone. - Each task gets its own ``ce``: ``NTASKS``, ``CPUS_PER_TASK``, ``MEM_PER_CPU`` and ``MEM_PER_NODE`` from its ``Resources``. - A task runs through the executor's ``_run_task()`` in a worker thread, and ``Local.exec`` starts it with ``Popen(start_new_session=True)``. - SIGTERM and SIGHUP to the evaluator kill the task process groups, then chain to the previous handler. The JobServer stops a job with ``process.terminate()``, so this covers it. A signal the evaluator ignores is left alone: under ``nohup`` SIGHUP stays ignored and the tasks survive, as the codes did before. - ``task_worker.py``: the pure-Python bundle worker, run by path. It applies ``success_text`` too. - ``base.py``: the old ``run()`` body is now ``_run_task()``, unchanged except for two opt-in knobs the concurrent pool uses: - ``set_umask=False``: the umask is process-wide, so concurrent tasks must not change it. - ``keep=[...]``: the in-situ cleanup must not prune the ``tasks/`` bookkeeping of the step directory. ``run()`` builds a one-task ``TaskSet`` with ``manifest=False`` and a synchronous one-slot ``LocalPool`` that runs in the caller's thread. It passes the caller's ``ce`` unchanged, leaves the environment untouched, and lets exceptions propagate. It writes no new files. - ``local.py``: under the pool only (a thread-local hook), ``Popen`` in a new session instead of ``subprocess.run``. The shim path is unchanged. Decisions taken with the design session ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ These are now in the design's task API section. - ``Task.directory``: None means ``tasks//``, the default only for fan-out. The ORCA and MOPAC call sites pass their own directories, so ``orca.out`` and ``mopac.out`` stay where users and the Dashboard expect them. - ``Task.config``: a local override of the ``.ini`` section. ORCA and MOPAC keep resolving their own configuration, and remote back ends will ignore it. ``LocalPool`` adds ``{code_dir}`` (the directory holding ``code``); ORCA's ``orca_2aim`` now uses it instead of an absolute path baked into the command. - ``Task.fingerprint``: the restart identity. The default hashes ``cmd`` and ``files``, never ``env``. ORCA gives one that ignores ``%pal`` and ``%maxcore``, so a different machine size does not force a rerun. - ``Task.success_text``: ``{file: text or [texts]}``. Added after the live test below found that **ORCA exits 0 after an error termination**, so a failed run was being marked ``DONE``. Both ORCA call sites set ``{"orca.out": "ORCA TERMINATED NORMALLY"}``, the test ``_parse_output`` already used. MOPAC sets none: its parser checks nothing today and its ``success.dat`` marks a run done whatever the outcome, so it is unchanged (consider-each-code-separately). ``mopac.out`` ends with ``== MOPAC DONE ==`` and ``mopac.arc`` has ``JOB ENDED NORMALLY`` if that is wanted later. - Restart: - Within a run a failed task is never retried, and a lost one is retried up to twice. - Across runs a failed or lost task is tried again until it has had three attempts in all, then reported failed with its history. Each history entry and the manifest record carry a ``reason``: the return code, the failed success check and the file, or lost. - New inputs (a changed fingerprint) reset the attempts. The old count and history are kept under ``previous`` in the manifest record. - A task refused for having used up its attempts gets a ``reason`` that starts "attempts exhausted" and names the manifest. Changing its input, or deleting its entry in ``tasks/manifest.json``, lets it run again. ORCA and MOPAC raise or stop on such a result instead of parsing its stale output, and print the reason of any other failure. - ``LocalPool`` does not reattach: it kills a leftover process group and reruns the task. Since a pid is reused once its process has gone, the group is killed only if its leader is the recorded process: same host, start time within 1 s, same working directory. The start (``pgid``, ``create_time``, ``cwd``) is written to the manifest at once, not at the next throttled save, so an evaluator killed a moment after a start still leaves enough to find the process. A laptop whose host name changed with its network leaves leftovers alone, the safe direction. - MPI binding (``OMPI_MCA_hwloc_base_binding_policy=none``) and ``OMP_NUM_THREADS = cpus_per_task`` are set **only when another task is queued or running in the pool as the task starts**, decided when it is dispatched, on or off SLURM, and never over a value the task sets. A task running alone gets exactly the environment it always had. - Bundling: ``TaskSet`` assigns bundles in the order tasks are added (``bundle_tasks``; one bundle by default). ``archive=True`` appends a bundle's ``tasks//`` directories to ``tasks/.tar`` once all of its tasks are done, removes the directories, and records the tar in the manifest. Only finished tasks are archived; failed ones stay as directories, since they may run again. Restart reads ``DONE`` and the returned files from the tar, opening each tar once per run with an index of its members. ``LocalPool`` does not go through the worker script. - Inline rule: ``TaskSet.route()`` sends a task with ``estimated_seconds < inline_below`` (60 s) to the evaluator's own pool when the program is installed here (``task.config`` given, or a ``[]`` section in ``/.ini``). Phase 1 has no target plumbing, so it is tested with a fake remote back end. - ``estimated_seconds`` heuristics, both rough: - ORCA: method class and basis family from the ``!`` line, times ``n_atoms**2.5``. Water at B3LYP/def2-SVP comes out at about 1.5 s. - MOPAC: a cube in the atom count, times 30 for an optimization. Validation ~~~~~~~~~~ - ``seamm_exec`` has 51 tests: 39 in ``tests/test_tasks.py`` (the task layer) and 12 in ``test_seamm_exec.py`` and ``test_credential_sections.py`` (the shim's in-situ cases, memory and credentials). They cover: - the shim's result dictionary and files, exception propagation, and ``@subdir+``; - concurrency, slot partitioning by cores and memory, whole-pool and oversized tasks, and binding only when concurrent; - ``{code_dir}`` and configuration from the ini file; - in-situ pruning per task and in the step directory, and scratch copy-back; - restart: skip finished tasks, rerun changed inputs, the fingerprint override, env not hashed, and the attempts cap; - SIGKILL and SIGTERM of a real evaluator subprocess, then rerun; - one tar per bundle, with restore from the tar; - the inline rule, the worker, ``success_text``, and a broken executor that fails its task instead of hanging; - from the review: the shim under ``SLURM_JOB_ID`` (``mkdtemp`` honouring ``$TMPDIR``, only ``return_files`` back, inputs also written to the directory) and with ``config``, ``env`` and ``ce`` passed through untouched; lost-task retries; the pid-reuse guard; the start recorded before any throttled save; separate start callbacks on one pool; a bare ``code`` without ``{code_dir}``; the attempts reset; an ignored SIGHUP; and failed tasks left out of the archive. All pass repeatedly in about 19 s. ``orca_step`` passes 166 and ``mopac_step`` 35, each with tests of the new helpers. - ``Testing/test.flow`` (ORCA energy, then DDEC6 through ``.wfx``) runs clean. Rerun in place without clearing ``tmp/``, ORCA reports "ORCA had already finished this calculation, with the same input ... using those results", and ``job.out`` differs from the first run only in timestamps and that line. - ``harness_water.flow`` and ``builder_loop.flow`` (MOPAC in a Loop) run clean. Each iteration has its own directory, so nothing is skipped between iterations; energies are C −13.43, CC −14.52, CCC −17.43. - ``phase1_driver.py`` (here) runs four ORCA (two cores each) and eight MOPAC calculations concurrently through one ``TaskSet``. After ``kill -9`` once two ORCA runs had finished, both running ORCA process groups were still alive. The rerun killed them, restored the two finished runs, ran the other ten, and left no processes behind. Found on the way ~~~~~~~~~~~~~~~~ - ``seamm-dev`` had a stale ``seamm`` (2026.9.29), so 115 ``orca_step`` tests failed with or without these changes (``applies_when``). Fixed with ``make install`` in ``seamm``. - A test double without ``_run_task`` made a pool thread die before it freed its slot, and the ``TaskSet`` then waited forever. The thread's setup is now inside its ``try``, so any error fails the task. - The pool's memory is the machine's *available* memory, which was 4 GB under load on this Mac. With 2 GB per ORCA task that allowed two at a time, which is correct but conservative. Total memory, or a site setting, may be the better capacity later. Review fixes (2026-10-02) ~~~~~~~~~~~~~~~~~~~~~~~~~ The design session's review of the three diffs found these, all fixed: 1. The attempts cap never reset when the input changed. 2. The orphan kill could hit a reused pid. 3. The start reached the manifest only at the next throttled save. 4. The SIGHUP handler replaced ``SIG_IGN``. 5. ``{code_dir}`` turned a bare ``orca`` into ``./orca_2aim``; ORCA now uses a bare ``orca_2aim`` then. 6. The start callback was shared on the pool; it now travels with each ``submit``. 7. A refused task slipped past the plug-ins' error check. 8. The definition of "concurrent". 9. The shim parsed ``ce``; it no longer reads it. 10. Archiving: failed tasks were tarred, and each tar was reopened per member. Deferred, for later phases ~~~~~~~~~~~~~~~~~~~~~~~~~~ - **Phase 2 requirement: the per-program resolver hook.** An entry point that, on the executing machine, turns ``program`` + ``Resources`` into the ``cmd`` prefix, ``env`` and configuration. It would replace what ORCA (``full_orca_path``, ``library-path`` → ``_mpi_env``) and MOPAC (default ini, ``which mopac``) do in the evaluator today. ``Task.config`` stays as the local override. - Phase 2: the worker's ``DONE`` records no ``files`` list. The scheduler back end's ``fetch`` must rewrite it in the ``TaskSet`` form after staging back, or ``_restore`` must list the returned files from ``return_files``. - Several evaluators on one machine each think they own it (four jobs × the whole pool). This is not a regression, since four ORCA jobs each use all cores today. A machine-wide slot lock would fix it. - A multi-node allocation is shared out as one pool of ``NTASKS`` cores; the per-task ``NODELIST`` is not narrowed. - Tasks of the ``docker`` installation and executor cannot be killed by the pool, which has no process to signal. - MOPAC's own ``success.dat`` skips a rerun of a completed step whatever its input now, ahead of the task layer's fingerprint. Retire it in favour of the fingerprint (a ``mopac_step`` change). 2026-10-02 -- released ---------------------- - Releases: ``seamm_exec`` 2026.10.2 (PR #33) first, then ``orca_step`` 2026.10.2.1 (PR #36) and ``mopac_step`` 2026.10.2 (PR #157). Each plug-in requires ``seamm-exec>=2026.10.2`` in ``requirements.txt``; neither has a ``test_env.yaml`` any more, so CI installs exactly those requirements. - Each plug-in's user guide has a "Rerunning a job" section: what a rerun reuses, when a calculation runs again, and where the record lives (``tasks/manifest.json``). MOPAC's says that a completed step with ``success.dat`` is skipped whatever its input, and that deleting ``success.dat`` forces a recompute. The A/B comparison in ``~/SEAMM_DEV`` ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The current environment had ``orca-step`` 2026.10.1, behind the base of this work, so two versions were built beside it, without switching (the JobServer was untouched). Each was a copy of the current ``uv pip freeze`` (what ``build_version`` does) plus: - A: ``orca-step`` 2026.10.2, i.e. everything released at the base of the work (``seamm-exec`` 2026.9.27, ``mopac-step`` 2026.10.1); - B: A plus the three checkouts (``f0bae06``, ``f43b8f1``, ``87feb48``), installed non-editable. ``seamm-manager --root ~/SEAMM_DEV compare -a A -b B`` on four flowcharts: - ``test.flow``: ORCA, then DDEC6 with chargemol going through the ``Base.run()`` shim; - ``harness_water.flow`` and ``builder_loop.flow``: MOPAC, the second in a Loop; - ``bsse.flow``: five ORCA counterpoise sub-jobs in subdirectories. All energies, charges and tables agree. The only differences were: - versions, and the paths and citation wrapping they change; - timestamps, and ORCA's process ids and timings; - the new ``tasks/manifest.json`` and ``DONE`` files, which exist only in B; - chargemol's last-digit ``Normalization`` noise in ``orca.output``. Two independent runs of A, compared with ``--no-run``, differ in exactly the same files, so that noise is run to run and B adds nothing. ``mlff_packmol`` was not used: its PACKMOL placement and MD are random on every run. Both versions were pruned afterwards. Lessons ~~~~~~~ - **Push a plug-in that uses a new library API only together with its minimum-version pin, or expect red CI until then.** ``orca_step`` dev, pushed before ``seamm-exec`` 2026.10.2 existed, failed its Branch CI: ``seamm_exec`` had no ``run_task`` or ``TaskResult``, because CI installed 2026.9.27 from PyPI. ``mopac_step`` passed only because its tests do not touch the new API; its code needs the pin just as much. The release PRs' ``seamm-exec>=2026.10.2`` fixed both. Run the pre-flight with the *released* library installed, not the development build, which hides exactly this. - ORCA exits 0 after an error termination, and a pid is reused once its process has gone. Both were caught only by running real codes and by review, not by the first round of tests.