.. _user-guide: ********** User Guide ********** The ORCA plug-in runs `ORCA `_ from a SEAMM flowchart. It is a *sub-flowchart* plug-in: dropping an **ORCA** step onto the canvas opens a small flowchart of ORCA sub-steps. Four are available: * **Energy** — a single-point energy, optionally with the gradient (forces) and a range of properties. * **Optimization** — a geometry optimization (the Energy step plus ORCA's ``Opt`` keyword). * **Frequencies** — the Hessian and harmonic vibrational frequencies, with thermochemistry (see below). * **BSSE** — the counterpoise-corrected energy and gradient of a two-fragment complex (see below). All four share the same controls for the level of theory, so the sections below apply to any of them. Choosing the level of theory ============================ Each sub-step can take its method and basis set either from a preceding **Model Chemistry** step (the default — leave *Use the global model chemistry* on) or set them explicitly. Turn *Use the global model chemistry* off to choose them in the step itself. Method ------ The **Method** pull-down offers Hartree–Fock, MP2/RI-MP2, coupled cluster (``CCSD(T)``, ``DLPNO-CCSD(T)``), the explicitly-correlated **F12** variants (``CCSD(T)-F12D/RI``, ``DLPNO-CCSD(T)-F12D``), and **DFT**. For everything except DFT the method *is* the ORCA keyword. Choosing **DFT** reveals two further, indented controls: * **Functional type** — ORCA's own classification of the functional: local, GGA, meta-GGA, (range-separated) hybrid, and (range-separated) double-hybrid. * **Functional** — the functional itself, filtered to the chosen type. Every functional ORCA documents is available, grouped by type; picking a type narrows the functional list so it stays readable. Double-hybrid functionals (for example ``REVDSD-PBEP86-D4/2021`` or ``B2PLYP``) need an auxiliary ``/C`` fitting basis for their MP2 part — leave the *Auxiliary (fitting) basis* on ``AutoAux`` and it is supplied automatically. Most double hybrids also come in a **DLPNO** variant, listed after the canonical ones (e.g. ``DLPNO-REVDSD-PBEP86-D4/2021``). It evaluates the MP2 part with near-linear-scaling DLPNO-MP2 instead of canonical RI-MP2, so it pays off for large systems and clusters; with the default thresholds it differs from the canonical result by roughly 0.1 kJ/mol. ORCA rejects some ``DLPNO-`` keywords even though its manual lists them (``DLPNO-REVDSD-PBEP86-D4/2021`` is one), so the step writes the canonical functional plus ``%mp2 DLPNO true end``, which works for all of them. ORCA 6.1 can compute DLPNO-MP2 gradients only for closed-shell systems. An open-shell DLPNO double hybrid can therefore give an energy but not forces, an optimization, or frequencies, and the step stops with an error if you ask for any of those. Energies of formation for a DLPNO double hybrid use the canonical functional's atomic reference energies (DLPNO changes an isolated atom's energy by less than 0.07 kJ/mol, and cannot treat H at all), and the thermochemistry report says so. Basis set --------- Type any basis name ORCA knows internally, pick one of the curated families from the list (Pople, Dunning correlation-consistent, and Karlsruhe def2, including the diffuse ``…D`` and minimally-augmented ``ma-…`` variants), or press **...** to browse the `Basis Set Exchange `_ on a periodic table, filtered to the elements in your system. Selecting **Basis Set Exchange** as the *Basis set source* opens the same picker. A basis chosen from the Exchange is stored as ``bse:NAME`` and embedded in the ORCA input, so the definition is identical across codes. The list is ordered by family and, within a family, into valence / polarization / diffuse ladders that each rise DZ → TZ → QZ → 5Z, so a sensible progression is a single ladder read top to bottom. Explicitly-correlated (F12) methods ----------------------------------- The F12 methods (``CCSD(T)-F12D/RI``, ``DLPNO-CCSD(T)-F12D``) reach near complete-basis-set accuracy from a modest basis, but they **must** be paired with one of ORCA's F12-optimized orbital bases — ``cc-pVDZ-F12``, ``cc-pVTZ-F12``, or ``cc-pVQZ-F12`` — and each of those needs a matching *complementary auxiliary basis set* (CABS). The step **adds the CABS automatically**: it derives ``-CABS`` from the F12 basis you choose (so ``cc-pVTZ-F12`` → ``cc-pVTZ-F12-CABS``, and likewise for DZ/QZ), unless you have already put a CABS in the extra keywords. Leave the auxiliary basis on ``AutoAux`` — it supplies the remaining RI fitting bases; the CABS is separate and is not something AutoAux can generate. To keep this foolproof, **selecting an F12 method narrows the basis-set list to just the F12 bases** (and forces the ORCA-internal source), so you can only pick a valid one; switching back to a non-F12 method restores the full list. If a hand-edited flowchart still pairs an F12 method with a non-F12 basis (and no CABS in the extra keywords), the run is stopped early with a clear message. .. tip:: F12 methods largely **eliminate basis-set superposition error**, so the counterpoise (BSSE) correction is essentially zero for them (often a few thousandths of a kcal/mol — at the numerical-noise level). With F12 you can skip the **BSSE** sub-step and take the interaction energy directly from a plain **Energy** run, which is ~5× cheaper. Reserve the BSSE step for conventional finite-basis methods (HF, DFT, MP2, canonical CCSD(T)), where the correction is real. Complete-basis-set (CBS) extrapolation -------------------------------------- Set **Basis-set extrapolation (CBS)** to ``2/3``, ``3/4``, or ``4/5`` to extrapolate the energy to the complete-basis-set limit from two successive cardinal numbers (double/triple, triple/quadruple, quadruple/quintuple), using the **Extrapolation family** (``cc``, ``aug-cc``, ``def2``, or ``ANO``). This is a *single* ORCA job (its ``Extrapolate`` keyword): ORCA runs both basis sets and extrapolates the SCF and correlation contributions separately. When it is on, the fixed basis set above is ignored. Note that an extrapolated energy has **no gradient** in ORCA, so extrapolation and the *gradients* result are mutually exclusive. Integration grid ================ The **Integration grid** control sets ORCA's numerical-integration grid preset (used for the DFT exchange-correlation integrals and the RIJCOSX/COSX grids): * ``default`` — leave ORCA's own default (``DEFGRID2``), which is robust for most work; * ``DEFGRID1`` — coarser and faster (roughly the old ORCA-4 default accuracy); * ``DEFGRID2`` — the current default; * ``DEFGRID3`` — finer and more conservative, for cases sensitive to the grid (e.g. dispersion-dominated energies, tight convergence, or numerical noise in forces). It only affects methods that use a grid (DFT, RIJCOSX) and is ignored otherwise. For a machine-learned-force-field training set, a finer grid (``DEFGRID3``) can be worth the cost to keep the energy/force surface smooth. When the control is left on ``default``, the step **automatically switches to ``DEFGRID3`` for basis sets with high angular momentum** — h functions or above, such as ``cc-pV5Z`` — which the default ``DEFGRID2`` integrates less accurately. The angular momentum is read from the Basis Set Exchange (which also covers ORCA's internal basis names), so it works for both internal and BSE bases; choosing a grid explicitly always overrides this. SCF convergence =============== The **SCF convergence** control sets ORCA's SCF convergence-tolerance preset, from loosest to tightest: ``SLOPPYSCF`` → ``LOOSESCF`` → ``NORMALSCF`` → ``STRONGSCF`` → ``TIGHTSCF`` → ``VERYTIGHTSCF`` → ``EXTREMESCF``. ``default`` leaves ORCA's own default (``NORMALSCF`` for a single point; ORCA already tightens it to ``TIGHTSCF`` for optimizations). The step defaults to **``TIGHTSCF``**, which is a good choice for smooth energies and forces (and matches what ORCA uses for optimizations). Loosen it only to save time when high precision is not needed; tighten it (``VERYTIGHTSCF`` / ``EXTREMESCF``) for very-high-accuracy or numerically delicate work. SCF SThresh (linear dependence) =============================== The **SCF SThresh** control sets ORCA's ``%scf SThresh`` value — the threshold below which an eigenvalue of the overlap matrix is treated as zero, so the corresponding (near-)linearly-dependent combination of basis functions is dropped from the calculation. Redundant, near-linearly-dependent basis functions cause numerical instabilities in the SCF, and diffuse-heavy sets (the augmented ``aug-cc-pVXZ`` / ``ma-…`` families in particular) are the usual culprits. Raising ``SThresh`` above the default removes more of these functions and can cure the resulting SCF trouble. * ``default`` — emit nothing, so the SCF-convergence preset above (or ORCA's own default of ``1.0e-07``) governs ``SThresh``. * an explicit value — written to ``%scf SThresh``, overriding whatever the preset would otherwise set. ORCA recommends keeping it between ``1e-8`` and ``1e-5``; raise it toward ``1e-6`` to shed near-dependent functions when a diffuse basis will not converge. .. caution:: Values beyond ``1e-6`` must be used carefully in **geometry optimizations** and when **comparing conformers**: because different basis functions can be cut off at different geometries, the final basis set — and hence the energy — can vary discontinuously from one structure to the next. See the ORCA manual's `basis-set section `_ for the full discussion of linear dependence and its automatic removal. Initial guess and wavefunction restart ======================================= Single atoms and open-shell transition metals are often the hardest systems to converge -- a superposition of atomic densities (ORCA's default ``SAD`` guess) has little meaning for one atomic center, and can converge to the wrong electronic state entirely. (Whatever the guess, a single-atom system is run with exact exchange, ``NoCOSX``: ORCA's default COSX approximation can build wrong virtual orbitals for a lone atom, which puts the MP2 part of a double hybrid off by up to ~5 kJ/mol, e.g. for Na, Na\ :sup:`+` and Mg, with no warning. This also applies to the bare single-atom fragments of a counterpoise correction. An exchange scheme given in the extra keywords is respected.) The **Initial guess** control sets ORCA's SCF starting guess (``Guess`` in the ``%scf`` block): * ``default`` -- leave ORCA's own default. * ``Hueckel``, ``HCore``, ``PAtom``, ``PModel``, ``SAD``, ``SADNO`` -- ORCA's built-in guess types. ``PModel`` is often much more stable than ``SAD`` for a single atom or ion. * **Previous wavefunction** -- seed from the orbitals (``orca.gbw``) left by the *nearest earlier ORCA step in this flowchart* -- for example, run a cheap, robust functional (``PBE``) immediately before a fragile double-hybrid (``REVDSD-PBEP86-D4/2021``) on the same atom, and turn this on for the double-hybrid step. This only reaches a *different, preceding* node in a roughly linear flowchart -- it cannot seed across iterations of a **Loop**, since a Loop gives each iteration its own directory (see below). * **Specified orbitals** -- seed from a named checkpoint file instead, named by the **Specified orbitals** control revealed underneath. This is the way to seed across loop iterations: for example, a basis-set escalation inside a Loop over basis sets, where each iteration reads the checkpoint the *previous* iteration (a smaller basis) wrote. Either wavefunction choice reveals **If wavefunction not found**, controlling what happens when there is nothing to read (e.g. the very first ORCA step, or the first pass through a basis-set-escalation loop): ``Throw an error`` (the default -- an explicit wavefunction request that cannot be honored is treated as a configuration problem) or one of the ``Use ... guess`` choices, which falls back to that guess type instead. Set this to a fallback guess to make "Specified orbitals" safe to leave on for every iteration of a loop. .. note:: **ORCA projects across basis sets.** ``MORead`` is not limited to restarting with the *same* basis set -- when the checkpoint's basis differs from the current job's, ORCA projects the old orbitals onto the new basis (the same trick ORCA's own ``Extrapolate`` (CBS) keyword uses internally). So this also works for a basis-set escalation, not just same-basis restarts across different functionals. It should still be the same atom/system in the same electronic state (charge and multiplicity) -- MORead supplies starting orbitals, it does not reinterpret the electronic state. Saving a checkpoint for a later step to read is a separate pair of controls: * **Save orbital checkpoint** -- ``yes`` copies this run's converged orbitals to the file named by **Checkpoint name**, revealed underneath, once the run finishes successfully. * **Checkpoint name** and **Specified orbitals** both resolve the same way: ``default`` (the default value of each) derives a label automatically from the system's composition, charge, and multiplicity (e.g. ``Co_q0_m4``) -- the usual choice, since it naturally gives one checkpoint per atom/electronic state, shared across an inner loop (e.g. over basis sets) for that atom, and reset automatically when an outer loop moves to a new atom. Any other bare name is stored under a ``checkpoints`` folder at the top of the job; an absolute path is used as-is, e.g. to keep a checkpoint outside this job and reuse it across separate flowchart runs. .. note:: **Reading a checkpoint from another job.** Only **Specified orbitals** can reference *another* job -- ``job:///`` -- since a job must never write into another job (so this is not available for **Checkpoint name**). Use ``job:///default`` when you want that other job's own auto-derived name for this same system but do not know what it resolved to -- for example, ``job://53/default`` picks up whatever ``Co_q0_m4``-style label job 53 auto-derived for this atom. This needs SEAMM's managed ``Jobs//Job_NNNNNN`` directory layout (i.e. jobs run through the dashboard/JobServer, not a bare local run) to locate the other job. For a basis-set-escalation loop (e.g. looping ``def2-SVP`` -> ``def2-TZVP`` -> ``def2-QZVP`` for one atom), turn on both **Save orbital checkpoint** and **Initial guess** = **Specified orbitals** (leaving **Checkpoint name** / **Specified orbitals** on ``default`` so they agree automatically), and set **If wavefunction not found** to a fallback guess so the first iteration -- which has nothing to read yet -- does not error. Loop from the smallest basis to the largest: projecting a converged small-basis density up into a larger virtual space is cheap and well-conditioned, which is the direction ORCA's own CBS extrapolation uses internally. Extra keywords and extra ORCA blocks ===================================== Two free-text escape hatches cover anything without a dedicated control: * **Extra keywords** -- additional ORCA ``!`` keywords, appended after everything the other controls generate (e.g. ``RIJCOSX``, ``NoFrozenCore``, ``SlowConv``). A keyword that repeats one the controls already produce (say ``TightSCF`` with the SCF convergence at ``TIGHTSCF``) is dropped, since ORCA refuses repeats. An SCF-convergence or integration-grid preset given here overrides the control's, with a note in the output. * **Extra ORCA blocks** -- literal ORCA input (one or more ``%`` blocks), inserted verbatim right before the geometry, after the blocks the controls above generate. Use this for one-off SCF-stabilization tricks with no dedicated control, for example on a difficult atom:: %scf MaxIter 400 Shift Shift 0.3 ErrStart 0.05 end DIISBfac 1.1 end Energies, gradients, and forces =============================== The **Results** tab lists everything the step can produce. Tick a result to save it to a variable, a table, or the structure's property database (see below). Requesting **gradients** makes the step compute the nuclear gradient (i.e. the forces). ORCA computes an *analytic* gradient (``EnGrad``) when one exists for the chosen method, and automatically falls back to a *numerical* gradient (``EnGrad NumGrad``) when it does not — for example ``DLPNO-CCSD(T)`` (no analytic ``(T)`` gradient), the non-self-consistent ``wB97M(2)`` / ``wB97X-2`` double hybrids, and ``PWPB95``, ``DSD-PBEB95``, ``KPR2SCAN50`` and ``WPR2SCAN50``, whose analytic gradients ORCA 6.1 lacks or cannot compute. A note is printed when the (more expensive) numerical gradient is used. Most functionals — including the standard double hybrids — have analytic gradients, so a single-point DFT run yields the energy **and** forces cheaply, which is convenient for generating machine-learned force-field training data. Geometry optimization ===================== The **Optimization** sub-step adds ORCA's ``Opt`` keyword to the energy run and, when it finishes, stores the optimized geometry. A **Convergence** control selects ORCA's optimization preset (``LooseOpt`` … ``VeryTightOpt``). A **Structure handling** control chooses where the optimized geometry goes. It defaults to *Overwrite the current configuration*, so the optimized structure flows straight into any following sub-step — for example a **Frequencies** calculation runs at the optimized geometry without any extra wiring. You can instead create a new configuration (or a new system and configuration) to keep the starting structure alongside the optimized one, or discard the optimized structure entirely. Frequencies (Hessian and thermochemistry) ========================================== The **Frequencies** sub-step computes the Hessian and the harmonic vibrational frequencies, the IR intensities, and the thermochemistry (zero-point energy, thermal enthalpy, entropy, and Gibbs free energy). The level-of-theory controls are the same as the Energy step, plus two extra controls: * **Second derivatives** — ``analytic`` (ORCA's ``AnFreq``; much faster, needs an analytic second derivative for the method — HF, most DFT functionals, MP2) or ``numerical`` (``NumFreq``; finite-differences the gradient, works for any method with a gradient but is considerably more expensive). * **Temperature** — the temperature for the thermochemistry (the pressure is ORCA's default of 1 atm). * **Structure handling** — where to store the structure and its properties. A frequency calculation does not change the geometry, so this defaults to *Overwrite the current configuration*; you can instead create a new configuration (or a new system and configuration) to hold the results, or discard the structure. Tick the results you want on the Results tab: the **frequencies**, the **IR intensities**, the **zero-point energy**, the **enthalpy**, the **Gibbs free energy**, the **number of imaginary frequencies**, and the **largest zero-mode frequency**. Imaginary (negative) frequencies are reported and flagged — a minimum has none, a transition state has one. The output lists the vibrational frequencies and their IR intensities in a table. The frequencies (and IR intensities) are also written to ``frequencies.csv`` in the step's directory for easy access, and an ``IR_spectrum.graph`` file is written with the IR spectrum as a stick trace plus a Lorentzian-broadened trace (FWHM ≈ 15 cm⁻¹) that mimics an experimental spectrum. Open the ``.graph`` file in the SEAMM dashboard to view it interactively. .. note:: **Units.** Following SEAMM's SI-based convention, the zero-point energy, enthalpy, and Gibbs free energy are reported in **kJ/mol**. (Orbital energies — HOMO/LUMO — are reported in **eV** throughout ORCA, the conventional unit for orbital energies.) The report also lists the **largest of the 5 or 6 nominally-zero translation/rotation frequencies**. These should be zero; how far the largest one departs from zero is a convenient gauge of the numerical accuracy of the Hessian. ORCA *projects* the translations and rotations out of its printed frequencies (it reports them as exactly ``0.00``), so this residual is instead obtained by diagonalizing the **raw, un-projected** mass-weighted Cartesian Hessian from ``orca.hess`` — the value you see is therefore the true numerical residual, not zero. .. note:: The same analytic second derivative also backs the ORCA **MDI engine**: it answers a custom ``; rest"`` and ``auto (2 molecules)`` becomes ``auto (molecules)``. Fragment charges ----------------- .. important:: **Set this explicitly for any ionic system, or check the printed log to confirm the automatic default did the right thing.** A charge given to the wrong fragment is not always caught by validation (see below) — it can silently compute the wrong physics, or surface only as a confusing ORCA SCF failure deep inside one sub-job. **Fragment charges** gives each fragment's formal charge, in the same order as the fragments themselves (the order ``auto`` finds molecules in, or the semicolon-separated groups in **Fragment atoms**) — a comma/space separated list of integers, e.g. ``1, -1`` for a Na+/Cl- pair. Leave it **empty** for the common case: * If every fragment is neutral (the usual case for an H-bonded complex), that is exactly what an empty field gives you. * If the input structure itself carries a per-atom formal charge — for example an ion in an SDF/MOL file, which records the charge on the specific atom it belongs to (an ``M CHG`` record) — each fragment's default charge is instead the **sum of its own atoms' formal charge**. A well-formed ``Na+..H2O`` or ``Na+..Cl-`` structure then gets the correct per-fragment charges (``0``/``+1``, or ``+1``/``-1``) with **no** ``Fragment charges`` entry needed at all — this covers monatomic ions (Na+, Cl-) and polyatomic ions (NH4+, BF4-) alike. An explicit value always overrides charges taken from the structure. .. caution:: **The fragment order is not necessarily the order you're thinking of the ions in, and it is easy to get backwards.** The step always prints exactly what it used, before submitting anything to ORCA:: Fragment 1 has 3 atom(s) (O, H, H), charge +0. Fragment 2 has 1 atom(s) (Na), charge +1. Check this against which fragment you actually meant to charge — especially when typing **Fragment charges** by hand. Swapping a charge between two same-parity fragments (e.g. two different monatomic ions) still sums to the right overall complex charge, so it is *not* always caught by the checks below; it just computes the wrong physics. Two checks run before anything is submitted to ORCA, stopping the run with a clear message if they fail: * the fragment charges must sum to the complex's own overall charge; * each fragment, at its assigned charge, must have an **even** number of electrons. Every fragment is closed-shell (multiplicity 1), and an odd-electron fragment cannot be — this is what catches a charge given to the wrong fragment when it does *not* happen to share parity with the fragment it should have gone to (e.g. a neutral water fragment mistakenly given an ion's +1 charge). Other controls --------------- * **Compute the gradient** — ``yes`` (the default) computes the counterpoise-corrected gradient (forces) as well as the energy, which is what MLFF training needs. ``no`` (energy only) is cheaper and, importantly, allows methods that have **no analytic gradient** in ORCA — notably ``CCSD(T)`` and ``DLPNO-CCSD(T)`` — for gold-standard counterpoise *interaction energies*. * **Optimize free monomers** — whether to relax each isolated fragment before taking the correction. Leave it ``no`` for a fixed-geometry PES / MLFF target. * **Write the wavefunction (wfx) file** — default ``no``. When ``yes``, the full cluster's density is kept and converted (via ``orca_2aim``) to an ``orca.wfx``, so a following **Atomic Charges** step can partition it into DDEC6 charges on the CP complex — exactly as after an Energy step. The step reports, and offers on the Results tab, the **BSSE-corrected energy** (the ``energy`` result), the **uncorrected** (raw) complex energy, the **BSSE correction** (corrected minus uncorrected, also shown in kcal/mol), and — when requested — the corrected **gradient**. Tick any of them on the Results tab to save it to a variable, table, or the property database. A ghost-centre gradient-noise warning --------------------------------------- Occasionally the step's output includes a warning like this:: WARNING: the BSSE-corrected gradient failed the translational-invariance guard (net force 0.001290 E_h/bohr > tolerance 0.000300 E_h/bohr), most likely a ghost-centre integration blowup. Falling back to the uncorrected cluster gradient for the forces; the energy is still fully counterpoise-corrected. The BSSE gradient correction being dropped had magnitude 0.001242 E_h/bohr -- small relative to typical forces means the fallback is almost certainly fine; if it is not small, treat this point as suspect (exclude/rerun) rather than trusting either gradient. **What causes it.** One of the ``fragment-in-cluster`` sub-jobs — a fragment's real atoms plus the other fragment(s) as ghost (basis-function only, no electron) centres — can occasionally pick up a small amount of numerical-integration noise in the Pulay (basis-derivative) force on a ghost centre, from ORCA's RIJCOSX/COSX approximation. This is a genuine, if uncommon, ORCA numerical artifact, not a SEAMM bug and (usually) not a sign that anything is chemically wrong with the point. A correctly-assembled counterpoise gradient always sums to exactly zero net force over the whole cluster (translational invariance — nothing external is acting on an isolated system); this noise breaks that, which is what makes it detectable at all — the energy stays fine, with no warning from ORCA itself. **What SEAMM does about it.** Every BSSE gradient is checked against this zero-net-force requirement. When the residual exceeds a small tolerance, the corrected **gradient** (only) falls back to the raw, uncorrected cluster gradient; the **energy** is unaffected either way, since the noise only shows up in the numerically far more sensitive gradient. **How to judge whether the fallback is safe for a given point.** The warning reports both the net-force residual that tripped the guard *and* the size of the BSSE gradient correction that got dropped. Compare them: * If the two are close in magnitude (as in the example above — 0.00129 vs. 0.00124 E_h/bohr), the "correction" being discarded was itself mostly noise, not real physics — the fallback is safe, and the point can be used as-is. * If the correction magnitude is **much larger** than the net-force residual, real physics is being discarded, not just noise — that point is genuinely suspect and worth excluding or rerunning (e.g. with a finer **Integration grid**, ``DEFGRID3`` — see *Integration grid* above) rather than trusting either gradient at face value. In practice this fires rarely, and mostly at short-to-moderate fragment separation (where the counterpoise correction is largest, so there is more for the noise to hide inside); a finer integration grid on the affected structures is usually enough to avoid it entirely. .. note:: **Current limitations.** Every fragment must be **closed-shell** (multiplicity 1); open-shell fragments are not yet supported. Only an **ORCA-internal** basis set works (not yet the Basis Set Exchange). Computing the **gradient** additionally requires an **analytic** gradient, so the numerical-gradient methods (``(DLPNO-)CCSD(T)``) are available in *energy-only* mode only. The SThresh control and the extra property analyses (bond orders, Hirshfeld charges, polarizability) available on the plain Energy step are not available here. Saving results to the database ============================== Both scalar results (the energies, HOMO/LUMO energies and gap, the dipole magnitude, ````, the polarizability) and array results (the **gradient**, the dipole-moment vector, the Mulliken/Löwdin/Hirshfeld charges, the Mayer valences, and the rotational constants) can be stored as properties on the configuration. In the Results tab, tick the database column for the result; it is stored under a name like ``gradients#ORCA#``, where ```` is the level of theory. Other properties and outputs ============================ The Energy step can also compute Mayer bond orders and Hirshfeld charges (and optionally apply them to the structure), the dipole polarizability, and an analytic wavefunction (``.wfx``, via ``orca_2aim``) for a following **Atomic Charges** step to partition into DDEC6 charges. Every run cites ORCA, the DFT functional, the basis set, and the supporting integral / exchange-correlation libraries. Running ORCA in parallel ======================== By default ORCA uses all the cores the machine or batch job provides. Settings come from two files with different jobs: * **How to run ORCA** lives in ``~/SEAMM/orca.ini`` -- the full path to the executable and the OpenMPI ``library-path`` for parallel runs: .. code-block:: ini [local] installation = local code = /path/to/orca library-path = /path/to/orca-mpi/lib * **User run options** live in the ``[orca-step]`` section of the main SEAMM configuration (``~/.seamm.d/seamm.ini``), and can also be given on the command line: .. code-block:: ini [orca-step] ncores = available # or an integer, or 1 to force serial memory = available # or 'all', or e.g. '3 GB' (per process) * **ncores** — how many processes ORCA may use (its ``%pal``). ``available`` (the default) uses all cores the machine/job provides; give an integer to cap it, or ``1`` to force serial. * **memory** — the per-process memory for ORCA's ``%maxcore``. ``available`` (the default) scales to the memory per core; ``all`` divides the whole node among the processes; or give an explicit amount such as ``3 GB``. * **library-path** (in ``orca.ini``) — the ``lib`` directory of the OpenMPI that **matches the version ORCA was built against**. ORCA 6.1 requires OpenMPI 4.1.x and does *not* support 5.x; mixing versions makes parallel runs abort with a ``BLAS-ERROR``. A dedicated conda env is the easy way to get the right one:: conda create -n orca-mpi -c conda-forge "openmpi=4.1" then set ``library-path`` to that env's ``lib`` directory. The step automatically puts the matching ``mpirun`` (the sibling ``bin`` directory) on ``PATH`` so ORCA launches its workers with the correct OpenMPI. .. note:: **macOS.** ORCA does not pass the ``DYLD_*`` loader variables to the MPI processes it spawns, so ``library-path`` alone does not let the dynamic loader find ``libmpi`` on a Mac. Put the OpenMPI libraries on the default search path once, for example by symlinking into ``/usr/local/lib``:: ln -s /path/to/orca-mpi/lib/libmpi.40.dylib /usr/local/lib/ On Linux the exported ``LD_LIBRARY_PATH`` is inherited normally, so ``library-path`` is sufficient there. If you do not have a matching OpenMPI, set ``ncores = 1`` to run serially. **Environment-modules clusters.** If ORCA is provided via a cluster's ``module`` system rather than a fixed path you manage yourself, use ``installation = modules`` instead of ``local``, naming the module(s) to load: .. code-block:: ini [local] installation = modules modules = ORCA code = /path/to/orca library-path = /path/to/orca-mpi/lib ``seamm_exec`` runs ``module load `` before invoking ORCA. This works for both serial and parallel runs -- as of this release, a parallel run's OpenMPI ``library-path`` is prepended onto whatever the module load already put on ``LD_LIBRARY_PATH``/``DYLD_LIBRARY_PATH``, rather than overwriting it (a real, fixed bug: earlier releases could silently discard the module-provided ORCA libraries on a parallel run, failing with ``error while loading shared libraries`` even though the module load itself had succeeded). Rerunning a job =============== ORCA runs through SEAMM's task layer, which keeps a record of each ORCA calculation in ``tasks/manifest.json`` in the step's directory, with a ``tasks//DONE`` file for each one that finished. When a job is run again in the same directory -- after it was stopped, lost its node, or ran out of time -- a calculation that had finished with the same input is not repeated: its results are read from the files it left, and the output says:: ORCA had already finished this calculation, with the same input, in ; using those results. A calculation is run again when * its input changed (the geometry, method, basis, keywords or blocks; the ``%pal`` and ``%maxcore`` lines are ignored, so a different number of cores or amount of memory does not count), or * it did not finish: it failed, or the job was stopped while it was running. ORCA reports success to the operating system even after an "error termination", so a run counts as finished only if ``orca.out`` ends with ``ORCA TERMINATED NORMALLY``. A calculation is tried at most three times in all; after that the step stops with a message saying so. Changing its input, or deleting its entry in ``tasks/manifest.json``, lets it run again. If the job was killed outright, an ORCA process from the earlier run may still be running; the rerun stops it before starting the calculation again. Where ORCA runs =============== An ORCA calculation names only the program; how to run ORCA is worked out on the machine that runs it, from that machine's ``orca.ini`` (``[local]``: the path to ORCA, or ``installation = modules`` with the modules to load, and ``library-path`` for the OpenMPI ORCA was built with), falling back to ``orca`` on the PATH. So a job whose target sends its calculations to a cluster runs ORCA with the cluster's own ORCA, even if ORCA is installed differently (or not at all) where the job runs. Calculations estimated to take under a minute stay on the job's own machine when ORCA is installed there. On a site that provides ORCA through environment modules, give ``code`` as ORCA's **full path** in ``orca.ini`` (with ``installation = modules``): ORCA must be invoked by its full path to find its sub-programs for parallel runs, and a bare ``orca`` cannot be expanded before the module is loaded. The sub-jobs of a counterpoise (BSSE) correction run together: concurrently on this machine when there are cores for more than one, or as separate calculations on the job's cluster queue. ORCA as a model chemistry ========================= Steps that set up a *model chemistry* and then evaluate it at many structures -- the **Energy** step, the **Dimer Builder**'s energy-based contact search, **Normal Mode Sampling** and LAMMPS QM-MD -- use ORCA without any settings in the ORCA step itself: put a **Model Chemistry** step in the flowchart and choose an ORCA model chemistry there (e.g. ``ORCA:DFT@B3LYP/def2-SVP``). How ORCA then runs is chosen for you, and gives the same numbers either way: * **As separate calculations**, one per structure, for the Energy step, the Dimer Builder and Normal Mode Sampling's finite-difference Hessian. They run several at a time on this machine, or on the cluster when the job's target sends its calculations to a queue, and a rerun of the job reuses those that had finished. ORCA starts afresh for each structure in any case, so running them side by side is usually faster than one after another in an engine. Periodic structures are not run this way. * **As a persistent** `MDI `_ **engine** for steps that drive ORCA one geometry after another (LAMMPS QM-MD, Normal Mode Sampling's analytic Hessian on this machine). The model chemistry's basis is used either way: any basis ORCA knows, or a basis from the Basis Set Exchange (``bse:NAME``), whose definition is fetched for the elements present and passed to ORCA exactly as the ORCA step does. A lone atom or ion runs without the COSX approximation (``NoCOSX``), as in the ORCA step, which avoids an artifact of several kJ/mol from COSX's grid on a single centre. Every calculation run this way uses tight SCF convergence and the fine integration grid (``TIGHTSCF DEFGRID3``). These calculations are mostly forces for training data or for derivatives, which need both, rather than ORCA's looser defaults. Functionals without a dispersion correction of their own are also offered with Grimme's D4 correction, as ``-D4`` -- e.g. ``ORCA:DFT@R2SCAN-D4/def2-TZVP``, which runs ``! R2SCAN D4``. They are offered for the functionals the D4 model has parameters for; double hybrids are not, and functionals that already include dispersion (``WB97X-D4``, ``B97M-V``, ``R2SCAN-3C``, ...) are offered only as themselves. Method names are matched as written, so use the capitals shown in the Model Chemistry step's list. Driving ORCA as an MDI engine ----------------------------- Because ORCA has no in-process interface, the engine runs the ``orca`` binary once per geometry in a persistent working directory, **reusing the previous geometry's orbitals** (``orca.gbw``) as the SCF guess -- the main saving for a series of nearby structures. Two consequences: * **Only methods with an analytic gradient are offered via MDI.** The engine always computes the energy and forces together (``EnGrad``), so ``DLPNO-CCSD(T)``, ``CCSD(T)`` and the non-self-consistent ``wB97M(2)`` / ``wB97X-2`` are *not* MDI-capable; choosing one gives a clear error. Everything else (HF, MP2, and the analytic-gradient functionals) works. * **ORCA is not start-up-dominated**, so the per-geometry cost is real -- the MDI benefit here is orbital reuse and a uniform interface, not the large speed-up that cheap engines see. Use an inexpensive functional for jobs that only need the energy surface to guide them (such as contact finding); reserve expensive, high-accuracy calculations for ordinary single-point steps. Settings that depend on each other ================================== Which settings apply depends on others: the functional and its type only for DFT, basis- set extrapolation not for F12 methods or for steps that need a gradient or Hessian, and only F12 bases for an F12 method. The step's dialogs show only the settings that apply with the current choices, and the same rules are used when a flowchart is built or edited without the editor (``seamm-flowchart`` or SEAMM's MCP server): a setting that would have no effect is refused, with the reason, and a value that contradicts another is refused too. See "Flowcharts without the editor" in SEAMM's user guide. Frequencies and BSSE do not use some of the Energy settings, and say so if one is set. Indices and tables ================== * :ref:`genindex` * :ref:`search`