.. _user-guide:
**********
User Guide
**********
The ORCA plug-in runs `ORCA `_ from a SEAMM
flowchart. It is a *sub-flowchart* plug-in: dropping an **ORCA** step onto the
canvas opens a small flowchart of ORCA sub-steps. Four are available:
* **Energy** — a single-point energy, optionally with the gradient (forces) and
a range of properties.
* **Optimization** — a geometry optimization (the Energy step plus ORCA's
``Opt`` keyword).
* **Frequencies** — the Hessian and harmonic vibrational frequencies, with
thermochemistry (see below).
* **BSSE** — the counterpoise-corrected energy and gradient of a two-fragment
complex (see below).
All four share the same controls for the level of theory, so the sections below
apply to any of them.
Choosing the level of theory
============================
Each sub-step can take its method and basis set either from a preceding
**Model Chemistry** step (the default — leave *Use the global model chemistry*
on) or set them explicitly. Turn *Use the global model chemistry* off to choose
them in the step itself.
Method
------
The **Method** pull-down offers Hartree–Fock, MP2/RI-MP2, coupled cluster
(``CCSD(T)``, ``DLPNO-CCSD(T)``), the explicitly-correlated **F12** variants
(``CCSD(T)-F12D/RI``, ``DLPNO-CCSD(T)-F12D``), and **DFT**. For everything except
DFT the method *is* the ORCA keyword. Choosing **DFT** reveals two further,
indented controls:
* **Functional type** — ORCA's own classification of the functional: local,
GGA, meta-GGA, (range-separated) hybrid, and (range-separated) double-hybrid.
* **Functional** — the functional itself, filtered to the chosen type.
Every functional ORCA documents is available, grouped by type; picking a type
narrows the functional list so it stays readable. Double-hybrid functionals
(for example ``REVDSD-PBEP86-D4/2021`` or ``B2PLYP``) need an auxiliary ``/C``
fitting basis for their MP2 part — leave the *Auxiliary (fitting) basis* on
``AutoAux`` and it is supplied automatically.
Most double hybrids also come in a **DLPNO** variant, listed after the canonical
ones (e.g. ``DLPNO-REVDSD-PBEP86-D4/2021``). It evaluates the MP2 part with
near-linear-scaling DLPNO-MP2 instead of canonical RI-MP2, so it pays off for
large systems and clusters; with the default thresholds it differs from the
canonical result by roughly 0.1 kJ/mol. ORCA rejects some ``DLPNO-`` keywords
even though its manual lists them (``DLPNO-REVDSD-PBEP86-D4/2021`` is one), so
the step writes the canonical functional plus ``%mp2 DLPNO true end``, which
works for all of them. ORCA 6.1 can compute DLPNO-MP2 gradients only for
closed-shell systems. An open-shell DLPNO double hybrid can therefore give an
energy but not forces, an optimization, or frequencies, and the step stops
with an error if you ask for any of those. Energies of formation for a DLPNO
double hybrid use the canonical functional's atomic reference energies (DLPNO
changes an isolated atom's energy by less than 0.07 kJ/mol, and cannot treat
H at all), and the thermochemistry report says so.
Basis set
---------
Type any basis name ORCA knows internally, pick one of the curated families
from the list (Pople, Dunning correlation-consistent, and Karlsruhe def2,
including the diffuse ``…D`` and minimally-augmented ``ma-…`` variants), or
press **...** to browse the `Basis Set Exchange
`_ on a periodic table, filtered to the
elements in your system. Selecting **Basis Set Exchange** as the *Basis set
source* opens the same picker. A basis chosen from the Exchange is stored as
``bse:NAME`` and embedded in the ORCA input, so the definition is identical
across codes. The list is ordered by family and, within a family, into
valence / polarization / diffuse ladders that each rise DZ → TZ → QZ → 5Z, so a
sensible progression is a single ladder read top to bottom.
Explicitly-correlated (F12) methods
-----------------------------------
The F12 methods (``CCSD(T)-F12D/RI``, ``DLPNO-CCSD(T)-F12D``) reach near
complete-basis-set accuracy from a modest basis, but they **must** be paired
with one of ORCA's F12-optimized orbital bases — ``cc-pVDZ-F12``,
``cc-pVTZ-F12``, or ``cc-pVQZ-F12`` — and each of those needs a matching
*complementary auxiliary basis set* (CABS). The step **adds the CABS
automatically**: it derives ``-CABS`` from the F12 basis you choose (so
``cc-pVTZ-F12`` → ``cc-pVTZ-F12-CABS``, and likewise for DZ/QZ), unless you have
already put a CABS in the extra keywords. Leave the auxiliary basis on
``AutoAux`` — it supplies the remaining RI fitting bases; the CABS is separate
and is not something AutoAux can generate.
To keep this foolproof, **selecting an F12 method narrows the basis-set list to
just the F12 bases** (and forces the ORCA-internal source), so you can only pick
a valid one; switching back to a non-F12 method restores the full list. If a
hand-edited flowchart still pairs an F12 method with a non-F12 basis (and no
CABS in the extra keywords), the run is stopped early with a clear message.
.. tip::
F12 methods largely **eliminate basis-set superposition error**, so the
counterpoise (BSSE) correction is essentially zero for them (often a few
thousandths of a kcal/mol — at the numerical-noise level). With F12 you can
skip the **BSSE** sub-step and take the interaction energy directly from a
plain **Energy** run, which is ~5× cheaper. Reserve the BSSE step for
conventional finite-basis methods (HF, DFT, MP2, canonical CCSD(T)), where the
correction is real.
Complete-basis-set (CBS) extrapolation
--------------------------------------
Set **Basis-set extrapolation (CBS)** to ``2/3``, ``3/4``, or ``4/5`` to
extrapolate the energy to the complete-basis-set limit from two successive
cardinal numbers (double/triple, triple/quadruple, quadruple/quintuple), using
the **Extrapolation family** (``cc``, ``aug-cc``, ``def2``, or ``ANO``). This
is a *single* ORCA job (its ``Extrapolate`` keyword): ORCA runs both basis sets
and extrapolates the SCF and correlation contributions separately. When it is
on, the fixed basis set above is ignored. Note that an extrapolated energy has
**no gradient** in ORCA, so extrapolation and the *gradients* result are
mutually exclusive.
Integration grid
================
The **Integration grid** control sets ORCA's numerical-integration grid preset
(used for the DFT exchange-correlation integrals and the RIJCOSX/COSX grids):
* ``default`` — leave ORCA's own default (``DEFGRID2``), which is robust for most
work;
* ``DEFGRID1`` — coarser and faster (roughly the old ORCA-4 default accuracy);
* ``DEFGRID2`` — the current default;
* ``DEFGRID3`` — finer and more conservative, for cases sensitive to the grid
(e.g. dispersion-dominated energies, tight convergence, or numerical noise in
forces).
It only affects methods that use a grid (DFT, RIJCOSX) and is ignored otherwise.
For a machine-learned-force-field training set, a finer grid (``DEFGRID3``) can
be worth the cost to keep the energy/force surface smooth.
When the control is left on ``default``, the step **automatically switches to
``DEFGRID3`` for basis sets with high angular momentum** — h functions or above,
such as ``cc-pV5Z`` — which the default ``DEFGRID2`` integrates less accurately.
The angular momentum is read from the Basis Set Exchange (which also covers
ORCA's internal basis names), so it works for both internal and BSE bases;
choosing a grid explicitly always overrides this.
SCF convergence
===============
The **SCF convergence** control sets ORCA's SCF convergence-tolerance preset,
from loosest to tightest: ``SLOPPYSCF`` → ``LOOSESCF`` → ``NORMALSCF`` →
``STRONGSCF`` → ``TIGHTSCF`` → ``VERYTIGHTSCF`` → ``EXTREMESCF``. ``default``
leaves ORCA's own default (``NORMALSCF`` for a single point; ORCA already
tightens it to ``TIGHTSCF`` for optimizations).
The step defaults to **``TIGHTSCF``**, which is a good choice for smooth
energies and forces (and matches what ORCA uses for optimizations). Loosen it
only to save time when high precision is not needed; tighten it (``VERYTIGHTSCF``
/ ``EXTREMESCF``) for very-high-accuracy or numerically delicate work.
SCF SThresh (linear dependence)
===============================
The **SCF SThresh** control sets ORCA's ``%scf SThresh`` value — the threshold
below which an eigenvalue of the overlap matrix is treated as zero, so the
corresponding (near-)linearly-dependent combination of basis functions is
dropped from the calculation. Redundant, near-linearly-dependent basis functions
cause numerical instabilities in the SCF, and diffuse-heavy sets (the augmented
``aug-cc-pVXZ`` / ``ma-…`` families in particular) are the usual culprits.
Raising ``SThresh`` above the default removes more of these functions and can
cure the resulting SCF trouble.
* ``default`` — emit nothing, so the SCF-convergence preset above (or ORCA's own
default of ``1.0e-07``) governs ``SThresh``.
* an explicit value — written to ``%scf SThresh``, overriding whatever the preset
would otherwise set. ORCA recommends keeping it between ``1e-8`` and ``1e-5``;
raise it toward ``1e-6`` to shed near-dependent functions when a diffuse basis
will not converge.
.. caution::
Values beyond ``1e-6`` must be used carefully in **geometry optimizations**
and when **comparing conformers**: because different basis functions can be
cut off at different geometries, the final basis set — and hence the energy —
can vary discontinuously from one structure to the next.
See the ORCA manual's `basis-set section
`_
for the full discussion of linear dependence and its automatic removal.
Initial guess and wavefunction restart
=======================================
Single atoms and open-shell transition metals are often the hardest systems to
converge -- a superposition of atomic densities (ORCA's default ``SAD`` guess)
has little meaning for one atomic center, and can converge to the wrong
electronic state entirely. (Whatever the guess, a single-atom system is run with
exact exchange, ``NoCOSX``: ORCA's default COSX approximation can build wrong
virtual orbitals for a lone atom, which puts the MP2 part of a double hybrid off
by up to ~5 kJ/mol, e.g. for Na, Na\ :sup:`+` and Mg, with no warning. This also
applies to the bare single-atom fragments of a counterpoise correction. An
exchange scheme given in the extra keywords is respected.) The **Initial
guess** control sets ORCA's SCF starting guess (``Guess`` in the ``%scf``
block):
* ``default`` -- leave ORCA's own default.
* ``Hueckel``, ``HCore``, ``PAtom``, ``PModel``, ``SAD``, ``SADNO`` -- ORCA's
built-in guess types. ``PModel`` is often much more stable than ``SAD`` for a
single atom or ion.
* **Previous wavefunction** -- seed from the orbitals (``orca.gbw``) left by the
*nearest earlier ORCA step in this flowchart* -- for example, run a cheap,
robust functional (``PBE``) immediately before a fragile double-hybrid
(``REVDSD-PBEP86-D4/2021``) on the same atom, and turn this on for the
double-hybrid step. This only reaches a *different, preceding* node in a
roughly linear flowchart -- it cannot seed across iterations of a **Loop**,
since a Loop gives each iteration its own directory (see below).
* **Specified orbitals** -- seed from a named checkpoint file instead, named by
the **Specified orbitals** control revealed underneath. This is the way to
seed across loop iterations: for example, a basis-set escalation inside a
Loop over basis sets, where each iteration reads the checkpoint the
*previous* iteration (a smaller basis) wrote.
Either wavefunction choice reveals **If wavefunction not found**, controlling
what happens when there is nothing to read (e.g. the very first ORCA step, or
the first pass through a basis-set-escalation loop): ``Throw an error`` (the
default -- an explicit wavefunction request that cannot be honored is treated
as a configuration problem) or one of the ``Use ... guess`` choices, which
falls back to that guess type instead. Set this to a fallback guess to make
"Specified orbitals" safe to leave on for every iteration of a loop.
.. note::
**ORCA projects across basis sets.** ``MORead`` is not limited to restarting
with the *same* basis set -- when the checkpoint's basis differs from the
current job's, ORCA projects the old orbitals onto the new basis (the same
trick ORCA's own ``Extrapolate`` (CBS) keyword uses internally). So this also
works for a basis-set escalation, not just same-basis restarts across
different functionals. It should still be the same atom/system in the same
electronic state (charge and multiplicity) -- MORead supplies starting
orbitals, it does not reinterpret the electronic state.
Saving a checkpoint for a later step to read is a separate pair of controls:
* **Save orbital checkpoint** -- ``yes`` copies this run's converged orbitals to
the file named by **Checkpoint name**, revealed underneath, once the run
finishes successfully.
* **Checkpoint name** and **Specified orbitals** both resolve the same way:
``default`` (the default value of each) derives a label automatically from
the system's composition, charge, and multiplicity (e.g. ``Co_q0_m4``) -- the
usual choice, since it naturally gives one checkpoint per atom/electronic
state, shared across an inner loop (e.g. over basis sets) for that atom, and
reset automatically when an outer loop moves to a new atom. Any other bare
name is stored under a ``checkpoints`` folder at the top of the job; an
absolute path is used as-is, e.g. to keep a checkpoint outside this job and
reuse it across separate flowchart runs.
.. note::
**Reading a checkpoint from another job.** Only **Specified orbitals** can
reference *another* job -- ``job:///`` -- since a job
must never write into another job (so this is not available for
**Checkpoint name**). Use ``job:///default`` when you want that
other job's own auto-derived name for this same system but do not know
what it resolved to -- for example, ``job://53/default`` picks up whatever
``Co_q0_m4``-style label job 53 auto-derived for this atom. This needs
SEAMM's managed ``Jobs//Job_NNNNNN`` directory layout (i.e. jobs
run through the dashboard/JobServer, not a bare local run) to locate the
other job.
For a basis-set-escalation loop (e.g. looping ``def2-SVP`` -> ``def2-TZVP`` ->
``def2-QZVP`` for one atom), turn on both **Save orbital checkpoint** and
**Initial guess** = **Specified orbitals** (leaving **Checkpoint name** /
**Specified orbitals** on ``default`` so they agree automatically), and set
**If wavefunction not found** to a fallback guess so the first iteration -- which
has nothing to read yet -- does not error. Loop from the smallest basis to the
largest: projecting a converged small-basis density up into a larger virtual
space is cheap and well-conditioned, which is the direction ORCA's own CBS
extrapolation uses internally.
Extra keywords and extra ORCA blocks
=====================================
Two free-text escape hatches cover anything without a dedicated control:
* **Extra keywords** -- additional ORCA ``!`` keywords, appended after
everything the other controls generate (e.g. ``RIJCOSX``, ``NoFrozenCore``,
``SlowConv``). A keyword that repeats one the controls already produce (say
``TightSCF`` with the SCF convergence at ``TIGHTSCF``) is dropped, since ORCA
refuses repeats. An SCF-convergence or integration-grid preset given here
overrides the control's, with a note in the output.
* **Extra ORCA blocks** -- literal ORCA input (one or more ``%`` blocks),
inserted verbatim right before the geometry, after the blocks the controls
above generate. Use this for one-off SCF-stabilization tricks with no
dedicated control, for example on a difficult atom::
%scf
MaxIter 400
Shift Shift 0.3 ErrStart 0.05 end
DIISBfac 1.1
end
Energies, gradients, and forces
===============================
The **Results** tab lists everything the step can produce. Tick a result to
save it to a variable, a table, or the structure's property database (see
below).
Requesting **gradients** makes the step compute the nuclear gradient (i.e. the
forces). ORCA computes an *analytic* gradient (``EnGrad``) when one exists for
the chosen method, and automatically falls back to a *numerical* gradient
(``EnGrad NumGrad``) when it does not — for example ``DLPNO-CCSD(T)`` (no
analytic ``(T)`` gradient), the non-self-consistent ``wB97M(2)`` / ``wB97X-2``
double hybrids, and ``PWPB95``, ``DSD-PBEB95``, ``KPR2SCAN50`` and
``WPR2SCAN50``, whose analytic gradients ORCA 6.1 lacks or cannot compute. A
note is printed when the (more expensive) numerical gradient is used.
Most functionals — including the standard double hybrids — have analytic
gradients, so a single-point DFT run yields the energy **and** forces cheaply,
which is convenient for generating machine-learned force-field training data.
Geometry optimization
=====================
The **Optimization** sub-step adds ORCA's ``Opt`` keyword to the energy run and,
when it finishes, stores the optimized geometry. A **Convergence** control
selects ORCA's optimization preset (``LooseOpt`` … ``VeryTightOpt``).
A **Structure handling** control chooses where the optimized geometry goes. It
defaults to *Overwrite the current configuration*, so the optimized structure
flows straight into any following sub-step — for example a **Frequencies**
calculation runs at the optimized geometry without any extra wiring. You can
instead create a new configuration (or a new system and configuration) to keep
the starting structure alongside the optimized one, or discard the optimized
structure entirely.
Frequencies (Hessian and thermochemistry)
==========================================
The **Frequencies** sub-step computes the Hessian and the harmonic vibrational
frequencies, the IR intensities, and the thermochemistry (zero-point energy,
thermal enthalpy, entropy, and Gibbs free energy). The level-of-theory controls
are the same as the Energy step, plus two extra controls:
* **Second derivatives** — ``analytic`` (ORCA's ``AnFreq``; much faster, needs
an analytic second derivative for the method — HF, most DFT functionals, MP2)
or ``numerical`` (``NumFreq``; finite-differences the gradient, works for any
method with a gradient but is considerably more expensive).
* **Temperature** — the temperature for the thermochemistry (the pressure is
ORCA's default of 1 atm).
* **Structure handling** — where to store the structure and its properties. A
frequency calculation does not change the geometry, so this defaults to
*Overwrite the current configuration*; you can instead create a new
configuration (or a new system and configuration) to hold the results, or
discard the structure.
Tick the results you want on the Results tab: the **frequencies**, the **IR
intensities**, the **zero-point energy**, the **enthalpy**, the **Gibbs free
energy**, the **number of imaginary frequencies**, and the **largest zero-mode
frequency**. Imaginary (negative) frequencies are reported and flagged — a
minimum has none, a transition state has one.
The output lists the vibrational frequencies and their IR intensities in a
table. The frequencies (and IR intensities) are also written to
``frequencies.csv`` in the step's directory for easy access, and an
``IR_spectrum.graph`` file is written with the IR spectrum as a stick trace plus
a Lorentzian-broadened trace (FWHM ≈ 15 cm⁻¹) that mimics an experimental
spectrum. Open the ``.graph`` file in the SEAMM dashboard to view it
interactively.
.. note::
**Units.** Following SEAMM's SI-based convention, the zero-point energy,
enthalpy, and Gibbs free energy are reported in **kJ/mol**. (Orbital energies
— HOMO/LUMO — are reported in **eV** throughout ORCA, the conventional unit
for orbital energies.)
The report also lists the **largest of the 5 or 6 nominally-zero
translation/rotation frequencies**. These should be zero; how far the largest
one departs from zero is a convenient gauge of the numerical accuracy of the
Hessian. ORCA *projects* the translations and rotations out of its printed
frequencies (it reports them as exactly ``0.00``), so this residual is instead
obtained by diagonalizing the **raw, un-projected** mass-weighted Cartesian
Hessian from ``orca.hess`` — the value you see is therefore the true numerical
residual, not zero.
.. note::
The same analytic second derivative also backs the ORCA **MDI engine**: it
answers a custom ``; rest"`` and ``auto (2 molecules)`` becomes
``auto (molecules)``.
Fragment charges
-----------------
.. important::
**Set this explicitly for any ionic system, or check the printed log to
confirm the automatic default did the right thing.** A charge given to
the wrong fragment is not always caught by validation (see below) — it
can silently compute the wrong physics, or surface only as a confusing
ORCA SCF failure deep inside one sub-job.
**Fragment charges** gives each fragment's formal charge, in the same order
as the fragments themselves (the order ``auto`` finds molecules in, or the
semicolon-separated groups in **Fragment atoms**) — a comma/space separated
list of integers, e.g. ``1, -1`` for a Na+/Cl- pair.
Leave it **empty** for the common case:
* If every fragment is neutral (the usual case for an H-bonded complex),
that is exactly what an empty field gives you.
* If the input structure itself carries a per-atom formal charge — for
example an ion in an SDF/MOL file, which records the charge on the
specific atom it belongs to (an ``M CHG`` record) — each fragment's
default charge is instead the **sum of its own atoms' formal charge**.
A well-formed ``Na+..H2O`` or ``Na+..Cl-`` structure then gets the
correct per-fragment charges (``0``/``+1``, or ``+1``/``-1``) with
**no** ``Fragment charges`` entry needed at all — this covers monatomic
ions (Na+, Cl-) and polyatomic ions (NH4+, BF4-) alike.
An explicit value always overrides charges taken from the structure.
.. caution::
**The fragment order is not necessarily the order you're thinking of the
ions in, and it is easy to get backwards.** The step always prints
exactly what it used, before submitting anything to ORCA::
Fragment 1 has 3 atom(s) (O, H, H), charge +0.
Fragment 2 has 1 atom(s) (Na), charge +1.
Check this against which fragment you actually meant to charge —
especially when typing **Fragment charges** by hand. Swapping a charge
between two same-parity fragments (e.g. two different monatomic ions)
still sums to the right overall complex charge, so it is *not* always
caught by the checks below; it just computes the wrong physics.
Two checks run before anything is submitted to ORCA, stopping the run with a
clear message if they fail:
* the fragment charges must sum to the complex's own overall charge;
* each fragment, at its assigned charge, must have an **even** number of
electrons. Every fragment is closed-shell (multiplicity 1), and an
odd-electron fragment cannot be — this is what catches a charge given to
the wrong fragment when it does *not* happen to share parity with the
fragment it should have gone to (e.g. a neutral water fragment
mistakenly given an ion's +1 charge).
Other controls
---------------
* **Compute the gradient** — ``yes`` (the default) computes the
counterpoise-corrected gradient (forces) as well as the energy, which is
what MLFF training needs. ``no`` (energy only) is cheaper and,
importantly, allows methods that have **no analytic gradient** in ORCA —
notably ``CCSD(T)`` and ``DLPNO-CCSD(T)`` — for gold-standard counterpoise
*interaction energies*.
* **Optimize free monomers** — whether to relax each isolated fragment
before taking the correction. Leave it ``no`` for a fixed-geometry PES /
MLFF target.
* **Write the wavefunction (wfx) file** — default ``no``. When ``yes``, the
full cluster's density is kept and converted (via ``orca_2aim``) to an
``orca.wfx``, so a following **Atomic Charges** step can partition it into
DDEC6 charges on the CP complex — exactly as after an Energy step.
The step reports, and offers on the Results tab, the **BSSE-corrected
energy** (the ``energy`` result), the **uncorrected** (raw) complex energy,
the **BSSE correction** (corrected minus uncorrected, also shown in
kcal/mol), and — when requested — the corrected **gradient**. Tick any of
them on the Results tab to save it to a variable, table, or the property
database.
A ghost-centre gradient-noise warning
---------------------------------------
Occasionally the step's output includes a warning like this::
WARNING: the BSSE-corrected gradient failed the translational-invariance
guard (net force 0.001290 E_h/bohr > tolerance 0.000300 E_h/bohr), most
likely a ghost-centre integration blowup. Falling back to the
uncorrected cluster gradient for the forces; the energy is still fully
counterpoise-corrected. The BSSE gradient correction being dropped had
magnitude 0.001242 E_h/bohr -- small relative to typical forces means
the fallback is almost certainly fine; if it is not small, treat this
point as suspect (exclude/rerun) rather than trusting either gradient.
**What causes it.** One of the ``fragment-in-cluster`` sub-jobs — a
fragment's real atoms plus the other fragment(s) as ghost (basis-function
only, no electron) centres — can occasionally pick up a small amount of
numerical-integration noise in the Pulay (basis-derivative) force on a ghost
centre, from ORCA's RIJCOSX/COSX approximation. This is a genuine, if
uncommon, ORCA numerical artifact, not a SEAMM bug and (usually) not a sign
that anything is chemically wrong with the point. A correctly-assembled
counterpoise gradient always sums to exactly zero net force over the whole
cluster (translational invariance — nothing external is acting on an
isolated system); this noise breaks that, which is what makes it
detectable at all — the energy stays fine, with no warning from ORCA
itself.
**What SEAMM does about it.** Every BSSE gradient is checked against this
zero-net-force requirement. When the residual exceeds a small tolerance,
the corrected **gradient** (only) falls back to the raw, uncorrected
cluster gradient; the **energy** is unaffected either way, since the noise
only shows up in the numerically far more sensitive gradient.
**How to judge whether the fallback is safe for a given point.** The
warning reports both the net-force residual that tripped the guard *and*
the size of the BSSE gradient correction that got dropped. Compare them:
* If the two are close in magnitude (as in the example above — 0.00129 vs.
0.00124 E_h/bohr), the "correction" being discarded was itself mostly
noise, not real physics — the fallback is safe, and the point can be used
as-is.
* If the correction magnitude is **much larger** than the net-force
residual, real physics is being discarded, not just noise — that point is
genuinely suspect and worth excluding or rerunning (e.g. with a finer
**Integration grid**, ``DEFGRID3`` — see *Integration grid* above) rather
than trusting either gradient at face value.
In practice this fires rarely, and mostly at short-to-moderate fragment
separation (where the counterpoise correction is largest, so there is more
for the noise to hide inside); a finer integration grid on the affected
structures is usually enough to avoid it entirely.
.. note::
**Current limitations.** Every fragment must be **closed-shell**
(multiplicity 1); open-shell fragments are not yet supported. Only an
**ORCA-internal** basis set works (not yet the Basis Set Exchange).
Computing the **gradient** additionally requires an **analytic**
gradient, so the numerical-gradient methods (``(DLPNO-)CCSD(T)``) are
available in *energy-only* mode only. The SThresh control and the extra
property analyses (bond orders, Hirshfeld charges, polarizability)
available on the plain Energy step are not available here.
Saving results to the database
==============================
Both scalar results (the energies, HOMO/LUMO energies and gap, the dipole
magnitude, ````, the polarizability) and array results (the **gradient**,
the dipole-moment vector, the Mulliken/Löwdin/Hirshfeld charges, the Mayer
valences, and the rotational constants) can be stored as properties on the
configuration. In the Results tab, tick the database column for the result; it
is stored under a name like ``gradients#ORCA#``, where ```` is the
level of theory.
Other properties and outputs
============================
The Energy step can also compute Mayer bond orders and Hirshfeld charges (and
optionally apply them to the structure), the dipole polarizability, and an
analytic wavefunction (``.wfx``, via ``orca_2aim``) for a following **Atomic
Charges** step to partition into DDEC6 charges. Every run cites ORCA, the DFT
functional, the basis set, and the supporting integral / exchange-correlation
libraries.
Running ORCA in parallel
========================
By default ORCA uses all the cores the machine or batch job provides. Settings
come from two files with different jobs:
* **How to run ORCA** lives in ``~/SEAMM/orca.ini`` -- the full path to the
executable and the OpenMPI ``library-path`` for parallel runs:
.. code-block:: ini
[local]
installation = local
code = /path/to/orca
library-path = /path/to/orca-mpi/lib
* **User run options** live in the ``[orca-step]`` section of the main SEAMM
configuration (``~/.seamm.d/seamm.ini``), and can also be given on the command
line:
.. code-block:: ini
[orca-step]
ncores = available # or an integer, or 1 to force serial
memory = available # or 'all', or e.g. '3 GB' (per process)
* **ncores** — how many processes ORCA may use (its ``%pal``). ``available``
(the default) uses all cores the machine/job provides; give an integer to cap
it, or ``1`` to force serial.
* **memory** — the per-process memory for ORCA's ``%maxcore``. ``available``
(the default) scales to the memory per core; ``all`` divides the whole node
among the processes; or give an explicit amount such as ``3 GB``.
* **library-path** (in ``orca.ini``) — the ``lib`` directory of the OpenMPI that
**matches the version ORCA was built against**. ORCA 6.1 requires OpenMPI
4.1.x and does *not* support 5.x; mixing versions makes parallel runs abort
with a
``BLAS-ERROR``. A dedicated conda env is the easy way to get the right one::
conda create -n orca-mpi -c conda-forge "openmpi=4.1"
then set ``library-path`` to that env's ``lib`` directory. The step
automatically puts the matching ``mpirun`` (the sibling ``bin`` directory) on
``PATH`` so ORCA launches its workers with the correct OpenMPI.
.. note::
**macOS.** ORCA does not pass the ``DYLD_*`` loader variables to the MPI
processes it spawns, so ``library-path`` alone does not let the dynamic
loader find ``libmpi`` on a Mac. Put the OpenMPI libraries on the default
search path once, for example by symlinking into ``/usr/local/lib``::
ln -s /path/to/orca-mpi/lib/libmpi.40.dylib /usr/local/lib/
On Linux the exported ``LD_LIBRARY_PATH`` is inherited normally, so
``library-path`` is sufficient there.
If you do not have a matching OpenMPI, set ``ncores = 1`` to run serially.
**Environment-modules clusters.** If ORCA is provided via a cluster's
``module`` system rather than a fixed path you manage yourself, use
``installation = modules`` instead of ``local``, naming the module(s) to
load:
.. code-block:: ini
[local]
installation = modules
modules = ORCA
code = /path/to/orca
library-path = /path/to/orca-mpi/lib
``seamm_exec`` runs ``module load `` before invoking ORCA. This
works for both serial and parallel runs -- as of this release, a
parallel run's OpenMPI ``library-path`` is prepended onto whatever the
module load already put on ``LD_LIBRARY_PATH``/``DYLD_LIBRARY_PATH``,
rather than overwriting it (a real, fixed bug: earlier releases could
silently discard the module-provided ORCA libraries on a parallel run,
failing with ``error while loading shared libraries`` even though the
module load itself had succeeded).
Rerunning a job
===============
ORCA runs through SEAMM's task layer, which keeps a record of each ORCA calculation in
``tasks/manifest.json`` in the step's directory, with a ``tasks//DONE`` file for
each one that finished. When a job is run again in the same directory -- after it was
stopped, lost its node, or ran out of time -- a calculation that had finished with the
same input is not repeated: its results are read from the files it left, and the output
says::
ORCA had already finished this calculation, with the same input, in ;
using those results.
A calculation is run again when
* its input changed (the geometry, method, basis, keywords or blocks; the ``%pal`` and
``%maxcore`` lines are ignored, so a different number of cores or amount of memory
does not count), or
* it did not finish: it failed, or the job was stopped while it was running. ORCA
reports success to the operating system even after an "error termination", so a run
counts as finished only if ``orca.out`` ends with ``ORCA TERMINATED NORMALLY``. A
calculation is tried at most three times in all; after that the step stops with a
message saying so. Changing its input, or deleting its entry in
``tasks/manifest.json``, lets it run again.
If the job was killed outright, an ORCA process from the earlier run may still be
running; the rerun stops it before starting the calculation again.
Where ORCA runs
===============
An ORCA calculation names only the program; how to run ORCA is worked out on the
machine that runs it, from that machine's ``orca.ini`` (``[local]``: the path to ORCA,
or ``installation = modules`` with the modules to load, and ``library-path`` for the
OpenMPI ORCA was built with), falling back to ``orca`` on the PATH. So a job whose
target sends its calculations to a cluster runs ORCA with the cluster's own ORCA, even
if ORCA is installed differently (or not at all) where the job runs. Calculations
estimated to take under a minute stay on the job's own machine when ORCA is installed
there.
On a site that provides ORCA through environment modules, give ``code`` as ORCA's
**full path** in ``orca.ini`` (with ``installation = modules``): ORCA must be invoked by
its full path to find its sub-programs for parallel runs, and a bare ``orca`` cannot be
expanded before the module is loaded.
The sub-jobs of a counterpoise (BSSE) correction run together: concurrently on this
machine when there are cores for more than one, or as separate calculations on the
job's cluster queue.
ORCA as a model chemistry
=========================
Steps that set up a *model chemistry* and then evaluate it at many structures --
the **Energy** step, the **Dimer Builder**'s energy-based contact search,
**Normal Mode Sampling** and LAMMPS QM-MD -- use ORCA without any settings in the
ORCA step itself: put a **Model Chemistry** step in the flowchart and choose an ORCA
model chemistry there (e.g. ``ORCA:DFT@B3LYP/def2-SVP``).
How ORCA then runs is chosen for you, and gives the same numbers either way:
* **As separate calculations**, one per structure, for the Energy step, the Dimer
Builder and Normal Mode Sampling's finite-difference Hessian. They run several at
a time on this machine, or on the cluster when the job's target sends its
calculations to a queue, and a rerun of the job reuses those that had finished.
ORCA starts afresh for each structure in any case, so running them side by side
is usually faster than one after another in an engine. Periodic structures are not run this way.
* **As a persistent** `MDI `_ **engine**
for steps that drive ORCA one geometry after another (LAMMPS QM-MD, Normal Mode
Sampling's analytic Hessian on this machine).
The model chemistry's basis is used either way: any basis ORCA knows, or a basis
from the Basis Set Exchange (``bse:NAME``), whose definition is fetched for the
elements present and passed to ORCA exactly as the ORCA step does. A lone atom or
ion runs without the COSX approximation (``NoCOSX``), as in the ORCA step, which
avoids an artifact of several kJ/mol from COSX's grid on a single centre.
Every calculation run this way uses tight SCF convergence and the fine integration
grid (``TIGHTSCF DEFGRID3``). These calculations are mostly forces for training
data or for derivatives, which need both, rather than ORCA's looser defaults.
Functionals without a dispersion correction of their own are also offered with
Grimme's D4 correction, as ``-D4`` -- e.g. ``ORCA:DFT@R2SCAN-D4/def2-TZVP``,
which runs ``! R2SCAN D4``. They are offered for the functionals the D4 model has
parameters for; double hybrids are not, and functionals that already include
dispersion (``WB97X-D4``, ``B97M-V``, ``R2SCAN-3C``, ...) are offered only as
themselves. Method names are matched as written, so use the capitals shown in the
Model Chemistry step's list.
Driving ORCA as an MDI engine
-----------------------------
Because ORCA has no in-process interface, the engine runs the ``orca`` binary
once per geometry in a persistent working directory, **reusing the previous
geometry's orbitals** (``orca.gbw``) as the SCF guess -- the main saving for a
series of nearby structures. Two consequences:
* **Only methods with an analytic gradient are offered via MDI.** The engine
always computes the energy and forces together (``EnGrad``), so
``DLPNO-CCSD(T)``, ``CCSD(T)`` and the non-self-consistent ``wB97M(2)`` /
``wB97X-2`` are *not* MDI-capable; choosing one gives a clear error. Everything
else (HF, MP2, and the analytic-gradient functionals) works.
* **ORCA is not start-up-dominated**, so the per-geometry cost is real -- the MDI
benefit here is orbital reuse and a uniform interface, not the large speed-up
that cheap engines see. Use an inexpensive functional for jobs that only need
the energy surface to guide them (such as contact finding); reserve expensive,
high-accuracy calculations for ordinary single-point steps.
Settings that depend on each other
==================================
Which settings apply depends on others: the functional and its type only for DFT, basis-
set extrapolation not for F12 methods or for steps that need a gradient or Hessian, and
only F12 bases for an F12 method. The step's dialogs show only the settings that apply
with the current choices, and the same rules are used when a flowchart is built or
edited without the editor (``seamm-flowchart`` or SEAMM's MCP server): a setting that
would have no effect is refused, with the reason, and a value that contradicts another
is refused too. See "Flowcharts without the editor" in SEAMM's user guide.
Frequencies and BSSE do not use some of the Energy settings, and say so if one is set.
Indices and tables
==================
* :ref:`genindex`
* :ref:`search`