User Guide#

The ORCA plug-in runs ORCA from a SEAMM flowchart. It is a sub-flowchart plug-in: dropping an ORCA step onto the canvas opens a small flowchart of ORCA sub-steps. Four are available:

  • Energy — a single-point energy, optionally with the gradient (forces) and a range of properties.

  • Optimization — a geometry optimization (the Energy step plus ORCA’s Opt keyword).

  • Frequencies — the Hessian and harmonic vibrational frequencies, with thermochemistry (see below).

  • BSSE — the counterpoise-corrected energy and gradient of a two-fragment complex (see below).

All four share the same controls for the level of theory, so the sections below apply to any of them.

Choosing the level of theory#

Each sub-step can take its method and basis set either from a preceding Model Chemistry step (the default — leave Use the global model chemistry on) or set them explicitly. Turn Use the global model chemistry off to choose them in the step itself.

Method#

The Method pull-down offers Hartree–Fock, MP2/RI-MP2, coupled cluster (CCSD(T), DLPNO-CCSD(T)), the explicitly-correlated F12 variants (CCSD(T)-F12D/RI, DLPNO-CCSD(T)-F12D), and DFT. For everything except DFT the method is the ORCA keyword. Choosing DFT reveals two further, indented controls:

  • Functional type — ORCA’s own classification of the functional: local, GGA, meta-GGA, (range-separated) hybrid, and (range-separated) double-hybrid.

  • Functional — the functional itself, filtered to the chosen type.

Every functional ORCA documents is available, grouped by type; picking a type narrows the functional list so it stays readable. Double-hybrid functionals (for example REVDSD-PBEP86-D4/2021 or B2PLYP) need an auxiliary /C fitting basis for their MP2 part — leave the Auxiliary (fitting) basis on AutoAux and it is supplied automatically.

Most double hybrids also come in a DLPNO variant, listed after the canonical ones (e.g. DLPNO-REVDSD-PBEP86-D4/2021). It evaluates the MP2 part with near-linear-scaling DLPNO-MP2 instead of canonical RI-MP2, so it pays off for large systems and clusters; with the default thresholds it differs from the canonical result by roughly 0.1 kJ/mol. ORCA rejects some DLPNO- keywords even though its manual lists them (DLPNO-REVDSD-PBEP86-D4/2021 is one), so the step writes the canonical functional plus %mp2 DLPNO true end, which works for all of them. ORCA 6.1 can compute DLPNO-MP2 gradients only for closed-shell systems. An open-shell DLPNO double hybrid can therefore give an energy but not forces, an optimization, or frequencies, and the step stops with an error if you ask for any of those. Energies of formation for a DLPNO double hybrid use the canonical functional’s atomic reference energies (DLPNO changes an isolated atom’s energy by less than 0.07 kJ/mol, and cannot treat H at all), and the thermochemistry report says so.

Basis set#

Type any basis name ORCA knows internally, pick one of the curated families from the list (Pople, Dunning correlation-consistent, and Karlsruhe def2, including the diffuse …D and minimally-augmented ma-… variants), or press … to browse the Basis Set Exchange on a periodic table, filtered to the elements in your system. Selecting Basis Set Exchange as the Basis set source opens the same picker. A basis chosen from the Exchange is stored as bse:NAME and embedded in the ORCA input, so the definition is identical across codes. The list is ordered by family and, within a family, into valence / polarization / diffuse ladders that each rise DZ → TZ → QZ → 5Z, so a sensible progression is a single ladder read top to bottom.

Explicitly-correlated (F12) methods#

The F12 methods (CCSD(T)-F12D/RI, DLPNO-CCSD(T)-F12D) reach near complete-basis-set accuracy from a modest basis, but they must be paired with one of ORCA’s F12-optimized orbital bases — cc-pVDZ-F12, cc-pVTZ-F12, or cc-pVQZ-F12 — and each of those needs a matching complementary auxiliary basis set (CABS). The step adds the CABS automatically: it derives <basis>-CABS from the F12 basis you choose (so cc-pVTZ-F12 → cc-pVTZ-F12-CABS, and likewise for DZ/QZ), unless you have already put a CABS in the extra keywords. Leave the auxiliary basis on AutoAux — it supplies the remaining RI fitting bases; the CABS is separate and is not something AutoAux can generate.

To keep this foolproof, selecting an F12 method narrows the basis-set list to just the F12 bases (and forces the ORCA-internal source), so you can only pick a valid one; switching back to a non-F12 method restores the full list. If a hand-edited flowchart still pairs an F12 method with a non-F12 basis (and no CABS in the extra keywords), the run is stopped early with a clear message.

Tip

F12 methods largely eliminate basis-set superposition error, so the counterpoise (BSSE) correction is essentially zero for them (often a few thousandths of a kcal/mol — at the numerical-noise level). With F12 you can skip the BSSE sub-step and take the interaction energy directly from a plain Energy run, which is ~5× cheaper. Reserve the BSSE step for conventional finite-basis methods (HF, DFT, MP2, canonical CCSD(T)), where the correction is real.

Complete-basis-set (CBS) extrapolation#

Set Basis-set extrapolation (CBS) to 2/3, 3/4, or 4/5 to extrapolate the energy to the complete-basis-set limit from two successive cardinal numbers (double/triple, triple/quadruple, quadruple/quintuple), using the Extrapolation family (cc, aug-cc, def2, or ANO). This is a single ORCA job (its Extrapolate keyword): ORCA runs both basis sets and extrapolates the SCF and correlation contributions separately. When it is on, the fixed basis set above is ignored. Note that an extrapolated energy has no gradient in ORCA, so extrapolation and the gradients result are mutually exclusive.

Integration grid#

The Integration grid control sets ORCA’s numerical-integration grid preset (used for the DFT exchange-correlation integrals and the RIJCOSX/COSX grids):

  • default — leave ORCA’s own default (DEFGRID2), which is robust for most work;

  • DEFGRID1 — coarser and faster (roughly the old ORCA-4 default accuracy);

  • DEFGRID2 — the current default;

  • DEFGRID3 — finer and more conservative, for cases sensitive to the grid (e.g. dispersion-dominated energies, tight convergence, or numerical noise in forces).

It only affects methods that use a grid (DFT, RIJCOSX) and is ignored otherwise. For a machine-learned-force-field training set, a finer grid (DEFGRID3) can be worth the cost to keep the energy/force surface smooth.

When the control is left on default, the step automatically switches to ``DEFGRID3`` for basis sets with high angular momentum — h functions or above, such as cc-pV5Z — which the default DEFGRID2 integrates less accurately. The angular momentum is read from the Basis Set Exchange (which also covers ORCA’s internal basis names), so it works for both internal and BSE bases; choosing a grid explicitly always overrides this.

SCF convergence#

The SCF convergence control sets ORCA’s SCF convergence-tolerance preset, from loosest to tightest: SLOPPYSCF → LOOSESCF → NORMALSCF → STRONGSCF → TIGHTSCF → VERYTIGHTSCF → EXTREMESCF. default leaves ORCA’s own default (NORMALSCF for a single point; ORCA already tightens it to TIGHTSCF for optimizations).

The step defaults to ``TIGHTSCF``, which is a good choice for smooth energies and forces (and matches what ORCA uses for optimizations). Loosen it only to save time when high precision is not needed; tighten it (VERYTIGHTSCF / EXTREMESCF) for very-high-accuracy or numerically delicate work.

SCF SThresh (linear dependence)#

The SCF SThresh control sets ORCA’s %scf SThresh value — the threshold below which an eigenvalue of the overlap matrix is treated as zero, so the corresponding (near-)linearly-dependent combination of basis functions is dropped from the calculation. Redundant, near-linearly-dependent basis functions cause numerical instabilities in the SCF, and diffuse-heavy sets (the augmented aug-cc-pVXZ / ma-… families in particular) are the usual culprits. Raising SThresh above the default removes more of these functions and can cure the resulting SCF trouble.

  • default — emit nothing, so the SCF-convergence preset above (or ORCA’s own default of 1.0e-07) governs SThresh.

  • an explicit value — written to %scf SThresh, overriding whatever the preset would otherwise set. ORCA recommends keeping it between 1e-8 and 1e-5; raise it toward 1e-6 to shed near-dependent functions when a diffuse basis will not converge.

Caution

Values beyond 1e-6 must be used carefully in geometry optimizations and when comparing conformers: because different basis functions can be cut off at different geometries, the final basis set — and hence the energy — can vary discontinuously from one structure to the next.

See the ORCA manual’s basis-set section for the full discussion of linear dependence and its automatic removal.

Initial guess and wavefunction restart#

Single atoms and open-shell transition metals are often the hardest systems to converge – a superposition of atomic densities (ORCA’s default SAD guess) has little meaning for one atomic center, and can converge to the wrong electronic state entirely. (Whatever the guess, a single-atom system is run with exact exchange, NoCOSX: ORCA’s default COSX approximation can build wrong virtual orbitals for a lone atom, which puts the MP2 part of a double hybrid off by up to ~5 kJ/mol, e.g. for Na, Na+ and Mg, with no warning. This also applies to the bare single-atom fragments of a counterpoise correction. An exchange scheme given in the extra keywords is respected.) The Initial guess control sets ORCA’s SCF starting guess (Guess in the %scf block):

  • default – leave ORCA’s own default.

  • Hueckel, HCore, PAtom, PModel, SAD, SADNO – ORCA’s built-in guess types. PModel is often much more stable than SAD for a single atom or ion.

  • Previous wavefunction – seed from the orbitals (orca.gbw) left by the nearest earlier ORCA step in this flowchart – for example, run a cheap, robust functional (PBE) immediately before a fragile double-hybrid (REVDSD-PBEP86-D4/2021) on the same atom, and turn this on for the double-hybrid step. This only reaches a different, preceding node in a roughly linear flowchart – it cannot seed across iterations of a Loop, since a Loop gives each iteration its own directory (see below).

  • Specified orbitals – seed from a named checkpoint file instead, named by the Specified orbitals control revealed underneath. This is the way to seed across loop iterations: for example, a basis-set escalation inside a Loop over basis sets, where each iteration reads the checkpoint the previous iteration (a smaller basis) wrote.

Either wavefunction choice reveals If wavefunction not found, controlling what happens when there is nothing to read (e.g. the very first ORCA step, or the first pass through a basis-set-escalation loop): Throw an error (the default – an explicit wavefunction request that cannot be honored is treated as a configuration problem) or one of the Use ... guess choices, which falls back to that guess type instead. Set this to a fallback guess to make “Specified orbitals” safe to leave on for every iteration of a loop.

Note

ORCA projects across basis sets. MORead is not limited to restarting with the same basis set – when the checkpoint’s basis differs from the current job’s, ORCA projects the old orbitals onto the new basis (the same trick ORCA’s own Extrapolate (CBS) keyword uses internally). So this also works for a basis-set escalation, not just same-basis restarts across different functionals. It should still be the same atom/system in the same electronic state (charge and multiplicity) – MORead supplies starting orbitals, it does not reinterpret the electronic state.

Saving a checkpoint for a later step to read is a separate pair of controls:

  • Save orbital checkpoint – yes copies this run’s converged orbitals to the file named by Checkpoint name, revealed underneath, once the run finishes successfully.

  • Checkpoint name and Specified orbitals both resolve the same way: default (the default value of each) derives a label automatically from the system’s composition, charge, and multiplicity (e.g. Co_q0_m4) – the usual choice, since it naturally gives one checkpoint per atom/electronic state, shared across an inner loop (e.g. over basis sets) for that atom, and reset automatically when an outer loop moves to a new atom. Any other bare name is stored under a checkpoints folder at the top of the job; an absolute path is used as-is, e.g. to keep a checkpoint outside this job and reuse it across separate flowchart runs.

Note

Reading a checkpoint from another job. Only Specified orbitals can reference another job – job://<job number>/<name> – since a job must never write into another job (so this is not available for Checkpoint name). Use job://<job number>/default when you want that other job’s own auto-derived name for this same system but do not know what it resolved to – for example, job://53/default picks up whatever Co_q0_m4-style label job 53 auto-derived for this atom. This needs SEAMM’s managed Jobs/<project>/Job_NNNNNN directory layout (i.e. jobs run through the dashboard/JobServer, not a bare local run) to locate the other job.

For a basis-set-escalation loop (e.g. looping def2-SVP -> def2-TZVP -> def2-QZVP for one atom), turn on both Save orbital checkpoint and Initial guess = Specified orbitals (leaving Checkpoint name / Specified orbitals on default so they agree automatically), and set If wavefunction not found to a fallback guess so the first iteration – which has nothing to read yet – does not error. Loop from the smallest basis to the largest: projecting a converged small-basis density up into a larger virtual space is cheap and well-conditioned, which is the direction ORCA’s own CBS extrapolation uses internally.

Extra keywords and extra ORCA blocks#

Two free-text escape hatches cover anything without a dedicated control:

  • Extra keywords – additional ORCA ! keywords, appended after everything the other controls generate (e.g. RIJCOSX, NoFrozenCore, SlowConv). A keyword that repeats one the controls already produce (say TightSCF with the SCF convergence at TIGHTSCF) is dropped, since ORCA refuses repeats. An SCF-convergence or integration-grid preset given here overrides the control’s, with a note in the output.

  • Extra ORCA blocks – literal ORCA input (one or more % blocks), inserted verbatim right before the geometry, after the blocks the controls above generate. Use this for one-off SCF-stabilization tricks with no dedicated control, for example on a difficult atom:

    %scf
      MaxIter 400
      Shift Shift 0.3 ErrStart 0.05 end
      DIISBfac 1.1
    end
    

Energies, gradients, and forces#

The Results tab lists everything the step can produce. Tick a result to save it to a variable, a table, or the structure’s property database (see below).

Requesting gradients makes the step compute the nuclear gradient (i.e. the forces). ORCA computes an analytic gradient (EnGrad) when one exists for the chosen method, and automatically falls back to a numerical gradient (EnGrad NumGrad) when it does not — for example DLPNO-CCSD(T) (no analytic (T) gradient), the non-self-consistent wB97M(2) / wB97X-2 double hybrids, and PWPB95, DSD-PBEB95, KPR2SCAN50 and WPR2SCAN50, whose analytic gradients ORCA 6.1 lacks or cannot compute. A note is printed when the (more expensive) numerical gradient is used. Most functionals — including the standard double hybrids — have analytic gradients, so a single-point DFT run yields the energy and forces cheaply, which is convenient for generating machine-learned force-field training data.

Geometry optimization#

The Optimization sub-step adds ORCA’s Opt keyword to the energy run and, when it finishes, stores the optimized geometry. A Convergence control selects ORCA’s optimization preset (LooseOpt … VeryTightOpt).

A Structure handling control chooses where the optimized geometry goes. It defaults to Overwrite the current configuration, so the optimized structure flows straight into any following sub-step — for example a Frequencies calculation runs at the optimized geometry without any extra wiring. You can instead create a new configuration (or a new system and configuration) to keep the starting structure alongside the optimized one, or discard the optimized structure entirely.

Frequencies (Hessian and thermochemistry)#

The Frequencies sub-step computes the Hessian and the harmonic vibrational frequencies, the IR intensities, and the thermochemistry (zero-point energy, thermal enthalpy, entropy, and Gibbs free energy). The level-of-theory controls are the same as the Energy step, plus two extra controls:

  • Second derivatives — analytic (ORCA’s AnFreq; much faster, needs an analytic second derivative for the method — HF, most DFT functionals, MP2) or numerical (NumFreq; finite-differences the gradient, works for any method with a gradient but is considerably more expensive).

  • Temperature — the temperature for the thermochemistry (the pressure is ORCA’s default of 1 atm).

  • Structure handling — where to store the structure and its properties. A frequency calculation does not change the geometry, so this defaults to Overwrite the current configuration; you can instead create a new configuration (or a new system and configuration) to hold the results, or discard the structure.

Tick the results you want on the Results tab: the frequencies, the IR intensities, the zero-point energy, the enthalpy, the Gibbs free energy, the number of imaginary frequencies, and the largest zero-mode frequency. Imaginary (negative) frequencies are reported and flagged — a minimum has none, a transition state has one.

The output lists the vibrational frequencies and their IR intensities in a table. The frequencies (and IR intensities) are also written to frequencies.csv in the step’s directory for easy access, and an IR_spectrum.graph file is written with the IR spectrum as a stick trace plus a Lorentzian-broadened trace (FWHM ≈ 15 cm⁻¹) that mimics an experimental spectrum. Open the .graph file in the SEAMM dashboard to view it interactively.

Note

Units. Following SEAMM’s SI-based convention, the zero-point energy, enthalpy, and Gibbs free energy are reported in kJ/mol. (Orbital energies — HOMO/LUMO — are reported in eV throughout ORCA, the conventional unit for orbital energies.)

The report also lists the largest of the 5 or 6 nominally-zero translation/rotation frequencies. These should be zero; how far the largest one departs from zero is a convenient gauge of the numerical accuracy of the Hessian. ORCA projects the translations and rotations out of its printed frequencies (it reports them as exactly 0.00), so this residual is instead obtained by diagonalizing the raw, un-projected mass-weighted Cartesian Hessian from orca.hess — the value you see is therefore the true numerical residual, not zero.

Note

The same analytic second derivative also backs the ORCA MDI engine: it answers a custom <HESSIAN command (AnFreq) so a driver — e.g. the Normal Mode Sampling step — can pull the analytic Hessian over a warm MDI connection. The engine advertises <HESSIAN only when ORCA has an analytic Hessian for the method (HF, MP2, ordinary DFT) and not for double hybrids or (DLPNO-)CCSD(T), so the driver’s capability check is truthful: it uses the analytic Hessian when offered and finite-differences the forces otherwise.

Counterpoise (BSSE) corrections#

The BSSE sub-step computes the counterpoise-corrected (Boys–Bernardi) energy and gradient of an N-fragment complex — removing the basis-set superposition error that artificially over-stabilizes the interaction. It is aimed at machine-learned-force-field (MLFF) training data that is BSSE-free on both the energy surface and the forces. For a complex of N fragments it runs 2N + 1 ordinary ORCA jobs (the full cluster; each fragment real in the full-cluster basis with the rest ghosted; each fragment alone in its own basis) and assembles the counterpoise correction from their results (the shared, engine-agnostic seamm_bsse library does the bookkeeping, so the same correction algebra also backs Psi4’s BSSE step).

Use it with the conventional finite-basis methods (HF, DFT, MP2, canonical CCSD(T)), where BSSE is a real effect. For explicitly-correlated F12 methods the superposition error is already negligible, so the counterpoise correction is essentially zero and this step is unnecessary — take the interaction energy directly from a plain Energy run instead (see Explicitly-correlated (F12) methods above).

The level-of-theory controls are the same as the Energy step. The counterpoise-specific controls are described below.

Fragments#

Fragments chooses how to split the complex:

  • auto (molecules) (the default) — every separate molecule found in the structure becomes its own fragment (an error if there are fewer than two — e.g. a Na+/Cl- pair is two molecules). Works directly on output from the Dimer Builder step, or any structure with two or more distinct molecules.

  • specified — the fragments come from Fragment atoms, shown only in this mode: one semicolon-separated group per fragment, each a comma/space list and/or ranges of 1-based atom numbers (as shown in the structure), e.g. 1-3; 4-6 for two fragments or 1-3; 4-6; 7 for three. The last group may be rest, meaning all the atoms not in an earlier group, e.g. 1-3; rest.

Any number of fragments ≥ 2 is supported. Flowcharts saved before these settings were generalized are read as before: their fragment A atoms becomes fragment atoms "<A>; rest" and auto (2 molecules) becomes auto (molecules).

Fragment charges#

Important

Set this explicitly for any ionic system, or check the printed log to confirm the automatic default did the right thing. A charge given to the wrong fragment is not always caught by validation (see below) — it can silently compute the wrong physics, or surface only as a confusing ORCA SCF failure deep inside one sub-job.

Fragment charges gives each fragment’s formal charge, in the same order as the fragments themselves (the order auto finds molecules in, or the semicolon-separated groups in Fragment atoms) — a comma/space separated list of integers, e.g. 1, -1 for a Na+/Cl- pair.

Leave it empty for the common case:

  • If every fragment is neutral (the usual case for an H-bonded complex), that is exactly what an empty field gives you.

  • If the input structure itself carries a per-atom formal charge — for example an ion in an SDF/MOL file, which records the charge on the specific atom it belongs to (an M  CHG record) — each fragment’s default charge is instead the sum of its own atoms’ formal charge. A well-formed Na+..H2O or Na+..Cl- structure then gets the correct per-fragment charges (0/+1, or +1/-1) with no Fragment charges entry needed at all — this covers monatomic ions (Na+, Cl-) and polyatomic ions (NH4+, BF4-) alike.

An explicit value always overrides charges taken from the structure.

Caution

The fragment order is not necessarily the order you’re thinking of the ions in, and it is easy to get backwards. The step always prints exactly what it used, before submitting anything to ORCA:

Fragment 1 has 3 atom(s) (O, H, H), charge +0.
Fragment 2 has 1 atom(s) (Na), charge +1.

Check this against which fragment you actually meant to charge — especially when typing Fragment charges by hand. Swapping a charge between two same-parity fragments (e.g. two different monatomic ions) still sums to the right overall complex charge, so it is not always caught by the checks below; it just computes the wrong physics.

Two checks run before anything is submitted to ORCA, stopping the run with a clear message if they fail:

  • the fragment charges must sum to the complex’s own overall charge;

  • each fragment, at its assigned charge, must have an even number of electrons. Every fragment is closed-shell (multiplicity 1), and an odd-electron fragment cannot be — this is what catches a charge given to the wrong fragment when it does not happen to share parity with the fragment it should have gone to (e.g. a neutral water fragment mistakenly given an ion’s +1 charge).

Other controls#

  • Compute the gradient — yes (the default) computes the counterpoise-corrected gradient (forces) as well as the energy, which is what MLFF training needs. no (energy only) is cheaper and, importantly, allows methods that have no analytic gradient in ORCA — notably CCSD(T) and DLPNO-CCSD(T) — for gold-standard counterpoise interaction energies.

  • Optimize free monomers — whether to relax each isolated fragment before taking the correction. Leave it no for a fixed-geometry PES / MLFF target.

  • Write the wavefunction (wfx) file — default no. When yes, the full cluster’s density is kept and converted (via orca_2aim) to an orca.wfx, so a following Atomic Charges step can partition it into DDEC6 charges on the CP complex — exactly as after an Energy step.

The step reports, and offers on the Results tab, the BSSE-corrected energy (the energy result), the uncorrected (raw) complex energy, the BSSE correction (corrected minus uncorrected, also shown in kcal/mol), and — when requested — the corrected gradient. Tick any of them on the Results tab to save it to a variable, table, or the property database.

A ghost-centre gradient-noise warning#

Occasionally the step’s output includes a warning like this:

WARNING: the BSSE-corrected gradient failed the translational-invariance
guard (net force 0.001290 E_h/bohr > tolerance 0.000300 E_h/bohr), most
likely a ghost-centre integration blowup. Falling back to the
uncorrected cluster gradient for the forces; the energy is still fully
counterpoise-corrected. The BSSE gradient correction being dropped had
magnitude 0.001242 E_h/bohr -- small relative to typical forces means
the fallback is almost certainly fine; if it is not small, treat this
point as suspect (exclude/rerun) rather than trusting either gradient.

What causes it. One of the fragment-in-cluster sub-jobs — a fragment’s real atoms plus the other fragment(s) as ghost (basis-function only, no electron) centres — can occasionally pick up a small amount of numerical-integration noise in the Pulay (basis-derivative) force on a ghost centre, from ORCA’s RIJCOSX/COSX approximation. This is a genuine, if uncommon, ORCA numerical artifact, not a SEAMM bug and (usually) not a sign that anything is chemically wrong with the point. A correctly-assembled counterpoise gradient always sums to exactly zero net force over the whole cluster (translational invariance — nothing external is acting on an isolated system); this noise breaks that, which is what makes it detectable at all — the energy stays fine, with no warning from ORCA itself.

What SEAMM does about it. Every BSSE gradient is checked against this zero-net-force requirement. When the residual exceeds a small tolerance, the corrected gradient (only) falls back to the raw, uncorrected cluster gradient; the energy is unaffected either way, since the noise only shows up in the numerically far more sensitive gradient.

How to judge whether the fallback is safe for a given point. The warning reports both the net-force residual that tripped the guard and the size of the BSSE gradient correction that got dropped. Compare them:

  • If the two are close in magnitude (as in the example above — 0.00129 vs. 0.00124 E_h/bohr), the “correction” being discarded was itself mostly noise, not real physics — the fallback is safe, and the point can be used as-is.

  • If the correction magnitude is much larger than the net-force residual, real physics is being discarded, not just noise — that point is genuinely suspect and worth excluding or rerunning (e.g. with a finer Integration grid, DEFGRID3 — see Integration grid above) rather than trusting either gradient at face value.

In practice this fires rarely, and mostly at short-to-moderate fragment separation (where the counterpoise correction is largest, so there is more for the noise to hide inside); a finer integration grid on the affected structures is usually enough to avoid it entirely.

Note

Current limitations. Every fragment must be closed-shell (multiplicity 1); open-shell fragments are not yet supported. Only an ORCA-internal basis set works (not yet the Basis Set Exchange). Computing the gradient additionally requires an analytic gradient, so the numerical-gradient methods ((DLPNO-)CCSD(T)) are available in energy-only mode only. The SThresh control and the extra property analyses (bond orders, Hirshfeld charges, polarizability) available on the plain Energy step are not available here.

Saving results to the database#

Both scalar results (the energies, HOMO/LUMO energies and gap, the dipole magnitude, <S^2>, the polarizability) and array results (the gradient, the dipole-moment vector, the Mulliken/Löwdin/Hirshfeld charges, the Mayer valences, and the rotational constants) can be stored as properties on the configuration. In the Results tab, tick the database column for the result; it is stored under a name like gradients#ORCA#<model>, where <model> is the level of theory.

Other properties and outputs#

The Energy step can also compute Mayer bond orders and Hirshfeld charges (and optionally apply them to the structure), the dipole polarizability, and an analytic wavefunction (.wfx, via orca_2aim) for a following Atomic Charges step to partition into DDEC6 charges. Every run cites ORCA, the DFT functional, the basis set, and the supporting integral / exchange-correlation libraries.

Running ORCA in parallel#

By default ORCA uses all the cores the machine or batch job provides. Settings come from two files with different jobs:

  • How to run ORCA lives in ~/SEAMM/orca.ini – the full path to the executable and the OpenMPI library-path for parallel runs:

    [local]
    installation = local
    code = /path/to/orca
    library-path = /path/to/orca-mpi/lib
    
  • User run options live in the [orca-step] section of the main SEAMM configuration (~/.seamm.d/seamm.ini), and can also be given on the command line:

    [orca-step]
    ncores = available        # or an integer, or 1 to force serial
    memory = available        # or 'all', or e.g. '3 GB' (per process)
    
  • ncores — how many processes ORCA may use (its %pal). available (the default) uses all cores the machine/job provides; give an integer to cap it, or 1 to force serial.

  • memory — the per-process memory for ORCA’s %maxcore. available (the default) scales to the memory per core; all divides the whole node among the processes; or give an explicit amount such as 3 GB.

  • library-path (in orca.ini) — the lib directory of the OpenMPI that matches the version ORCA was built against. ORCA 6.1 requires OpenMPI 4.1.x and does not support 5.x; mixing versions makes parallel runs abort with a BLAS-ERROR. A dedicated conda env is the easy way to get the right one:

    conda create -n orca-mpi -c conda-forge "openmpi=4.1"
    

    then set library-path to that env’s lib directory. The step automatically puts the matching mpirun (the sibling bin directory) on PATH so ORCA launches its workers with the correct OpenMPI.

Note

macOS. ORCA does not pass the DYLD_* loader variables to the MPI processes it spawns, so library-path alone does not let the dynamic loader find libmpi on a Mac. Put the OpenMPI libraries on the default search path once, for example by symlinking into /usr/local/lib:

ln -s /path/to/orca-mpi/lib/libmpi.40.dylib /usr/local/lib/

On Linux the exported LD_LIBRARY_PATH is inherited normally, so library-path is sufficient there.

If you do not have a matching OpenMPI, set ncores = 1 to run serially.

Environment-modules clusters. If ORCA is provided via a cluster’s module system rather than a fixed path you manage yourself, use installation = modules instead of local, naming the module(s) to load:

[local]
installation = modules
modules = ORCA
code = /path/to/orca
library-path = /path/to/orca-mpi/lib

seamm_exec runs module load <modules> before invoking ORCA. This works for both serial and parallel runs – as of this release, a parallel run’s OpenMPI library-path is prepended onto whatever the module load already put on LD_LIBRARY_PATH/DYLD_LIBRARY_PATH, rather than overwriting it (a real, fixed bug: earlier releases could silently discard the module-provided ORCA libraries on a parallel run, failing with error while loading shared libraries even though the module load itself had succeeded).

Rerunning a job#

ORCA runs through SEAMM’s task layer, which keeps a record of each ORCA calculation in tasks/manifest.json in the step’s directory, with a tasks/<name>/DONE file for each one that finished. When a job is run again in the same directory – after it was stopped, lost its node, or ran out of time – a calculation that had finished with the same input is not repeated: its results are read from the files it left, and the output says:

ORCA had already finished this calculation, with the same input, in <directory>;
using those results.

A calculation is run again when

  • its input changed (the geometry, method, basis, keywords or blocks; the %pal and %maxcore lines are ignored, so a different number of cores or amount of memory does not count), or

  • it did not finish: it failed, or the job was stopped while it was running. ORCA reports success to the operating system even after an “error termination”, so a run counts as finished only if orca.out ends with ORCA TERMINATED NORMALLY. A calculation is tried at most three times in all; after that the step stops with a message saying so. Changing its input, or deleting its entry in tasks/manifest.json, lets it run again.

If the job was killed outright, an ORCA process from the earlier run may still be running; the rerun stops it before starting the calculation again.

Where ORCA runs#

An ORCA calculation names only the program; how to run ORCA is worked out on the machine that runs it, from that machine’s orca.ini ([local]: the path to ORCA, or installation = modules with the modules to load, and library-path for the OpenMPI ORCA was built with), falling back to orca on the PATH. So a job whose target sends its calculations to a cluster runs ORCA with the cluster’s own ORCA, even if ORCA is installed differently (or not at all) where the job runs. Calculations estimated to take under a minute stay on the job’s own machine when ORCA is installed there.

On a site that provides ORCA through environment modules, give code as ORCA’s full path in orca.ini (with installation = modules): ORCA must be invoked by its full path to find its sub-programs for parallel runs, and a bare orca cannot be expanded before the module is loaded.

The sub-jobs of a counterpoise (BSSE) correction run together: concurrently on this machine when there are cores for more than one, or as separate calculations on the job’s cluster queue.

ORCA as a model chemistry#

Steps that set up a model chemistry and then evaluate it at many structures – the Energy step, the Dimer Builder’s energy-based contact search, Normal Mode Sampling and LAMMPS QM-MD – use ORCA without any settings in the ORCA step itself: put a Model Chemistry step in the flowchart and choose an ORCA model chemistry there (e.g. ORCA:DFT@B3LYP/def2-SVP).

How ORCA then runs is chosen for you, and gives the same numbers either way:

  • As separate calculations, one per structure, for the Energy step, the Dimer Builder and Normal Mode Sampling’s finite-difference Hessian. They run several at a time on this machine, or on the cluster when the job’s target sends its calculations to a queue, and a rerun of the job reuses those that had finished. ORCA starts afresh for each structure in any case, so running them side by side is usually faster than one after another in an engine. Periodic structures are not run this way.

  • As a persistent MDI engine for steps that drive ORCA one geometry after another (LAMMPS QM-MD, Normal Mode Sampling’s analytic Hessian on this machine).

The model chemistry’s basis is used either way: any basis ORCA knows, or a basis from the Basis Set Exchange (bse:NAME), whose definition is fetched for the elements present and passed to ORCA exactly as the ORCA step does. A lone atom or ion runs without the COSX approximation (NoCOSX), as in the ORCA step, which avoids an artifact of several kJ/mol from COSX’s grid on a single centre.

Every calculation run this way uses tight SCF convergence and the fine integration grid (TIGHTSCF DEFGRID3). These calculations are mostly forces for training data or for derivatives, which need both, rather than ORCA’s looser defaults.

Functionals without a dispersion correction of their own are also offered with Grimme’s D4 correction, as <functional>-D4 – e.g. ORCA:DFT@R2SCAN-D4/def2-TZVP, which runs ! R2SCAN D4. They are offered for the functionals the D4 model has parameters for; double hybrids are not, and functionals that already include dispersion (WB97X-D4, B97M-V, R2SCAN-3C, …) are offered only as themselves. Method names are matched as written, so use the capitals shown in the Model Chemistry step’s list.

Driving ORCA as an MDI engine#

Because ORCA has no in-process interface, the engine runs the orca binary once per geometry in a persistent working directory, reusing the previous geometry’s orbitals (orca.gbw) as the SCF guess – the main saving for a series of nearby structures. Two consequences:

  • Only methods with an analytic gradient are offered via MDI. The engine always computes the energy and forces together (EnGrad), so DLPNO-CCSD(T), CCSD(T) and the non-self-consistent wB97M(2) / wB97X-2 are not MDI-capable; choosing one gives a clear error. Everything else (HF, MP2, and the analytic-gradient functionals) works.

  • ORCA is not start-up-dominated, so the per-geometry cost is real – the MDI benefit here is orbital reuse and a uniform interface, not the large speed-up that cheap engines see. Use an inexpensive functional for jobs that only need the energy surface to guide them (such as contact finding); reserve expensive, high-accuracy calculations for ordinary single-point steps.

Settings that depend on each other#

Which settings apply depends on others: the functional and its type only for DFT, basis- set extrapolation not for F12 methods or for steps that need a gradient or Hessian, and only F12 bases for an F12 method. The step’s dialogs show only the settings that apply with the current choices, and the same rules are used when a flowchart is built or edited without the editor (seamm-flowchart or SEAMM’s MCP server): a setting that would have no effect is refused, with the reason, and a value that contradicts another is refused too. See “Flowcharts without the editor” in SEAMM’s user guide.

Frequencies and BSSE do not use some of the Energy settings, and say so if one is set.

Indices and tables#