seamm_thermochemistry package#

Submodules#

seamm_thermochemistry.db module#

SQLite-backed store for atomic reference energies.

Every SEAMM code that reports a formation-referenced energy (Gaussian, Psi4, ORCA, VASP, …) needs the same two pieces of data per element: an experimental reference (heat of formation, entropy, standard-state description) and one or more computed atomic reference energies (one per code/method/settings combination). Today each plugin carries its own copy of both as a wide, mostly-empty CSV. This module replaces that with a single, shared, relational store – one element row per element, one atom_energy row per (element, code, method, ref_type, settings) combination – with room for the provenance a shared reference dataset needs (who computed it, when, with what).

Two reference conventions are both first-class (ref_type):

  • "atom": energy of the isolated, gas-phase atom. This is the Gaussian/Psi4/ORCA convention, and the one to use for a physically anchored, cross-code-comparable energy of formation.

  • "element_phase": energy per atom of the element’s standard-state phase (bulk metal, graphite, O2(g), …). This is VASP’s existing convention (element_energies.csv) – no experimental anchor needed, but the result is referenced to the computed elemental phases, not real formation energies. It also serves as a fallback for elements (e.g. Mn) where the free atom is a poor DFT reference.

See formation.py for the arithmetic that consumes this store.

class seamm_thermochemistry.db.ThermoDB(path=None, read_only=False)[source]#

Bases: object

Helper around the shared atomic reference-energy SQLite database.

Parameters:
  • path (str or Path, optional) – Database file. Defaults to DEFAULT_DB_PATH: the installer- managed location registered in ~/.seamm.d/seamm.ini’s [thermochemistry] section (database-path, normally ~/SEAMM/Parameters/thermochemistry/thermochemistry.db – see installer.py) if the installer has been run, else the bundled data/thermochemistry.db shipped with the package (a stale prototype snapshot, useful only before the installer has run).

  • read_only (bool) – Open without creating/migrating the schema, and without permission to write. Use for a package-installed, already-built database.

add_atom_energy(symbol, code, method, energy, *, ref_type='atom', settings='', units='kJ/mol', correction=None, spin_multiplicity=None, source=None, computed_date=None, code_version=None, notes=None)[source]#

Insert or update one computed atomic reference energy.

Parameters:
  • symbol (str) – Element symbol, e.g. “O”. Must already exist via add_element.

  • code (str) – Originating code: “gaussian”, “psi4”, “orca”, “vasp”, …

  • method (str) – Method/functional label, e.g. “CBS-QB3”, “PBE-D3BJ”. Where the basis is folded into the label (as in the legacy Gaussian/Psi4 tables, e.g. “CCD-FC/6-31++G(2d,2p)”), pass the full label.

  • energy (float) – The atom’s energy, in units (converted to kJ/mol on storage).

  • ref_type ({"atom", "element_phase"}) – “atom” = isolated gas-phase atom. “element_phase” = energy per atom of the standard-state phase (VASP’s existing convention).

  • settings (str) – Free-form disambiguator that is part of the identity key, e.g. “encut=700eV”. Defaults to “” (not NULL, so the UNIQUE constraint dedups consistently).

  • units (str) – Units of energy and correction: one of “kJ/mol”, “kcal/mol”, “eV”, “hartree”/”E_h”.

  • correction (float, optional) – Additive correction (e.g. Gaussian’s “<method> correction” column), same units as energy.

add_element(atomic_number, symbol, *, multiplicity=None, term_symbol=None, standard_state=None, dfH0_0K=None, dfH0_298K=None, dfH0_298K_stderr=None, h298_minus_h0_atom=None, h298_minus_h0_std_state=None, s298_gas=None, s298_gas_stderr=None, s298_std_state=None, s298_std_state_reference=None, reference=None, reference_note=None)[source]#

Insert or update the experimental reference row for one element.

batch()[source]#

Defer commits until the end of the block, for bulk imports.

add_element/add_atom_energy normally commit immediately – the right default for interactive use, but one fsync per row makes a bulk import of hundreds of thousands of rows (e.g. the full Gaussian/Psi4 composite-method grids) prohibitively slow. Wrap those calls in with db.batch(): to commit once at the end instead.

close()[source]#
coverage(code, method, *, ref_type='atom', settings='')[source]#

Return (n_elements_present, max_atomic_number) for one (code, method).

dfH0(symbol, *, at_0K=True)[source]#

Experimental atomic heat of formation, kJ/mol.

Parameters:

at_0K (bool) – If True (default, and the recommended anchor for an “energy of formation” – see the design doc), return DfH(X, 0K). If False, return the 298 K gas-phase value used by the existing enthalpy-of-formation code paths.

doi()[source]#

The Zenodo DOI of the published snapshot this database file IS, or None for a working copy that has not been published yet.

Unlike the schema version (which just says “the table layout this code understands”), the DOI identifies the actual data – the citable, permanent reference for reproducibility. Set once, right before the file is uploaded to Zenodo – see scripts/publish_to_zenodo.py and set_doi.

dump_atom_energies_csv(path)[source]#

Write the atom_energy table (joined to element symbols) to CSV.

dump_elements_csv(path)[source]#

Write the element table to CSV, ordered by atomic number.

elements()[source]#

Return all element rows, ordered by atomic number.

get_atom_energy(symbol, code, method, *, ref_type='atom', settings='', units='kJ/mol')[source]#

Return one atomic reference energy (including its correction), or None.

get_atom_energy_row(symbol, code, method, *, ref_type='atom', settings='')[source]#

Return the full atom_energy row (energy, correction, source, computed_date, code_version, notes, …) as a dict, or None.

For provenance/reporting – get_atom_energy returns just the combined number get_reference_energies/atomization_energy need.

get_element(symbol=None, atomic_number=None)[source]#

Return the experimental reference row for one element as a dict, or None.

get_reference_energies(code, method, *, ref_type='atom', settings='', units='kJ/mol')[source]#

Return {symbol: energy} for every element tabulated for this (code, method, ref_type, settings) combination.

This is the lookup table formation.formation_energy needs.

list_methods(code=None)[source]#

Return distinct (code, method, ref_type, settings) tuples present.

missing(code, method, symbols, *, ref_type='atom', settings='')[source]#

Return the subset of symbols with no tabulated (code, method) energy.

s298_gas(symbol)[source]#

Experimental standard molar entropy of the isolated gas atom, S°(X, g, 298.15 K), J/(mol K), or None.

s298_std_state(symbol)[source]#

Experimental standard molar entropy of the element’s standard-state phase, PER ATOM of X (e.g. S°(H2,g,298.15K)/2 for H), J/(mol K), or None. Needed (together with s298_gas) for a Gibbs-energy-of- formation anchor – see formation.formation_gibbs_energy.

schema_version()[source]#

The schema version stamped in this database (an int), for reports/provenance – see _SCHEMA_VERSION.

set_doi(doi)[source]#

Stamp this database file with the Zenodo DOI of the snapshot it is about to become. Call this right before uploading the file – Zenodo prereserves a DOI as soon as a deposit/version is created, before any file is attached, so the DOI can be baked into the file itself rather than only living in the Zenodo record around it.

set_s298_std_state(symbol, value, *, reference=None)[source]#

Set just the standard-state-phase entropy (and its citation) for an already-add_element-ed element, without needing to re-supply every other column (unlike add_element, which replaces the whole row).

seamm_thermochemistry.derive module#

Derive DfH(X, 0K) from the 298 K experimental data already in the store.

The master reference sheet only has ΔfH°(0K) filled in for H and He; every other element has ΔfH°(298K) plus the two H(298K)-H(0K) corrections needed to convert it, via the standard thermochemical identity:

DfH(X, 0K) = DfH(X, 298K) - [H298-H0]_atom + [H298-H0]_standard_state

h298_minus_h0_std_state is tabulated per formula unit of the standard state, not per atom – for a diatomic standard state (H2, N2, O2, F2, Cl2, Br2) it must be divided by 2 to get a per-atom contribution; for the monatomic/solid standard states used by every other element in range, no division is needed. Verified against the two elements the master sheet already gives directly: H (218 - 6.197 + 8.468/2 = 216.03, matches 216.034) and He (0 - 6.197 + 6.197 = 0, matches 0), both to within the sheet’s own rounding.

seamm_thermochemistry.derive.derive_dfH0_0K(db, *, overwrite=False)[source]#

Fill in element.dfH0_0K for every element where it’s derivable.

Parameters:
  • db (ThermoDB)

  • overwrite (bool) – If False (default), skip elements that already have dfH0_0K – the master sheet’s own H/He values are treated as authoritative over a re-derivation (though the derivation reproduces them to within rounding; see module docstring).

Returns:

Symbols of the elements filled in.

Return type:

list of str

seamm_thermochemistry.formation module#

The shared arithmetic: atomization and formation energies from a ThermoDB.

Every code (Gaussian, Psi4, ORCA, VASP, …) that reports a formation energy is doing the same subtraction:

atomization = sum(n_X * E_ref(X)) - E(system)

and, when an experimental anchor is available,

formation = sum(n_X * DfH(X)) - atomization

They differ only in the choice of E_ref (an isolated gas atom, or an atom of the element’s standard-state phase – ref_type on ThermoDB) and whether an experimental anchor is added at all. See ~/Sites/reference-energy/2026-07-24_reference-energy/ for the full design rationale, including why the arbitrary code-dependent zero of a raw total energy is the problem this solves, and the “no thermochemistry -> energy of formation, not enthalpy” case this module is aimed at.

Sign convention, checked against the existing gaussian_step / psi4_step / vasp_step implementations:

  • atomization_energy matches Gaussian/Psi4’s “E atomization” (>0 for a bound system: separated atoms sit higher in energy than the molecule).

  • formation_energy(..., anchor=True) matches the existing DfH_at - H_atomization construction (Gaussian/Psi4’s enthalpy of formation), generalized to also give a ZPE-free “energy of formation” when the 0 K anchor is used and system_energy excludes ZPE.

  • formation_energy(..., anchor=False) matches VASP’s existing DfE0 = E(system) - sum(n_X * mu_X) (materials formation-energy convention, no experimental data required) – note this is -atomization_energy, not atomization_energy itself; the two conventions report the quantity with opposite sign, which is why they are two functions rather than one flag.

exception seamm_thermochemistry.formation.MissingReferenceData[source]#

Bases: KeyError

Raised when a composition needs reference data the ThermoDB doesn’t have.

seamm_thermochemistry.formation.atomization_energy(composition, system_energy, db, code, method, *, ref_type='atom', settings='', units='kJ/mol')[source]#

sum(n_X * E_ref(X)) - E(system): the energy to separate all atoms.

This is exactly Gaussian/Psi4’s “Atomization Energy” (electronic-only, no ZPE, if system_energy is itself ZPE-free) and VASP’s “cohesive energy”. Positive for a bound system.

Parameters:
  • composition (dict or collections.Counter) – {element_symbol: count}, e.g. {“C”: 2, “H”: 6}.

  • system_energy (float) – The system’s electronic energy, in units.

  • db (ThermoDB) – Open reference-energy database.

  • code (str) – Identify which tabulated atom energies to use.

  • method (str) – Identify which tabulated atom energies to use.

  • ref_type ({"atom", "element_phase"}) – See ThermoDB.add_atom_energy.

  • settings (str) – Must match the settings used when the atom energies were stored.

  • units (str) – A pint-parseable unit string for system_energy and the return value, e.g. “kJ/mol”, “eV”, “E_h”.

seamm_thermochemistry.formation.formation_energy(composition, system_energy, db, code, method, *, ref_type='atom', settings='', anchor=True, anchor_at_0K=True, units='kJ/mol')[source]#

The formation-referenced energy of a system.

Parameters:
  • composition – As in atomization_energy.

  • system_energy – As in atomization_energy.

  • db – As in atomization_energy.

  • code – As in atomization_energy.

  • method – As in atomization_energy.

  • ref_type – As in atomization_energy.

  • settings – As in atomization_energy.

  • units – As in atomization_energy.

  • anchor (bool) – If True (default), add the experimental atomic heat of formation so the result is a true energy/enthalpy of formation – the physically meaningful, cross-code-comparable quantity this package exists for. If False, return the bare, no-experimental-data-needed quantity: VASP’s existing DfE0 convention (ref_type="element_phase" is the natural pairing, but any ref_type is accepted).

  • anchor_at_0K (bool) – Use the 0 K experimental anchor (recommended: pairs with a ZPE-free system_energy to give an “energy of formation”, degrades gracefully to DfH(0K) if ZPE is added back in, and to DfH(298K) if the molecule’s thermal correction is added) or the 298 K anchor (matches the existing enthalpy-of-formation code paths directly). Ignored if anchor is False.

Returns:

Energy/enthalpy of formation (anchor=True) or the bare atomization-referenced energy in VASP’s DfE0 sign convention (anchor=False), in units.

Return type:

float

seamm_thermochemistry.formation.formation_enthalpy(composition, H_system, temperature, db, code, method, *, ref_type='atom', settings='', units='kJ/mol')[source]#

DfH(T): enthalpy of formation from the elements, at temperature T.

Generalizes formation_energy(..., anchor=True, anchor_at_0K=True) from the electronic-only, 0 K quantity to a real temperature, by treating each reference atom as a monatomic ideal gas: its H(T)-H(0) = (5/2)RT exactly, so (from the elements -> atoms -> molecule Hess cycle, atom energies at T = E_ref(X) + (5/2)RT)

DfH(T) = sum(n_X * DfH0(X, 0K)) - n_atoms*(5/2)*R*T
  • [sum(n_X * E_ref(X)) - H_system]

Approximation: only the ATOM’s own H(T)-H(0) is added; the standard- state element’s own H(T)-H(0) away from 298.15 K is not separately corrected for (there is no single cheap closed form for it the way there is for a monatomic ideal gas – it depends on the standard state’s own heat capacity). This is the same approximation the legacy 298 K-only gaussian_step/psi4_step formation-enthalpy code made implicitly (it used the exact tabulated 298 K anchor rather than deriving one); the error it introduces is small near 298.15 K and grows (slowly) away from it.

Parameters:
  • composition – As in atomization_energy.

  • db – As in atomization_energy.

  • code – As in atomization_energy.

  • method – As in atomization_energy.

  • ref_type – As in atomization_energy.

  • settings – As in atomization_energy.

  • units – As in atomization_energy.

  • H_system (float) – The system’s total enthalpy H(T) – electronic + ZPE + thermal correction – on the SAME absolute energy scale as the electronic energy stored for the reference atoms (not a thermal correction alone), in units.

  • temperature (float) – Temperature, in K.

Returns:

DfH(T), in units.

Return type:

float

seamm_thermochemistry.formation.formation_gibbs_energy(composition, G_system, temperature, db, code, method, *, ref_type='atom', settings='', units='kJ/mol')[source]#

DfG(T): Gibbs energy of formation from the elements, at T.

Same Hess-cycle construction as formation_enthalpy, extended to G. Requires each element’s s298_std_state (the standard-state phase’s molar entropy, per atom – see ThermoDB.s298_std_state) in addition to the usual 0 K atom energies and heats of formation; raises MissingReferenceData for any element missing either.

Derivation note: the reference atom’s OWN entropy cancels exactly out of the final formula (it is only an intermediate state in the elements -> atoms -> molecule cycle, and appears with opposite sign in each half-reaction), so – pleasantly – no atomic entropy data is needed at all, only the standard-state element’s:

DfG(T) = sum(n_X * DfH0(X,0K))
  • T * sum(n_X * S298_std_state(X)) / 1000

  • n_atoms*(5/2)*R*T

  • [sum(n_X * E_ref(X)) - G_system]

Approximation: S298_std_state is used at its tabulated 298.15 K value regardless of T (a standard-state solid’s entropy has no cheap closed-form T-dependence the way a monatomic ideal gas’s does – this is the dominant source of error away from 298.15 K, layered on top of the same 0 K-anchor approximation formation_enthalpy makes). At T=0 this reduces exactly to formation_energy(..., anchor=True, anchor_at_0K=True) evaluated with G_system in place of the electronic energy – a useful sanity check (G_system at T=0 with no thermal terms IS the electronic energy).

Parameters:
  • composition – As in atomization_energy.

  • db – As in atomization_energy.

  • code – As in atomization_energy.

  • method – As in atomization_energy.

  • ref_type – As in atomization_energy.

  • settings – As in atomization_energy.

  • units – As in atomization_energy.

  • G_system (float) – The system’s total Gibbs free energy G(T), on the same absolute scale as the electronic energy, in units.

  • temperature (float) – Temperature, in K.

Returns:

DfG(T), in units.

Return type:

float

seamm_thermochemistry.import_orca_cli module#

Import one or more ORCA atom-energy results CSVs into thermochemistry.db.

Installed as the seamm-thermochemistry-import-orca console script.

Usage#

seamm-thermochemistry-import-orca results1.csv results2.csv … seamm-thermochemistry-import-orca –dry-run results.csv seamm-thermochemistry-import-orca –force results.csv # overwrite every conflict seamm-thermochemistry-import-orca –prefer-lower-energy results.csv # keep lower

See seamm_thermochemistry.importers.import_orca_atom_results for the CSV shape expected and the vetting/duplicate-handling rules – this module is just the CLI + pretty-printing around that function.

seamm_thermochemistry.import_orca_cli.main()[source]#
seamm_thermochemistry.import_orca_cli.report(csv_path, summary)[source]#

seamm_thermochemistry.importers module#

Importers from the legacy master files into a ThermoDB.

Requires the import extra (pandas, openpyxl) – the core db/formation modules do not need either.

Three legacy sources exist today, all sharing the same 12-column experimental-reference block (by position; header text drifts slightly between them – “ΔfH°(0)” vs “ΔfH°gas(0)” vs “DfH0(0)” for the same column):

  • Paul’s master workbook, “Atom Reference Energies and States.xlsx” – the clean experimental block only, one row per element.

  • “VASP element_energies.xlsx” (sheet “element_energies”) – the same experimental block, plus paired "<method>@<encut> atom energy" (isolated atom, ref_type=”atom”) and "<method>@<encut>" (per-atom standard-state energy, ref_type=”element_phase”) columns.

  • Each of gaussian_step / psi4_step’s data/atom_energies.csv – the same experimental block, plus one column per method (isolated atom, kJ/mol) and an optional "<method> correction" column.

A fourth, ongoing source needs no xlsx/openpyxl at all:

  • import_orca_atom_results reads a per-job ORCA atom-energy results CSV (one row per element, one E DFT@<method>/<basis> (kJ/mol) / S^2 DFT@<method>/<basis> column pair per method+basis run) and vets each cell against its expected spin before adding it – see that function’s docstring.

seamm_thermochemistry.importers.import_orca_atom_results(db, csv_path, *, code='orca', ref_type='atom', warn_threshold=0.02, reject_threshold=0.2, energy_rel_tol=1e-06, energy_abs_tol=0.01, force=False, prefer_lower_energy=False, dry_run=False)[source]#

Import ORCA atom-energy results from a per-job results CSV.

Expects the shape produced by the atom-energy scan flowchart: one row per element (Atomic Number, Element, Multiplicity), and one "E DFT@<method>/<basis> (kJ/mol)" / "S^2 DFT@<method>/<basis>" column pair per method+basis combination that has been run so far (blank cells are jobs not finished/reached yet, not errors). method and basis are stored separately – method as the ThermoDB method, basis as settings – since ORCA (unlike Gaussian’s composite-method columns) always keeps them as independent axes.

Vetting, since a bad SCF root (wrong occupation, not just unconverged) is the actual failure mode seen for hard atoms (lanthanides etc.), not just a missing value:

  • A computed S^2 is compared to the multiplicity’s expected S(S+1). Within warn_threshold (relative) it’s added silently; beyond that but within reject_threshold it’s added but flagged in “spin_warnings” for manual review; beyond reject_threshold it’s not added at all (“rejected”).

  • A missing S^2 is fine for a singlet (trivially 0, commonly not reported) but rejected for any open-shell multiplicity – no way to vet it.

  • If the CSV’s own multiplicity for an element disagrees with what’s already in the element table (when known), that’s noted in “multiplicity_mismatches” but does not block the import – it may just mean the reference table’s own value needs a look.

Duplicate handling, since re-running elements/methods is expected across “hundreds of these”: an existing (element, code, method, ref_type, settings) value that’s already in the database is left alone if the new value is close (“unchanged”); if it differs by more than energy_rel_tol/energy_abs_tol it is a conflict, resolved one of three ways (force and prefer_lower_energy are mutually exclusive – pick one):

  • Neither set (default): reported in “conflicts”, not overwritten. Safest, but needs a human to look at every conflict.

  • force=True: always overwritten with the new value, listed in “updated”. Only right when you’re confident the new run is simply better across the board – a blanket “latest wins” that can silently make things worse (see prefer_lower_energy).

  • prefer_lower_energy=True: for hard atoms (partially-filled d/f shells – lanthanides, some transition metals) two independently converged runs of the same nominal state can land on genuinely different self-consistent solutions, invisible to the S^2 gate since both have the right total spin. The tell is a conflict whose delta is nearly constant across several basis sets for the same element – not a basis-incompleteness effect, a different orbital occupation. There’s no rigorous variational principle for DFT, but “prefer whichever energy is lower” is the standard practical criterion for picking between competing self-consistent solutions of the same state. If the new value is lower it’s written and listed in “updated”; if the existing value is already lower it is kept and listed in “kept_existing” – so a blanket rerun can’t silently make a good stored value worse. Still reported either way, never silent.

Parameters:
  • db (ThermoDB)

  • csv_path (str or Path)

  • code (str) – Defaults to “orca”; parameterized in case this CSV shape is ever reused for another code.

  • ref_type (str) – Defaults to “atom” (isolated gas-phase atom energies – the only thing this scan computes).

  • warn_threshold (float) – Relative S^2 deviation thresholds (see above). Tune per how strict you want the gate; there’s no universally “right” value, the deviations here are your own to judge (e.g. an ~20%-high S^2 on some lanthanides is real but was still accepted on review in this project).

  • reject_threshold (float) – Relative S^2 deviation thresholds (see above). Tune per how strict you want the gate; there’s no universally “right” value, the deviations here are your own to judge (e.g. an ~20%-high S^2 on some lanthanides is real but was still accepted on review in this project).

  • energy_rel_tol (float) – math.isclose tolerances for treating a duplicate energy value as “the same result” rather than a conflict.

  • energy_abs_tol (float) – math.isclose tolerances for treating a duplicate energy value as “the same result” rather than a conflict.

  • force (bool) – Overwrite every conflicting existing value unconditionally.

  • prefer_lower_energy (bool) – Resolve conflicts by keeping whichever energy is lower (see above). Mutually exclusive with force.

  • dry_run (bool) – Classify and report everything, but never call db.add_atom_energy – preview a CSV before committing it.

Returns:

  • dict with keys “added”, “updated”, “unchanged”, “conflicts”,

  • ”kept_existing”, “rejected”, “spin_warnings”,

  • ”multiplicity_mismatches” – each a list of small dicts describing

  • what happened, for the caller to report.

seamm_thermochemistry.importers.import_reference_xlsx(db, path, sheet='Reference Data', reference_note='JANAF')[source]#

Load the experimental block from Paul’s master workbook.

Returns the list of element symbols imported.

seamm_thermochemistry.importers.import_vasp_workbook(db, path, sheet='element_energies', reference_note='JANAF', import_elements=True)[source]#

Load the VASP element/atom energy workbook.

Imports the shared experimental block (optional, import_elements) plus every "<method>@<encut> atom energy" column as ref_type=”atom” and every matching "<method>@<encut>" column as ref_type=”element_phase”, with settings=f"encut={encut}eV". Both column families are already in kJ/mol in this workbook (checked against its own “Testing” sheet, which carries the raw VASP eV values alongside their kJ/mol conversion) – not eV, despite the encut being in eV.

Returns a dict {“elements”: […], “atom_energy”: n, “element_phase”: n}.

seamm_thermochemistry.importers.import_wide_method_csv(db, path, code, *, methods=None, ref_type='atom', reference_note=None, import_elements=True)[source]#

Load one of the wide per-code CSVs (gaussian_step/psi4_step style).

Parameters:
  • code (str) – “gaussian” or “psi4” (or whatever this file belongs to).

  • methods (list of str, optional) – Restrict to these method columns (by exact header text). If None, import every non-empty method column found – can be thousands of columns for the full Gaussian composite-method grid; pass an explicit list for anything but a one-off full import.

  • {"elements" (Returns a dict)

seamm_thermochemistry.installer module#

Installer for seamm_thermochemistry: fetches the shared atomic reference-energy database from its Zenodo record.

This package has no external executable and no conda environment – the only “installation” step is downloading thermochemistry.db, which is not shipped in the pip package (see the design doc at ~/Sites/reference-energy/2026-07-24_reference-energy/: the database is regenerated from external master files and will keep growing). It lives outside the package instead, in a configured SEAMM-root directory – the same pattern used for DFTB+’s Slater-Koster parameter sets (~/SEAMM/Parameters/slako) – registered in seamm.ini’s [thermochemistry] section as database-path.

Because there is no executable/environment to manage, this installer does NOT call super().check() / super().install() – those are entirely about the executables/conda-environment machinery InstallerBase provides for a typical plug-in, which doesn’t apply here. check/install/ update/uninstall are implemented directly instead.

class seamm_thermochemistry.installer.Installer(logger=<Logger seamm_thermochemistry.installer (WARNING)>)[source]#

Bases: InstallerBase

Install/update/remove the seamm_thermochemistry reference database.

check()[source]#

Check that the reference database is installed, offering to install it if not.

Returns:

True if the database is present (installing it first if asked and needed), False if it’s missing and installation was declined.

Return type:

bool

database_filename = 'thermochemistry.db'#
install()[source]#

Download the reference database from Zenodo and register its location in seamm.ini’s [thermochemistry] section.

An existing database is never overwritten here; see update.

uninstall()[source]#

Remove the installed reference database and clear the config.

A database with local changes is left in place (only the configuration is cleared) unless uninstall --force is given.

update()[source]#

Bring the database up to the latest published version – unless it has local changes.

The database is replaced only if the local file is exactly one of the published versions (or the copy this installer installed). A database that has been edited, e.g. with curated reference energies, is kept as it is and reported; update --force replaces it, keeping a backup.

zenodo_concept_id = 21612187#

seamm_thermochemistry.report module#

A detailed, citable text report of a formation-energy calculation.

Every code that reports a formation-referenced energy used to build (and duplicate) its own ASCII table of “here is exactly what atom energies went into this number” – see e.g. the ~200-line calculate_enthalpy_of_formation this package’s formation.py replaced the arithmetic of. This module is the shared replacement for the reporting half: one place that turns a composition + a ThermoDB + a set of system energies into the explicit, per-atom breakdown table plus references, so a chemist can see exactly which database, which atomic energy numbers, and which experimental citations produced the headline number – not just trust it.

seamm_thermochemistry.report.format_report(composition, db, code, method, *, ref_type='atom', settings='', units='kJ/mol', name=None, level_label=None, system_energy=None, system_enthalpy=None, system_gibbs_energy=None, temperature=None)[source]#

Build the detailed formation-energy report text.

Parameters:
  • composition – As in atomization_energy.

  • db – As in atomization_energy.

  • code – As in atomization_energy.

  • method – As in atomization_energy.

  • ref_type – As in atomization_energy.

  • settings – As in atomization_energy.

  • units – As in atomization_energy.

  • name (str, optional) – A human-readable name for the system (falls back to a generic label).

  • level_label (str, optional) – A human-readable level of theory (falls back to f"{code}/{method}").

  • system_energy (float, optional) – The system’s electronic energy (0 K, ZPE-free). If given, the Atomization Energy and Energy-of-Formation (DfE0) sections are included.

  • system_enthalpy (float, optional) – The system’s total enthalpy H(T). If given (together with temperature), the Enthalpy-of-Formation (DfH(T)) section is included.

  • system_gibbs_energy (float, optional) – The system’s total Gibbs free energy G(T). If given (together with temperature), the Gibbs-Energy-of-Formation (DfG(T)) section is included.

  • temperature (float, optional) – Temperature, in K, for the DfH(T)/DfG(T) sections.

Returns:

The formatted report. Never raises: any missing reference data is reported as a plain-text note in the relevant section instead of stopping the whole report (a molecule that is missing, say, the standard-state entropy for one of its elements should still get an atomization-energy and DfE0 section).

Return type:

str

Module contents#

Shared atomic reference-energy database and formation-energy arithmetic for SEAMM.

See ~/Sites/reference-energy/2026-07-24_reference-energy/ for the design rationale. In short: a raw total energy from Gaussian/ORCA/Psi4/VASP has an arbitrary, code-dependent zero that means nothing to a non-expert user and cannot be compared across codes. This package holds the one shared table of atomic reference energies, and the arithmetic to turn a system’s electronic energy into a physically meaningful energy or enthalpy of formation.

exception seamm_thermochemistry.MissingReferenceData[source]#

Bases: KeyError

Raised when a composition needs reference data the ThermoDB doesn’t have.

class seamm_thermochemistry.ThermoDB(path=None, read_only=False)[source]#

Bases: object

Helper around the shared atomic reference-energy SQLite database.

Parameters:
  • path (str or Path, optional) – Database file. Defaults to DEFAULT_DB_PATH: the installer- managed location registered in ~/.seamm.d/seamm.ini’s [thermochemistry] section (database-path, normally ~/SEAMM/Parameters/thermochemistry/thermochemistry.db – see installer.py) if the installer has been run, else the bundled data/thermochemistry.db shipped with the package (a stale prototype snapshot, useful only before the installer has run).

  • read_only (bool) – Open without creating/migrating the schema, and without permission to write. Use for a package-installed, already-built database.

add_atom_energy(symbol, code, method, energy, *, ref_type='atom', settings='', units='kJ/mol', correction=None, spin_multiplicity=None, source=None, computed_date=None, code_version=None, notes=None)[source]#

Insert or update one computed atomic reference energy.

Parameters:
  • symbol (str) – Element symbol, e.g. “O”. Must already exist via add_element.

  • code (str) – Originating code: “gaussian”, “psi4”, “orca”, “vasp”, …

  • method (str) – Method/functional label, e.g. “CBS-QB3”, “PBE-D3BJ”. Where the basis is folded into the label (as in the legacy Gaussian/Psi4 tables, e.g. “CCD-FC/6-31++G(2d,2p)”), pass the full label.

  • energy (float) – The atom’s energy, in units (converted to kJ/mol on storage).

  • ref_type ({"atom", "element_phase"}) – “atom” = isolated gas-phase atom. “element_phase” = energy per atom of the standard-state phase (VASP’s existing convention).

  • settings (str) – Free-form disambiguator that is part of the identity key, e.g. “encut=700eV”. Defaults to “” (not NULL, so the UNIQUE constraint dedups consistently).

  • units (str) – Units of energy and correction: one of “kJ/mol”, “kcal/mol”, “eV”, “hartree”/”E_h”.

  • correction (float, optional) – Additive correction (e.g. Gaussian’s “<method> correction” column), same units as energy.

add_element(atomic_number, symbol, *, multiplicity=None, term_symbol=None, standard_state=None, dfH0_0K=None, dfH0_298K=None, dfH0_298K_stderr=None, h298_minus_h0_atom=None, h298_minus_h0_std_state=None, s298_gas=None, s298_gas_stderr=None, s298_std_state=None, s298_std_state_reference=None, reference=None, reference_note=None)[source]#

Insert or update the experimental reference row for one element.

batch()[source]#

Defer commits until the end of the block, for bulk imports.

add_element/add_atom_energy normally commit immediately – the right default for interactive use, but one fsync per row makes a bulk import of hundreds of thousands of rows (e.g. the full Gaussian/Psi4 composite-method grids) prohibitively slow. Wrap those calls in with db.batch(): to commit once at the end instead.

close()[source]#
coverage(code, method, *, ref_type='atom', settings='')[source]#

Return (n_elements_present, max_atomic_number) for one (code, method).

dfH0(symbol, *, at_0K=True)[source]#

Experimental atomic heat of formation, kJ/mol.

Parameters:

at_0K (bool) – If True (default, and the recommended anchor for an “energy of formation” – see the design doc), return DfH(X, 0K). If False, return the 298 K gas-phase value used by the existing enthalpy-of-formation code paths.

doi()[source]#

The Zenodo DOI of the published snapshot this database file IS, or None for a working copy that has not been published yet.

Unlike the schema version (which just says “the table layout this code understands”), the DOI identifies the actual data – the citable, permanent reference for reproducibility. Set once, right before the file is uploaded to Zenodo – see scripts/publish_to_zenodo.py and set_doi.

dump_atom_energies_csv(path)[source]#

Write the atom_energy table (joined to element symbols) to CSV.

dump_elements_csv(path)[source]#

Write the element table to CSV, ordered by atomic number.

elements()[source]#

Return all element rows, ordered by atomic number.

get_atom_energy(symbol, code, method, *, ref_type='atom', settings='', units='kJ/mol')[source]#

Return one atomic reference energy (including its correction), or None.

get_atom_energy_row(symbol, code, method, *, ref_type='atom', settings='')[source]#

Return the full atom_energy row (energy, correction, source, computed_date, code_version, notes, …) as a dict, or None.

For provenance/reporting – get_atom_energy returns just the combined number get_reference_energies/atomization_energy need.

get_element(symbol=None, atomic_number=None)[source]#

Return the experimental reference row for one element as a dict, or None.

get_reference_energies(code, method, *, ref_type='atom', settings='', units='kJ/mol')[source]#

Return {symbol: energy} for every element tabulated for this (code, method, ref_type, settings) combination.

This is the lookup table formation.formation_energy needs.

list_methods(code=None)[source]#

Return distinct (code, method, ref_type, settings) tuples present.

missing(code, method, symbols, *, ref_type='atom', settings='')[source]#

Return the subset of symbols with no tabulated (code, method) energy.

s298_gas(symbol)[source]#

Experimental standard molar entropy of the isolated gas atom, S°(X, g, 298.15 K), J/(mol K), or None.

s298_std_state(symbol)[source]#

Experimental standard molar entropy of the element’s standard-state phase, PER ATOM of X (e.g. S°(H2,g,298.15K)/2 for H), J/(mol K), or None. Needed (together with s298_gas) for a Gibbs-energy-of- formation anchor – see formation.formation_gibbs_energy.

schema_version()[source]#

The schema version stamped in this database (an int), for reports/provenance – see _SCHEMA_VERSION.

set_doi(doi)[source]#

Stamp this database file with the Zenodo DOI of the snapshot it is about to become. Call this right before uploading the file – Zenodo prereserves a DOI as soon as a deposit/version is created, before any file is attached, so the DOI can be baked into the file itself rather than only living in the Zenodo record around it.

set_s298_std_state(symbol, value, *, reference=None)[source]#

Set just the standard-state-phase entropy (and its citation) for an already-add_element-ed element, without needing to re-supply every other column (unlike add_element, which replaces the whole row).

seamm_thermochemistry.atomization_energy(composition, system_energy, db, code, method, *, ref_type='atom', settings='', units='kJ/mol')[source]#

sum(n_X * E_ref(X)) - E(system): the energy to separate all atoms.

This is exactly Gaussian/Psi4’s “Atomization Energy” (electronic-only, no ZPE, if system_energy is itself ZPE-free) and VASP’s “cohesive energy”. Positive for a bound system.

Parameters:
  • composition (dict or collections.Counter) – {element_symbol: count}, e.g. {“C”: 2, “H”: 6}.

  • system_energy (float) – The system’s electronic energy, in units.

  • db (ThermoDB) – Open reference-energy database.

  • code (str) – Identify which tabulated atom energies to use.

  • method (str) – Identify which tabulated atom energies to use.

  • ref_type ({"atom", "element_phase"}) – See ThermoDB.add_atom_energy.

  • settings (str) – Must match the settings used when the atom energies were stored.

  • units (str) – A pint-parseable unit string for system_energy and the return value, e.g. “kJ/mol”, “eV”, “E_h”.

seamm_thermochemistry.derive_dfH0_0K(db, *, overwrite=False)[source]#

Fill in element.dfH0_0K for every element where it’s derivable.

Parameters:
  • db (ThermoDB)

  • overwrite (bool) – If False (default), skip elements that already have dfH0_0K – the master sheet’s own H/He values are treated as authoritative over a re-derivation (though the derivation reproduces them to within rounding; see module docstring).

Returns:

Symbols of the elements filled in.

Return type:

list of str

seamm_thermochemistry.format_report(composition, db, code, method, *, ref_type='atom', settings='', units='kJ/mol', name=None, level_label=None, system_energy=None, system_enthalpy=None, system_gibbs_energy=None, temperature=None)[source]#

Build the detailed formation-energy report text.

Parameters:
  • composition – As in atomization_energy.

  • db – As in atomization_energy.

  • code – As in atomization_energy.

  • method – As in atomization_energy.

  • ref_type – As in atomization_energy.

  • settings – As in atomization_energy.

  • units – As in atomization_energy.

  • name (str, optional) – A human-readable name for the system (falls back to a generic label).

  • level_label (str, optional) – A human-readable level of theory (falls back to f"{code}/{method}").

  • system_energy (float, optional) – The system’s electronic energy (0 K, ZPE-free). If given, the Atomization Energy and Energy-of-Formation (DfE0) sections are included.

  • system_enthalpy (float, optional) – The system’s total enthalpy H(T). If given (together with temperature), the Enthalpy-of-Formation (DfH(T)) section is included.

  • system_gibbs_energy (float, optional) – The system’s total Gibbs free energy G(T). If given (together with temperature), the Gibbs-Energy-of-Formation (DfG(T)) section is included.

  • temperature (float, optional) – Temperature, in K, for the DfH(T)/DfG(T) sections.

Returns:

The formatted report. Never raises: any missing reference data is reported as a plain-text note in the relevant section instead of stopping the whole report (a molecule that is missing, say, the standard-state entropy for one of its elements should still get an atomization-energy and DfE0 section).

Return type:

str

seamm_thermochemistry.formation_energy(composition, system_energy, db, code, method, *, ref_type='atom', settings='', anchor=True, anchor_at_0K=True, units='kJ/mol')[source]#

The formation-referenced energy of a system.

Parameters:
  • composition – As in atomization_energy.

  • system_energy – As in atomization_energy.

  • db – As in atomization_energy.

  • code – As in atomization_energy.

  • method – As in atomization_energy.

  • ref_type – As in atomization_energy.

  • settings – As in atomization_energy.

  • units – As in atomization_energy.

  • anchor (bool) – If True (default), add the experimental atomic heat of formation so the result is a true energy/enthalpy of formation – the physically meaningful, cross-code-comparable quantity this package exists for. If False, return the bare, no-experimental-data-needed quantity: VASP’s existing DfE0 convention (ref_type="element_phase" is the natural pairing, but any ref_type is accepted).

  • anchor_at_0K (bool) – Use the 0 K experimental anchor (recommended: pairs with a ZPE-free system_energy to give an “energy of formation”, degrades gracefully to DfH(0K) if ZPE is added back in, and to DfH(298K) if the molecule’s thermal correction is added) or the 298 K anchor (matches the existing enthalpy-of-formation code paths directly). Ignored if anchor is False.

Returns:

Energy/enthalpy of formation (anchor=True) or the bare atomization-referenced energy in VASP’s DfE0 sign convention (anchor=False), in units.

Return type:

float

seamm_thermochemistry.formation_enthalpy(composition, H_system, temperature, db, code, method, *, ref_type='atom', settings='', units='kJ/mol')[source]#

DfH(T): enthalpy of formation from the elements, at temperature T.

Generalizes formation_energy(..., anchor=True, anchor_at_0K=True) from the electronic-only, 0 K quantity to a real temperature, by treating each reference atom as a monatomic ideal gas: its H(T)-H(0) = (5/2)RT exactly, so (from the elements -> atoms -> molecule Hess cycle, atom energies at T = E_ref(X) + (5/2)RT)

DfH(T) = sum(n_X * DfH0(X, 0K)) - n_atoms*(5/2)*R*T
  • [sum(n_X * E_ref(X)) - H_system]

Approximation: only the ATOM’s own H(T)-H(0) is added; the standard- state element’s own H(T)-H(0) away from 298.15 K is not separately corrected for (there is no single cheap closed form for it the way there is for a monatomic ideal gas – it depends on the standard state’s own heat capacity). This is the same approximation the legacy 298 K-only gaussian_step/psi4_step formation-enthalpy code made implicitly (it used the exact tabulated 298 K anchor rather than deriving one); the error it introduces is small near 298.15 K and grows (slowly) away from it.

Parameters:
  • composition – As in atomization_energy.

  • db – As in atomization_energy.

  • code – As in atomization_energy.

  • method – As in atomization_energy.

  • ref_type – As in atomization_energy.

  • settings – As in atomization_energy.

  • units – As in atomization_energy.

  • H_system (float) – The system’s total enthalpy H(T) – electronic + ZPE + thermal correction – on the SAME absolute energy scale as the electronic energy stored for the reference atoms (not a thermal correction alone), in units.

  • temperature (float) – Temperature, in K.

Returns:

DfH(T), in units.

Return type:

float

seamm_thermochemistry.formation_gibbs_energy(composition, G_system, temperature, db, code, method, *, ref_type='atom', settings='', units='kJ/mol')[source]#

DfG(T): Gibbs energy of formation from the elements, at T.

Same Hess-cycle construction as formation_enthalpy, extended to G. Requires each element’s s298_std_state (the standard-state phase’s molar entropy, per atom – see ThermoDB.s298_std_state) in addition to the usual 0 K atom energies and heats of formation; raises MissingReferenceData for any element missing either.

Derivation note: the reference atom’s OWN entropy cancels exactly out of the final formula (it is only an intermediate state in the elements -> atoms -> molecule cycle, and appears with opposite sign in each half-reaction), so – pleasantly – no atomic entropy data is needed at all, only the standard-state element’s:

DfG(T) = sum(n_X * DfH0(X,0K))
  • T * sum(n_X * S298_std_state(X)) / 1000

  • n_atoms*(5/2)*R*T

  • [sum(n_X * E_ref(X)) - G_system]

Approximation: S298_std_state is used at its tabulated 298.15 K value regardless of T (a standard-state solid’s entropy has no cheap closed-form T-dependence the way a monatomic ideal gas’s does – this is the dominant source of error away from 298.15 K, layered on top of the same 0 K-anchor approximation formation_enthalpy makes). At T=0 this reduces exactly to formation_energy(..., anchor=True, anchor_at_0K=True) evaluated with G_system in place of the electronic energy – a useful sanity check (G_system at T=0 with no thermal terms IS the electronic energy).

Parameters:
  • composition – As in atomization_energy.

  • db – As in atomization_energy.

  • code – As in atomization_energy.

  • method – As in atomization_energy.

  • ref_type – As in atomization_energy.

  • settings – As in atomization_energy.

  • units – As in atomization_energy.

  • G_system (float) – The system’s total Gibbs free energy G(T), on the same absolute scale as the electronic energy, in units.

  • temperature (float) – Temperature, in K.

Returns:

DfG(T), in units.

Return type:

float