Provenance. This is a proposed v2 revision of the Model state proposal: a distillation of its five-document set — the main document, evidence, walkthrough, prior art, and the spec-v2 review memo — into one self-contained document. The originals are left untouched alongside this file, the same convention the
_specV2memo used: split out so the reviewed set stays stable and the revision can be argued as a whole. Where this document and the original appendices disagree (the E3/E4/E6/E8/E9/E10 details in particular), the re-verified figures here supersede. Every checkable claim was re-verified on 2026-08-10 against icon4pyorigin/mainatde151fad8— which matters, because the landscape moved since the originals were written: PR 1301 merged on 2026-08-03, PR 1404 fixed the flagship defect, and several figures and characterizations below have been corrected accordingly. Corrections and disagreements with the original set are marked (revised). This revision also absorbs the verified findings of the parallel consolidation attempt (knowledge-base PR #18, verified against4c858a6a, 2026-08-06) after independent re-verification here where possible, so that PR can be closed. Figures that could not be re-run are labelled reported; everything else is verified. Status isdraftbecause this consolidation has not yet been human-reviewed; the pre-consolidation main document had been reviewed.A meta-lesson this set keeps re-teaching: every internal inconsistency across its revisions was one fact represented in several places and updated in some — DRY applied to prose. Mitigation here: each load-bearing count has one authoritative statement (repeats are pointers; drift between them is a bug to file), and against the originals kept alongside, the supersession note above applies. If this revision is accepted, the superseded originals should be retired rather than left to drift.
TL;DR Several incompatible designs for “how components get their fields” have been in flight at once, none stating its requirements. This document states the requirements first — split into four non-negotiable constraints and seven ranked goals — then argues that the field container should be a setup-time wiring step that emits ordinary typed dataclasses, not a global bucket passed around at run time. The mechanisms are separated so they can be adopted one at a time, and an honest cheaper alternative (“declare and check, never generate”) is presented next to the full design so the choice between them is deliberate.
1. The problem, and why now
Four designs for the component/state interface exist inside the icon4py orbit, mutually aware only after the fact — and a fifth exists outside the tree as a working experimental prototype (ICON-sc, §10):
| Design | Component signature | State shape | Status (2026-08-10) |
|---|---|---|---|
| PR 1301 / 1360 | __call__(dict[str, DataField], datetime) -> dict | per-process PhysicsState gather/scatter adapter | 1301 merged 2026-08-03 (now physics_driver package on main); 1360 open, +17.5k lines, zero reviews |
| egparedes architecture refactor | __call__(state: ModelState, step: StepInfo) -> None | one shared typed ModelState, in-place writes | draft (PR 1358 closed; doc lives here) |
| msimberg revive-components | v3: run(state: InputT) -> OutputT, dtime a field of InputT | typed frozen dataclasses + a graph-composition layer (chain/loop/when, CarrySpec) | v3 spec supersedes the v2 this document set originally argued against |
| OngChia design | __call__(state: StateView, time) -> dict | run-time StateProvider + per-field freshness | draft |
| ICON-sc (egparedes) | sympl property contracts: components declare input_properties/output_properties, dict-based array_call | dict[str, DataArray] at the public boundary, compiled at bind time into a frozen execution plan over a slotted, index-addressed StateVault | experimental architectural prototype, hosting icon4py granules; published at github.com/grAItools/ICON-sc — see §10 for references and lessons |
(revised) The original document called these “four incompatible designs” (and, inconsistently, “five”). That framing conflates two mostly orthogonal axes:
- the state axis — who allocates fields, who wires them into containers, what metadata exists, who can query it;
- the protocol axis — what a component’s call signature is and who orchestrates calls.
Rows 1 and 4 share the same dict signature and differ only in machinery; msimberg’s v3 is
mostly an orchestration design (its CarrySpec(..., initial=...) takes caller-allocated
state and says nothing about who allocates it); ICON-sc spans both axes (its own state design
and a coupling algebra); and this proposal deliberately changes no granule signature
(§4). The genuine conflicts are narrower than “five incompatible designs”: they are (a)
whether a run-time, component-reachable container exists at all — only OngChia’s requires
one, egparedes’ passes one whole, and ICON-sc keeps a run-time vault but proves components
cannot reach it and no name is resolved on the step path — and (b) allocation, on which
the four in-tree designs are silent. That silence is why the E1 class of defect survives
every one of them unchanged. (ICON-sc is the exception that proves the point: it does own
allocation — one buffer per contracted name — and thereby violates C2 instead; §10.)
Reconciling the protocol axis is still necessary — but it is a different decision, owned by
whoever owns the component protocol, and this proposal is compatible with any of the typed
outcomes. PR 1301’s merge has meanwhile made the dict protocol the de-facto incumbent on
main; see open question 4.
The design space at a glance
Before the evidence, a map — for building intuition rather than making arguments. Every design above, and every system surveyed in §10, is an arrangement of the same four elements standing on a shared foundation, inside two hard walls, with its tensions resolved somewhere on one timeline.
The elements.
- Model state — the storage for every field the discretization needs: the five prognostics, tracers, tendencies, diagnostics, static geometry/metric/interpolation coefficients, and granule-private scratch.
- Model components — dycore, diffusion, advection, physics schemes. Each reads and writes a specific subset of the state — its working set — and that subset is the component’s public interface, whether declared or merely implied by its code.
- The main loop — orders component execution (including sub-stepping and per-component call frequency), owns time-level swaps, and applies or delegates state updates.
- Cross-cutting services — output, checkpoint/restart, halo exchange, nudging. Not components: they need to select fields by role (“everything restartable”), not by dataflow, and today each keeps its own hand-written list.
The foundation all four stand on is the vocabulary: a field name must mean exactly one thing before any interface can be declared or any subset selected. This is a design element in its own right — icon4py currently has five parallel namespaces per field and a live name collision (E8, E9), and the honest cost estimate of the metadata mechanism (M2) is “small mechanism, large vocabulary”.
The two hardest walls — constraints, not tensions; they delete whole design families outright (§3 holds two further constraints: C3, which deletes whole-state arguments, and C4):
- C1: whatever reaches a
gtx.programmust be a static named collection — a name-keyed map can never cross the stencil boundary, so some typed container always exists at the last mile. - C2: in the Fortran-embedded path ICON owns the memory, so the state layer must be able to adopt buffers it did not allocate — designs that insist on owning allocation (including ICON-sc’s) do not fit this deployment.
The timeline is the dimension that actually separates the designs. Each tension below can be answered at one of four phases — declaration (in the source), setup (once, before the loop), per step, or per stencil call — and the surveyed systems differ less in structure than in when they answer: CCPP resolves working sets at build time, ICON-sc at bind time, sympl and the merged physics driver per call. The thesis of this document, in one line: push every answer to the earliest phase that can hold it — and for everything here except the derived-quantity barrier (M5-lite) and halo exchange, that phase is setup or earlier (M6-structural’s wiring-validity checks fire only when a binding changes, not per step).
The map, with each tension’s options and where its evidence and mechanism live (E/M references are forward pointers into §2 and §5 — this table doubles as a reading guide):
| Element | Tension | Options (phase in parentheses) | See |
|---|---|---|---|
| Vocabulary | what is a field’s identity? | flat string vs (quantity, placement) key; CF vs icon:; declared once vs re-derived per subsystem | E8, E9 / M2 |
| State | is the element set static or config-dependent? | fixed type with config-dependent allocation (setup) vs dynamic registry (run) | §4 / M11 |
| State | who allocates, who owns? | registry allocates (setup) vs every caller allocates (the status quo) vs adopt external buffers (C2) | E1, E2 / M1 |
| State | lifetime and visibility | persistent vs granule-private scratch vs first-class absence | M10, G6 |
| Components | how is the interface declared? | explicit metadata (sympl, CCPP) vs bare Python signature vs metadata on the signature (NDSL, this proposal) | M2, M3 |
| Components | who builds the working set, and when? | component gathers per call (run) vs orchestrator emits a typed container (setup) | E6, E7 / M4 |
| Components | update policy | in-place mutation vs returned tendencies applied by the loop | ADR 0001; the five designs |
| Loop | how is the graph assembled? | hand-coded driver vs combinators vs config/IR — ordering contracts declared or implicit | M13; msimberg v3 |
| Loop | at what rates do parts run? | one loop vs dyn substeps vs per-component frequency | G7; OngChia |
| Loop | derived-quantity consistency | per-consumer re-derivation (run, the status quo) vs one named barrier over a closed set (run, bounded) vs lazy recompute (rejected) | E3, E4 / M5-lite |
| Services | how do sweeps select their sets? | hand-written lists (the status quo) vs label queries materialized at setup | G4, G5 / M7 |
Note the split hiding inside “slicing”, because it decides two different mechanisms: a component’s working set is dataflow subsetting — static per component, resolvable once at setup into a typed container (M4) — while a service’s sweep is role subsetting — a config-dependent query over field metadata (M7). Conflating them is how “which container a field sits in” ends up standing in for “what role it has”, which is exactly the confusion G5 exists to remove.
2. Evidence: what the current shape has cost
Defects verified in icon4py main; status re-checked at de151fad8 (2026-08-10). Paths
relative to the icon4py repo root. This list is not a migration bill — it is the requirement
source; each defect forces a requirement in §3.
| # | Defect | Status | Severity |
|---|---|---|---|
| E1 | PrepAdvection and AdvectionPrepAdvState held the same 3 quantities, allocated separately, with nothing copying one into the other — standalone-driver tracer advection ran on identically-zero trajectory velocity and mass fluxes | fixed by PR 1404 (2026-07-30) — see below | correctness |
| E2 | Their mass_flx_ic disagreed on vertical extent (nlev+1 vs nlev), so even a copy would not have been shape-compatible; fa.CellKField[float] cannot express the difference | fixed with E1; PR 1404’s zero-field fallback now also allocates extend={KDim: 1}, settling the extent as half-levels | correctness |
| E3 | T/Tv/p derived in at least three places on main from different inputs — and worse than first stated: the IO path runs the diagnosis fully dry. driver_io.py:146-147 allocates qv, qc, qi, qr, qs, qg once, all permanently zero (“dry air: all hydrometeors stay zero (never written, so allocated once)”) — qv included — so on the output path virtual_temperature ≡ temperature identically, and the temperature icon4py publishes is a dry-air temperature no physics component ever uses | live (driver_io.py:181, muphys state.py, jablonowski_williamson.py; PR 1360 adds a fourth) | science |
| E4 | vn → u,v computed at two production sites into two different buffers — driver_states.py:290 writes diagnostic_state.u/.v and halo-exchanges after (:302); driver_io.py:199 writes its private _u/_v and does not. The domain bounds do not differ (both run lateral_boundary_level_2 → END); the defect is the duplicated buffers plus the missing exchange. driver_io.py:115 carries TODO(kotsaloscv) asking for exactly this refactor | live | correctness |
| E5 | dycore.InterpolationState (16 fields) is a strict superset of DiffusionInterpolationState (8); geofac_div declared in 3 containers; ddqz_z_full in 3; byte-identical one-field dataclasses in microphysics | live | duplication |
| E6 | driver_utils.initialize_granules: 91 hand-typed .get() mappings (count drifts upward with each merge), replicated at ~12 further sites — the union of every file hand-constructing a granule interpolation/metric container (the earlier “7 sites” was an undercount; 11–13 depending on tree date and definition), of which 2 are the production Fortran binding wrappers | live | boilerplate |
| E7 | Wrong-key bugs in exactly that replicated hand-mapping: d2dexdz2_fac2_mc=…get(D2DEXDZ2_FAC1_MC) and edge-normal into a dual-normal slot. The intended mapping is confirmed independently at three other sites (driver_utils, test_diffusion, test_parallel_geometry), so these are filable today with the evidence attached | live (integration_tests/test_benchmark_solve_nonhydro.py:98,163; test_benchmark_diffusion.py) — both type-check, both run; benchmark paths, not the production driver, and the primary site driver_utils is correct, which argues for deleting the replication rather than distrusting it | correctness |
| E8 | Placement recorded three times per field: name string (…_at_cells_on_half_levels), dims=(CellDim, KHalfDim), is_on_half_levels: bool — the factory rewrites KHalfDim → KDim at allocation (factory.py:540,545,823, “remove once gt4py supports vertically staggered dimension”), so half-levelness survives only in a metadata tuple nothing validates. And the triplication is now tested into place: is_on_half_levels is read at driver_io.py:83,242, and a test_driver_io docstring states that dims and is_on_half_levels are “stated independently” — deliberately | live | drift |
| E9 | Five parallel namespaces per field (catalog key / CF name / ICON Fortran name / dataclass attribute / port name); four disjoint metadata dicts plus one orphan; none of the 49 metrics entries is a real CF name; units="" for essentially all metrics and interpolation entries (though not geometry — the earlier “most” was too broad); and a live collision: metrics_attributes.py:106-107 declares two distinct fields (cell and edge) under one standard_name, which the IO writer’s filter_by_standard_name resolves variable identity by — they would silently alias into one netCDF variable | live | drift |
| E10 | The trend, verified on origin/physics_driver_tmx (2026-08-05): PR 1360 adds seven state dataclasses and 92 declared fields for one component (TmxDiagnosticState 31, TmxMetricState 17, TmxInputState 16, …). TmxInputState re-declares all six fields of the common DiagnosticState, all six tracers, and rho/w; TmxMetricState re-declares five metrics fields; TmxInterpolationState makes geofac_div a fourth container; gather_from_prognostic re-derives T/Tv/p a fourth time | live (PR open) | trend |
E10 is the important row: duplication is being created faster than it is removed, because every new component pays the full adapter-stack tax.
E1/E2’s fix is itself the strongest argument in this document. PR 1404 added
initialize_prep_tracer_advection (driver_states.py:222): a hand-written function whose
whole job is to alias three of the dycore’s buffers into a second container, plus an identity
test asserting the aliasing holds. Both are now permanent maintained surface, and every
future producer→consumer pair starts from the same footing. The fallback branch is the
sharper point: with no dycore, the same function allocates fresh zeros, and to get
mass_flx_ic right it must restate the half-level extent in a second file, with a
comment explaining why (“one more level than KDim, like the dycore’s … it stands in for”)
— E2’s knowledge, one quantity’s vertical extent, now represented a third time, and that
representation was created by the fix. Under M1 (one quantity → one buffer, §5) neither
the function nor the test would need to exist, and the class of defect would become
unrepresentable on the registry path — see §6 for the honest scope of that claim.
Claims from the original investigation that remain unverified — treat with suspicion:
- All memory numbers (~185 MB redundancy, ~290 MB granule scratch, etc.) were derived from an assumed grid not read from any file, and conflict with other estimates. Order-of-magnitude only. The load-bearing relative claim — granule-private scratch exceeds cross-granule duplication — should be measured before being cited.
- Micro-benchmarks (label filtering ~4 µs/300 fields,
gtx.as_fieldcopy cost) were measured on one laptop. - Prior-art line numbers for sympl, climt, NDSL/Pace, ClimaAtmos, MPAS and CCPP are paraphrase-grade; only ICON, LFRic, gt4py and ICON-sc were read from local checkouts.
- A claimed tracer double-buffering bug tied to
ndyn_substeps_varparity was derived by counting swaps; no test, no observed wrong answer.
3. Constraints and ranked goals
(revised) The original R1–R11 was a flat, unweighted list — the Working Principles’ “advocate-less wish list” red flag. Split and ranked as the principles demand:
Constraints — non-negotiable; violating one makes a design wrong, not merely worse.
| # | Constraint | Forced by |
|---|---|---|
| C1 | Whatever reaches a gtx.program must be a static named collection | gt4py, structural — see §4 |
| C2 | The container must adopt externally-owned buffers — at solve_nh_run ICON owns the memory | the Fortran-embedded path |
| C3 | Granule call sites must keep naming their actual inputs | the havogt/msimberg stamp-coupling objection; also implied by C1 |
| C4 | No new per-stencil-call overhead | ~100 stencil calls per 20–50 ms timestep |
C2 holds today with high confidence — py2fgen’s whole type model is
ParamDescriptor = ArrayParamDescriptor | ScalarParamDescriptor (no record descriptor, so
no struct can cross the ABI), and bindings/tests/test_codegen_references.py is a
golden-file test of the generated Fortran/C bindings, so any wrapper-signature change breaks
a checked-in artifact. What is unexamined is the assumption that the embedded path is
permanent: it is why the registry may not own allocation and why the adopt-external seam
exists, and if it ever stops being true the design gets materially simpler. Worth re-asking
explicitly before committing (open question 1).
C4 is real but nearly satisfied by construction in the setup-time reading: the two run-time admissions are M5-lite (one profiled per-step barrier) and M6-structural’s checks, which fire only at wiring/rebinding events plus an optional debug-mode sweep — nothing touches the per-stencil path. Listing C4 as a peer constraint invites a performance debate the design does not need and, per §4, should not be argued on.
Goals — ranked; a team taking “G1–G3 and stopping” is a coherent outcome.
| rank | # | Goal | Forced by |
|---|---|---|---|
| 1 | G1 | One quantity → one buffer; shape and placement declared once | E1, E2, E5, E8 |
| 2 | G2 | A cross-granule producer→consumer handoff must be expressible, so it cannot silently become two allocations | E1 |
| 3 | G3 | A derivation (vn→u,v; theta_v,exner→T,p) must have exactly one implementation, with domain and halo semantics part of the declaration | E3, E4 |
| 4 | G4 | Cross-cutting sweeps (output, restart, halo sets) must be queries, not hand-written lists | restart lists hand-picked; tracers have no IO path |
| 5 | G5 | A field’s role is not implied by which container it sits in | the restart inventory, §9 |
| 6 | G6 | Absence must be first-class — optional IAU increments, inactive tracers — not a zero allocation | dummy_field_factory |
| 7 | G7 | Multi-buffer/time-level must be expressible, at different rates (dyn substep vs tracer step) | TimeStepPair, PR 1404 |
G1 ranks first because it is the only goal that makes defects unrepresentable rather than
merely detected, and because G2, G4 and G5 lean on it. G7 ranks last because TimeStepPair
already works and nothing there is broken. G6 is not hypothetical: the bindings already fake
absence with dummy allocations (wrapper_common.cached_dummy_field_factory, used for
hdef_ic/div_ic/dwdx/dwdy and the optional IAU increments
vn_incr/rho_incr/exner_incr).
The budgeted resource is reviewer attention, not runtime. PR 1301 took two months and 102 review events to merge; PR 1360 sits at +17.5k lines with zero reviews. Two consequences: the adoption order matters more than the end state, because each rung must survive review independently; and performance was the wrong axis to argue on (§4).
The user model, stated explicitly (better wrong than vague): the user is a physics or
dynamics developer adding or modifying a scheme — fluent in ICON and gt4py stencils, not a
software architect, and not someone who reads a framework manual first. Success means adding a
field edits one place, and a mistake fails at setup naming the field, not at timestep 3000
with wrong numbers. The second user is the reviewer, who must be able to answer “who writes
this field?” from the diff — that reviewer is the entire justification for recording intent
(§5, M2).
4. The design decision: the container is a setup-time emitter, not a run-time bucket
The objection from havogt and msimberg — “one global bucket passed around in its entirety when only part of it would be enough” — is correct and stronger than it sounds: it is stamp coupling escalating toward common coupling, and the lazy-shared-store variant is the Blackboard pattern, whose own POSA liabilities list reads “difficulty of testing, difficulty of establishing a control strategy, low efficiency, no support for parallelism.”
But the objection applies to a run-time bucket. The decisive observation, checked defect by defect:
Every defect in §2 is created at setup time — in allocation and wiring code that runs once before the time loop. None of them requires run-time name resolution to fix. (revised: the original said “not one of them recurs per timestep”, which overstates E3/E4 — the duplicated derivations do execute every step; what is setup-time is the choice of which implementation and which inputs each site is wired to.)
So the container consumes declarations and emits bindings: it assembles the typed containers granules already take, then it is done — a factory floor, not a warehouse. Three tests for any proposal in this space:
- Schema test — can the schema be settled before the time loop? For icon4py yes: the
tracer set is a pure function of
TracerConfig, the output set of the namelist, the component set of config. A hard freeze is not required: ICON-sc (§10) shows a staleness guard beats a freeze — mutation stays legal, and running against a stale wiring raises (~100 LOC, forbids nothing). - Reachability test — can a granule reach a name-keyed store at run time? If yes with
string lookups on the compute path, you have rebuilt MPAS’s pools, which MPAS-Ocean deleted
for GPU and whose successor Omega dropped entirely. (revised) A typed whole-state
argument — egparedes’
ModelState— is a different, milder failure: no string lookup, no dynamic hash, gt4py-compatible per member. What it violates is C3: the call site no longer names its inputs, so the reviewer cannot answer “what does this component read?” from the signature. Strictly these are two tests — reachability of a name-keyed store (structural), and interface honesty (= C3) — and the two designs fail one each. - Emission test — is the output an ordinary dataclass gt4py accepts?
gt4py forces the last mile (verified from source)
gt4py/next/named_collections.py:36-49 — CustomDataclassNamedCollectionABC.__subclasshook__
accepts a type only if it is a dataclass (outside gt4py.*) whose every
__dataclass_fields__ entry has init=True, default is MISSING,
default_factory is MISSING, and is a regular field. Consequences, checked against the live
tree:
- A
dictcan never be agtx.programargument. Whatever is built, the last mile is always an explicit typed collection — so C3 is not a compromise, it is mandatory. - Any defaulted field disqualifies the container.
TracerState(tracer_states.py:106-116,qv: … | None = None×6) is disqualified today by its defaults;PrognosticStatebecame conformant when PR 1404 removed its defaultedtracerfield. - (revised — mechanics the original set had subtly wrong) A
| Noneannotation alone does not disqualify — the hook never inspects annotations, only defaults. And aClassVarmember does disqualify (the regular-field clause;__dataclass_fields__retains ClassVar pseudo-fields). Practical consequence: attaching class-level metadata to a state dataclass silently strips its named-collection status — the msimberg specs putinputs_properties/outputs_propertiesasClassVaron the Component, which is safe; putting them onInputTwould not be. And M11’s optionality fails in the wrong place: a default-free container holding aNonevalue passes the structural check and fails later, inside extraction — so whether a container is a wiring object or a program argument must be part of its declaration, checked atseal()(None-free at build for program arguments), not discovered when gt4py rejects a value (§5, M11).
Do not argue this on speed. ICON-sc measured the entire benefit of moving all negotiation/lookup out of the step loop at ~6.8 % of step time on a real model (JW R02B04×35, gtfn_cpu, 3.68 → 3.43 s/step; reported — ICON-sc’s own measurement, not re-run here); the eye-catching 64–101× figures circulated earlier are from a kernel-free toy, and ICON-sc’s own architecture doc is blunt: “a dict lookup is ~40–60 ns and slotted attribute access ~20–40 ns, but those were never the real cost.” The case for typed dataclasses rests on the structural prohibition above (C1), on Fortran buffer adoption (C2), and on type-checkability and explicit ownership — not on lookup cost.
The strongest objection, answered
egparedes, with our exact use case in view: “The schema is configuration-dependent (the tracer
set alone varies), so no static dataclass can be the public state type.” Correct about a
static type — and answered by a setup-time emitter only if M11 (conditional allocation
from config predicates) is actually built, which is why M11 sits at step 2 of the adoption
order rather than being a footnote. The type is fixed; the allocation set is config-dependent;
inactive means no buffer and a build-time-checked absence, not a zero field. The limit of
this answer should be stated too: it covers closed config spaces (the fixed qv…qg tracer
set). A genuinely open-ended tracer list (chemistry, ART aerosols) cannot be a fixed dataclass
and stays on the wiring side, with per-tracer fields extracted at call sites — which is
already how tracer_advection.run(p_tracer_now=…) works today.
Everyone who shipped this converged on the same answer
Every surveyed production system either resolves state wiring before the loop or regrets not doing so: ICON’s compute path never queries its registry, CCPP resolves at build time, NDSL/Pace stayed with rigid dataclasses plus field metadata, ClimaAtmos is migrating back to explicit structs, and LFRic’s own retrospective rejects its global store. The per-system verdicts and quotes are in §10’s table.
Political corollary: the setup-time reading changes no granule call signature. The
wiring intervention lands in driver_utils.initialize_granules,
driver_states.assemble_driver_states and the test-side repacking — precisely the code that
is already replicated and already contains the shipped wrong-key bugs (E6, E7). The
declarations — one spec() per field — land in the state-class modules across the granule
packages; §6 costs that honestly.
5. Mechanisms
Separated so they can be adopted independently. Bucket? = needs a globally reachable mutable name→field map at run time.
| # | Mechanism | Phase | Bucket? | Requires | Cost |
|---|---|---|---|---|---|
| M2 | Metadata on dataclass fields (standard_name, icon_name, units, dims, origin, intent, scope, restart) | setup | No | — | S mech / L vocabulary |
| M10 | Scope/lifetime tag + granule-private Local scratch type | setup | No | M2 | S |
| M11 | Conditional allocation from config predicates | setup | No | M2 | S |
| M3 | Declared I/O used for validation only | setup | No | M2 | S |
| M1 | Canonical allocation registry, signatures unchanged | setup, sealed | No | (M10) | M |
| M4 | Declared I/O → automatic wiring (emit the dataclass) | setup | No | M2 | M |
| M7 | Labels/groups, materialized at setup | setup | No | M2 | S / M vocab |
| M8 | Units: validate, don’t convert | setup | No | (M2) | S |
| M12 | Declared handoff + consumer/producer arity check | setup | No | M1, M2 | S (~90 LOC) |
| M13 | Ordering constraints as declared data (must_follow / must_precede) | setup | No | M2 | S |
| M14 | Parameters as a structure distinct from state | setup | No | M2 | S/M |
| M5-lite | One named update_derived_quantities() barrier over a closed, enumerated set | run, per step | No | M2 | S/M |
| M6-structural | Wiring-validity counters (epoch/generation/schema_hash) | setup; checks fire on rebinding events | No | M1, M2 | S (~100 LOC) |
| M5 | Lazy derived-field computation | run | No | M2, M6-scientific | H — deferred |
| M6-scientific | Derived-field consistency tracking | run | Yes | M1, M2 | H — deferred |
| M9 | Automatic regridding as registered rules | run | No | M2, M5, M6-scientific | H — deferred |
A parenthesized requirement is soft — the mechanism works without it but is materially
better with it: M1 without M10 makes granule-private scratch globally reachable; M8 can
bootstrap over the existing *_attributes.py dicts before M2 lands (which is why it also
appears as a step-0 free win).
Notes on the ones that matter most:
- M1 kills the E1 class by construction. One buffer named per
(quantity, placement)key, however many container fields reference it. Half already works: the static-field factories memoize, so the driver path already aliases metrics/interpolation fields — but nothing declares that sharing and nothing checks it (coincidental correctness: the savepoint test path already breaks it by copying). The gap is that the time-varying half of the state never touches any factory. - M2 is the enabler for everything, and its hard part is not code. The mechanism is
NDSL’s, verbatim:
spec(...)wrappingdataclasses.field(metadata=…), which sets no default and is therefore free under the named-collection rules (§4). Includeorigin/K-domain from the start — gt4py fields carry a domain, not a shape, and omitting it cost ICON-sc two work units of misdiagnosis. The vocabulary is the real cost: ~80 % of an atmospheric model’s fields have no CF name (ICON-sc measured 18 CF / 72icon:-namespaced). Adopt the two-way invariant — unprefixed ⇒ claims CF identity; no CF name ⇒ must beicon:<name>— enforced at registration. Two precisions that keep the vocabulary from becoming yet another namespace: the canonical identity is the(standard_name, dims)pair —standard_namebeing the one canonical quantity name (CF where one exists, elseicon:-prefixed; §6’squantity=is this same key, not a sixth namespace) — and G1’s declared once is realized through the name file of adoption step 1: it is the single authority for name → (dims, units, origin), per-containerspec()entries repeat the name, andseal()rejects any entry that contradicts the file. Declared once, referenced N times, checked mechanically. Who owns the vocabulary is open question 6. - M3 is the highest value per line (~150 lines) and the one thing every proposal already
agrees on. The current
Componentprotocol’s declared properties are completely inert — re-verified after PR 1301’s merge: the shippedphysics_driver.runnever consultsinputs_properties/outputs_properties; they remain decorative onmaintoday, and the protocol’s own TODOs (unit matching, dimension consistency) are still open. One nuance makes M3 cheaper than greenfield:muphys/component.py:49does declare both properties as plain class attributes — nothing reads them, so M3 has a first consumer waiting. - M10 covers the larger memory number. Granule-private scratch (~29 full-3D fields across
SolveNonhydro/Diffusion/VelocityAdvection) plausibly exceeds cross-granule duplication (unmeasured, §2). A shared container without a scope tag makes memory strictly worse by making every private buffer globally reachable and never freed. NDSL’sLocalpoisons the buffer at init, sets DaCetransient=Trueso the compiler can elide it, and enforces scope at run time — the single most transferable idea in NDSL. - M11 answers the config-dependence objection (§4) and, with M10, is one of only two
mechanisms that reduce memory. Enforcement point (revised):
reg.build(T, config=cfg)allocates only fields whoseactive_whenpredicate holds, binds the rest toNone— and a container declared as a program argument is validated None-free at build, because gt4py’s structural check will not catch a None value, only a defaulted field. - M12 gives G2 a mechanism. Declare each producer→consumer handoff; at seal, 0 consumers
and ≥2 consumers both reject (a dangling tendency silently loses physics; a double consumer
double-applies it), and 0 producers is E1 exactly. No runtime object — the check runs once.
It catches the E1 class only on top of M1+M2 (one quantity ⇒ one name ⇒ one buffer);
without those, two same-named slots in different containers are both legitimate and the check
is decorative. Do not assume one producer: ICON genuinely sums multiple publishers into
ddt_*slots, so publisher multiplicity must be declared rather than assumed. And as with M13, the production API should take references rather than free-form strings — the string form in §6’s sketch is illustrative only. - M13 is the safety net a configurable component order needs and does not have. ICON’s
fast-physics ordering carries implicit contracts (saturation adjustment twice per step,
surface transfer last, turbulence on old-time-level inputs) that live today in tutorial
prose. Declared on the component (
must_follow/must_precedeas references, not strings — ICON-sc matched free-form strings, so a typo silently passes), they become a build-time assertion. Directly relevant to OngChia’s user-configurable ordering and msimberg-v3’s graph combinators, both of which make reordering structurally easy and scientifically treacherous. - M14 keeps calibration constants out of state. Tunable scheme parameters (entrainment coefficients, autoconversion thresholds) declared as a structure separate from state, so they are never smuggled through state fields. Right for ensembles, perturbed physics and namelist provenance; needs none of the differentiability machinery it originated in.
- M6 splits in two, and only one half is worth building. Structural staleness — “is my
wiring still valid?” — is cheap and replaces the freeze:
epoch(a field’s identity changed ⇒ wiring stale ⇒ raise),generation(a time-level swap ⇒ only cached views drop), plus a debug-build renegotiate-and-diff every N steps. The rule: values are the caller’s business, identities are the wiring’s — in-place writes stale nothing; rebinding raises instead of silently computing on a dead buffer.prognostic_states.swap()therefore invalidates nothing. Scientific staleness — “is this derived field consistent with its inputs?” — has no working prior art anywhere in the survey, can never be a correctness guarantee (gt4py fields hand out writable.ndarraybuffers), and is deferred; M5-lite’s barrier removes the need. For calibration, icon4py today has zero invalidation machinery of either kind: a grep forinvalidat|stale|dirty|recomputeoverstates/and the factories returns nothing, memoized providers return a cached field forever, and there is no evict API. - M5/M9 carry the loudest warnings. CCPP has wanted
theta_v,exner→T,pderivation for years, has not built it, and scopes its issue “this is not an open-ended task!”; thetheta_v,exner ↔ T,pcycle is real (ClimaAtmos documents the identical cycle and breaks it with a physical approximation, not a solver). Ship M5-lite instead: one named, profiledupdate_derived_quantities()barrier over a closed, enumerated set of derivations. That kills E3 and E4 at ~10 % of the cost and stays compatible with full M5 later. Full laziness is additionally barred by bit-reproducibility (§8).
On labels (M7), from ICON’s twenty years: the label is declared at the field’s definition
site, by its owner — new field + right group string ⇒ it appears automatically in output,
restart-analysis, IAU, LBC prefetch, meteograms, plugins; no central list to edit; that is why
each of those services is ~200 lines instead of ~5000, and the namelist even gets set algebra
('group:atmo_ml_vars', '-qg'). Three amendments from ICON’s own scars: unknown label must
raise (ICON auto-creates on first use — a documented typo trap; MPAS’s silent-null
equivalent shipped bugs for a decade); materialize buckets at setup, never query per call;
labels are a selection mechanism, never an addressing or placement mechanism.
On naming and placement: key on (quantity, placement) where placement is the dims tuple,
and treat the flat string theta_v_at_cells_on_half_levels as its rendering, not the primary
key. dims and is_on_half_levels become derived rather than independently maintained (fixes
E8), and regridding becomes an edge over a fixed quantity — impossible if placement is welded
into an opaque string. CF names stay as output metadata only. Escape hatch required: not
everything factorises (rbf_vec_coeff_e; vn vs u,v differ by more than placement) — those
stay plain named fields with no derivation rules.
6. Concretely
Today’s pipeline has three setup phases and a loop; every §2 defect is created in the setup phases. The proposal changes only those.
# ─── TODAY ──────────────────────────────────────────────────────────
# A: static fields — three memoized factory sources (already lazy, already aliasing;
# this half of the problem is solved, but nothing declares or checks the sharing)
# B: hand-wire granule containers — 91 hand-typed `.get()` keyword mappings in
# driver_utils, replicated at ~12 sites (E5, E6, E7)
# C: allocate mutable state — driver_states.assemble_driver_states; before PR 1404
# this allocated the same three quantities twice, disconnected (E1, E2, E8)
# D: the time loop — dycore accumulates into prep_adv over ndyn_substeps_var substeps;
# advection reads it once per step; prognostic_states.swap()
# ─── PROPOSED ───────────────────────────────────────────────────────
# A′: DECLARE — the container classes stay exactly the frozen dataclasses granules
# take today; each field additionally says what it IS (M2, NDSL's mechanism):
@dataclasses.dataclass(frozen=True)
class DiffusionInterpolationState:
geofac_div: gtx.Field = spec(
quantity="geofac_div", # canonical name — ONE per quantity (same
# string the dycore's container uses)
icon_name="geofac_div", # the ICON Fortran name, for the bindings
dims=(CellDim, C2EDim), # placement, single source of truth → E8
units="1", intent=READ, scope=STATIC,
)
...
# B′: REGISTER + EMIT — setup, once:
reg = FieldRegistry(grid, vertical_grid, allocator)
reg.adopt_sources(static_sources) # adopt the memoized factories, don't re-allocate
reg.declare(PrognosticState, DiagnosticStateNonHydro, PrepAdvection,
AdvectionPrepAdvState, DiffusionInterpolationState, ...)
reg.declare_handoff("vn_traj", producer="dycore", consumer="advection") # M12
reg.seal() # the QUANTITY SET is now fixed; rebinding a quantity to a different
# buffer stays legal (bumps epoch), and values are unrestricted
prep_adv = reg.build(PrepAdvection)
tracer_prep_adv = reg.build(AdvectionPrepAdvState) # same quantities ⇒ SAME buffers
...
# C′: gone — allocation happened inside build, once per quantity.
# D′: the time loop — IDENTICAL to today. `reg` is not mentioned; no lookup happens;
# nothing is lazy; no dict crosses a stencil boundary.What each defect costs then: E1’s class becomes unrepresentable (one buffer per quantity —
build hands both containers the same object, no aliasing function, no identity test); E2/E8
become a contradiction the registry rejects at seal() (two containers claiming different
dims for one quantity); E6 collapses to one build line per container; E7 becomes
impossible (no keyword list left to mistype); E5’s sharing becomes declared instead of
coincidental.
The honest cost side:
- One
spec()per field across the model —MetricStateNonHydroalone is 32 fields; a few hundred declaration lines total. Real work. But it replaces more than it adds: the declaration is written once per field; the hand-mapping it deletes was written once per field per site (~12 sites), and the ~50-line repack in each Fortran binding wrapper becomes a table walk over the same declarations. - A wrong
specfails atseal()naming the field — today the same mistake is E7: it type-checks, runs, and produces wrong numbers. - Tests keep working untouched. The containers remain ordinary constructors;
reg.buildis an additional path, not a replacement. That matters across ~269 test files. The flip side scopes the headline claim: because direct construction stays legal, E1’s class is unrepresentable only on the registry path, and prevention off that path is by convention. The backstop must be named, and it is cheap: a test asserting the production drivers construct no state class except throughreg.build— a grep-level check that cannot be delegated totach, which currently enforces nothing (§8). - New machinery on the setup path that must itself be reviewed and maintained.
The Fortran-embedded path survives (C2). At solve_nh_run (dycore_wrapper.py:306) ICON
owns the memory: 37 raw fields + 10 scalars under ICON Fortran names, re-packed into five
containers on every call (:370-421) — a third hand-maintained copy of the ICON↔icon4py
name map (the other two: the icon_var_name entries in the *_attributes.py vocabularies,
and the container docstrings). Under the proposal the wrapper calls reg.adopt_buffers({...})
(zero-copy wrapping is already icon4py’s own technique) and build reads the same
declarations. Two trade-offs to state rather than bury: per-call adopt_buffers + build
adds validation work to a per-timestep call, so its cost must be measured before this
path ships — memoizing on pointer identity (ICON’s are module-level allocatables, so very
likely stable) is the likely fix, not a requirement; and because nothing is held across
calls on this path, it deliberately forgoes the staleness guard (epoch is never consulted
there) — an accepted asymmetry, not an oversight. Two hard consequences stand: the registry
must adopt buffers it did not allocate, and no struct can cross the ABI (py2fgen has no
record descriptor; the wrapper keeps flattening).
The distributed path is a requirement on determinism, not a new mechanism. Everything
above must also hold under the GHEX/MPI decomposition, which adds four rules. The registry
is built per rank from the rank-local grid and decomposition — dims metadata describes
the rank-local allocation, while a field’s identity stays global (the quantity key). Every
seal-time check (M3, M11, M12, M13) must be a deterministic function of configuration and
schema alone, so it passes or fails identically on every rank — a raise on a single rank
inside otherwise-collective setup is a deadlock, and the same determinism requirement binds
M6-structural’s raise-on-stale. G4’s “halo sets as label queries” selects fields; the
exchange objects and their scheduling stay owned by the decomposition layer. And M5-lite’s
barrier owns the halo exchanges its derivations declare — which is the half of E4 (the
missing exchange) that consolidating the derivation alone does not fix.
The alternative that must stay on the page: declare and check, never generate
Design-it-twice demands the genuine competitor, and it is not a straw man — it is the honest stopping point: adoption steps 0–1 plus step 3, i.e. M2 + M3 only (B takes neither M10/M11 nor anything after):
Alternative B. Add per-field metadata (M2) and a validator (M3), then keep
driver_utils exactly as it is. The hand-written wiring stays and becomes checked: declared
dims that contradict another container’s raise at startup; a ~10-line field-coverage test per
site catches drift. One check must be named explicitly for B’s E7 claim to hold: the checker
compares the bound value’s declared identity against the receiving slot’s declaration —
coverage and dims checks alone pass both shipped wrong-key bugs.
| A — emit the wiring | B — check the wiring | |
|---|---|---|
| E7 (wrong keys) | yes — no keyword list survives | yes — the check catches it |
| E2/E8/E9 (contradictions) | yes | yes |
| E6 (boilerplate × ~12 sites) | yes | no — all of it stays |
| E1’s class | yes — one buffer per quantity | no — a checker can report that two containers disagree; it cannot make them one buffer |
| E3/E4 (divergent derivations) | needs M5-lite either way | needs M5-lite either way |
| Cost | M1 + M4, a few hundred declaration lines, allocation moves | M2 + M3 only |
| Risk | new machinery on the setup path | almost none |
The E3/E4 row deserves its own sentence, because it bounds what either design can claim: those two defects are not fully setup-time — the divergent values are produced every step — so no amount of wiring, emitted or checked, fixes them. They need M5-lite’s per-step barrier, the single place where this design admits a run-time mechanism.
If the team never goes past B, that is a coherent outcome: the correctness defects are fenced and the boilerplate remains. A wins only if E6 and E1’s class are judged worth the extra machinery. Stating this makes the A-vs-B choice deliberate instead of something that happens by drift.
7. Adoption order
Each step independently shippable, each with standalone value — which is what the budgeted resource (reviewer attention, §3) demands.
- Free wins, no design commitment. File E7’s wrong-key bugs now — verified, live, and
one-line fixes; the per-site field-coverage test (~10 lines each — E6/E7 drift turns red
today); units-as-identity-validation (~110 LOC, no dependencies); the
icon:namespace two-way invariant. - M2 — metadata on dataclass fields,
originincluded; start the name file. Reuse thekindkeystates/model.pyalready defines rather than inventing a parallel one. - M10 + M11 — scope tag and config-predicate allocation: the only two mechanisms that reduce memory, and M11 is what answers the strongest objection (§4).
- M3 — validation at class creation. Highest value per line; agreed by every proposal. — Alternative B is steps 0–1 plus this step, coherently. —
- M1 — canonical allocation, settled at setup. Kills the E1 class. No signature changes.
- M12 — handoff arity check; cheap once M1+M2 exist.
- M4 — auto-wiring, gated on M2’s vocabulary being real: auto-wiring keyed on today’s
vocabulary will silently bind the wrong field (
metrics_attributes.py:106-107already declares two distinct fields under onestandard_name). Deletes E6. - M7 — labels: unlocks output/restart/checkpoint sets; gives tracers an IO path at all.
- M5-lite — one derived-quantities barrier over a closed set, owning the halo exchanges its derivations declare (§6, distributed path). Kills E3 and E4; the only per-step compute addition in the whole sequence.
- M6-structural — adopt whenever setup-time wiring lands; it replaces the freeze.
- M13, M14 — independent of everything else. M14 lands opportunistically; M13 should land with whichever composition layer wins the protocol question.
- M6-scientific, M9, M5-full — deferred; each needs a written justification.
Acceptance criterion for every step, demonstrated feasible by ICON-sc over 288 composed steps / 1440 dycore substeps: the old and new wiring agree bitwise, as a release blocker, never a tolerance to widen.
8. What this deliberately does not solve
- Granule-private scratch — the (plausibly) larger memory number — is fixed by a type (M10), not a container.
nlevvsnlev+1—fa.CellKField[float]cannot express it and the factory erasesKHalfDim → KDimat allocation (factory.py:540,545,823). A gt4py type-system gap; no container fixes it, M2’sdims+originmetadata only fences it.- Halo-exchange placement — needs declared access (PSyclone derives exchanges statically
from access mode × function space). Record
intentnow (one word per field); building the consumer is a separate project — and ICON-sc is the cautionary tale, having declared halo metadata and never built its consumer. - The prognostic double buffer — intentional and already optimal;
swapis a pointer rebind. - Bit-reproducibility across refactors is kept, not solved: laziness would make evaluation order data-dependent, which is one reason M5-full is deferred.
- Integration control state (
ndyn_substeps_var, CFL-watch mode, elapsed time, random seeds) is restart state but not fields; the container holds fields, and conflating the two is a scope error. - Module boundaries.
tach checkcurrently enforces nothing: it resolves zero first-party imports, the tach ≥ 0.27 namespace-package regression documented in the layered-architecture refactor’s Phase 0 (re-confirmed by the parallel verification in PR #18). Any argument of the form “the boundary check will stop a shared container from landing in the wrong package” is false today.
9. Restart: the consumer that settles G4 and G5
restart doc is the best-specified consumer of field metadata we have, and it decides two arguments.
State of play (verified): main has a serialbox-based read path
(initial_condition/from_file.py) restoring prognostics + exner_pr + the advective
tendencies; no write path; origin/ibm_02 has a serial pickle-based prototype writer.
- G5 comes straight from the restart inventory. What must be checkpointed is prognostics
plus the dycore diagnostics carried across steps (
exner_pr,ddt_vn_apc_pc,ddt_w_adv_pc) — while metrics, interpolation coefficients and compiled stencils must not be. “Is this restartable” is orthogonal to “which container is this in”; the prog/diag split cannot express it, andibm_02is consequently forced to hand-pick threeDiagnosticStateNonHydromembers by name — precisely the hand-written list ICON’s field groups abolish. - Exact vs scientific restart is a per-field decision — a label set chosen once, not a code path.
- Restart must reconstruct fields correctly in a fresh process, and metadata read off
a live field at write time cannot guarantee that.
ibm_02’s writer stores{data, dims}per field (a sixth independent reinvention of field metadata) and can rebuild arrays viagtx.as_field— but the reconstruction is incomplete and backend-bound: half-levelness never reaches the file because the factory already erased it at allocation (E8), and origin and dtype are not recorded. This is the declare-dims-once argument arriving from a second direction. - Tracer restart fails today for a naming reason, not a technical one — the savepoint
grouping, stated as a
NotImplementedError.
If restart ships first with a hand-picked field list, that list becomes the de-facto role
vocabulary. The restart doc has a 2-week appetite; this design does not. So the one
conversation worth having before that work starts is restart: bool in M2’s metadata —
even if nothing else here is adopted (open question 7).
10. Prior art, compressed
One-line verdicts; steal/avoid distilled. Line numbers were read from local checkouts for ICON, LFRic, gt4py and ICON-sc; the rest are paraphrase-grade.
| System | Container | Resolution | Verdict for icon4py |
|---|---|---|---|
| ICON-sc (egparedes’ prototype) | boundary dict → slotted vault at run time | bind-time, frozen execution plan | closest test of this document’s thesis: confirms it, and corrects it in three places (see the dedicated subsection below) |
| ICON (Fortran) | typed derived types + parallel add_var registry | run time, I/O only | steal metadata-at-definition-site + labels; never pay the dual declaration (~15k lines) |
| MPAS | string-keyed pools | run time | avoid — its own team deleted it for GPU; successor Omega dropped it |
| CCPP | host model’s, unchanged | build time, generated glue | the framework is not in the executable; but its #1 regret is the vocabulary, and forbidding shared derivation made 66 % of one suite interstitial glue |
| sympl / climt | dict[str, DataArray] | run time, per call | icon4py’s Component protocol descends from it; its maintainers abstracted the container away for performance |
| NDSL / Pace | rigid dataclasses + field metadata | setup | closest technical analogue — the other GT4Py model, and it stayed rigid; source of M2, M4 (~40 lines, no registry), M10’s Local, and QuantityFactory/GridSizer |
| ClimaAtmos | prognostic Y + cache p, one explicit refresh barrier | — | migrating back to explicit structs; source of the M5-lite barrier pattern |
| LFRic / PSyclone | field collections + global store | compile time | steal declared-access → derived halo exchange (the biggest structural prize, needs intent not a container); its retrospective rejects the global store |
CAM pbuf / Omega | runtime buffer | init-time index, run-time access | resolve names once into typed handles; never a string lookup in a kernel |
| WRF Registry | text table → generated types | build time | one declaration driving many services — as a bespoke DSL, the cautionary form of M4 |
| NUOPC | advertise → realize | init-time negotiation | unconnected exports cost zero memory (M11’s shape); errors on ambiguity, never guesses |
ICON-sc: an experimental architectural prototype, with its own model-state proposal
What it is. An experimental architectural prototype of a full alternative Python
architecture for ICON — sympl/Tasmania-lineage composition over a zero-copy device-field
boundary — hosting icon4py granules rather than forking them, built in six agent-driven
days (work units 001–014, 2026-07-08→13). Its model-state design is the fifth alternative in
§1: state crosses the public boundary as dict[str, DataArray] under sympl property
contracts; at bind time a compiler resolves all names and contracts once into a frozen
execution plan; what survives into the loop is a slotted, index-addressed StateVault
no component can reach (a test proves zero name lookups per step). A staleness guard —
epoch / generation / schema_hash counters plus a debug renegotiate-and-diff — replaces
a freeze; buffer adoption is the primary ingress path (from_state never allocates);
units are validated as identity, never converted; non-CF fields must carry the icon:
prefix (measured split: 18 CF / 72 icon:). Built against our exact problem on our
codebase, its evidence is the most directly transferable in this survey — and needs the
most calibration.
References (published; this supersedes the original appendix’s “nothing pushed to any
remote”): repository ·
documentation ·
architecture doc
(for model state: §2, state/contracts/naming, and §8.2, the negotiation/execution split) ·
plan/bind.py
(~1730 LOC, the plan compiler) vs
state/vault.py
(203 LOC, the vault) ·
REGISTRY.md ·
lock.toml
(the SHA-pinned provenance ledger — itself worth stealing as a process artifact). The
README’s validation claims (L2 parity at upstream tolerances; 9-day bitwise-zero vs the
icon4py driver; T0 ≡ T1 through the dycore) are self-reported. Calibration: zero GPU
execution ever, zero MPI, 2 of ~11 NWP schemes, no real data; its halo validator — the
most-quoted idea in its own architecture — is unbuilt; internal review discipline genuinely
strong, production contact none.
Lessons learned, in decreasing order of weight:
- Confirms the central thesis of §4 — “nothing about the interfaces changes during execution … every lookup performed in the loop is recomputing an invariant” — and demonstrated §7’s bitwise old≡new acceptance criterion over 288 composed steps.
- Corrects three earlier claims (folded into §4): the container may survive into the loop — the right test is component-unreachability, not non-existence; the performance stake is ~6.8 %, not orders of magnitude; and a static dataclass cannot be the public state type without conditional allocation (M11) — which ICON-sc itself lacks.
- The dict is what costs. The ~1730-line compiler exists chiefly to erase a dict its own interpreted tier introduced — 8.5× the size of the container it feeds. icon4py never has to introduce that dict; the strongest single argument for emitting typed dataclasses.
- It does not solve C2. It owns allocation (one buffer per contracted name — its answer to G1), so its two hosted granules copy ~17 full fields per Δt (~100 MB) in and out. An icon4py-native design must adopt buffers instead.
- Its transferable residue is small and already absorbed into §5: M6-structural, M12’s
arity check (plus closing its own publisher-count hole), M14, units-as-identity, the
icon:invariant,origin/K-domain as first-class metadata, and thelock.tomlprocess — ≈300 LOC out of ~3700 of compiler + tests. - Do not adopt its unbuilt or unused parts: the coupling algebra (7 combinators, 2 used); the F-tier/JAX lowering (a second physics implementation that abolishes component privacy); the halo story (declared everywhere, consumed nowhere); ping-pong SSA time levels (its own dycore opted out).
Steal (beyond the mechanisms already in §5 and the ICON-sc items above): units as
identity-validation with the conversion path quarantined (sympl’s per-call conversion is its
own documented performance regret; ICON’s post_op converts only at the file boundary);
capability-vs-request separation (vert_interp says how a field could be interpolated, the
namelist says whether); revision counters over content hashes for any invalidation (hashing
is O(field bytes) and its payoff never fires in floating-point dynamics).
Avoid: a run-time bucket reachable from compute code; silent lookup failure or auto-creation of unknown names; dual declaration kept in sync by hand; a general derivation planner and the opposite extreme of forbidding shared derivation; per-call unit conversion; assuming CF names cover model-internal fields (~80 % have none); adopting the unbuilt parts of any prototype.
11. Relation to the other proposals (updated 2026-08-10)
- msimberg’s revive-components:
the original conflict analysis targeted the v2 spec (
run(state, dtime),convert_state,setflags(write=False)read-only enforcement — the cupy objection applied to its AC14 and is moot for v3, which already downgrades read-only to best-effort). v3 supersedes all of that and moves the design to a graph-composition layer aboveComponent. Two readings of v3 were argued in the two parallel consolidations of this document, and both are partly right, so state the synthesis precisely. On the declaration side they compose: v3 owns control flow (chain/loop/when, schedules), this proposal owns allocation/wiring/metadata, andreg.build(...)emits exactly the caller-allocated containers v3’sCarrySpec(..., initial=…)expects. On the execution side v3’s D1/D2 is a run-time, name-addressed store: every step reads and writes a shared mutableCarry, the component adapters rebuildInputTfrom carry slots per call, andsamplerkeeps a name-keyed recycle cache — and v3 concedes the consequence itself (“read-only is a debugging aid … not a hard guarantee on a shared mutable carry”). Components never touch the carry directly (the adapters do), so this is not MPAS pools — but the reachability and emission tests of §4 apply to its executor, and the resolution is the bind-once treatment ICON-sc applied: resolve carry-slot → dataclass-field bindings once at setup and the per-step repacking disappears, dissolving the conflict. Three further points stand: v3 says nothing about who allocates, so the E1 class survives it intact; itsFlowKindconflatesrole(prognostic/tendency/diagnostic) withintent(in-place) with arity (parameter) — still, even after v3 resolved v2’s O1/O4/O5. The refinement this document recommends keepsintentin M2’s metadata and carriesroleas labels (M7), never as container membership (G5); and its whole-graphvalidate()and M13 are the same idea approached from two directions — build it once. - OngChia’s design is the only one
requiring a run-time container. Two specific problems: “each component derives its own
inputs” is the rule that produced E3 (CCPP ran the same experiment and got 66 % interstitial
glue); and
is_freshanswers “was this written this timestep”, not “is this consistent” — a hole the design documents itself (derived fields are not invalidated when their inputs change). Its per-component call frequency and Jacobi/Gauss-Seidel selection are genuinely covered nowhere else and should be kept — as driver/orchestration features, which is also where msimberg-v3’ssampleroverlaps them. - egparedes’ layered-architecture
refactor independently reaches the same duplication findings and proposes merging
PrepAdvectioninto a sharedModelState. Its-> Nonein-place contract is the most GPU-honest of the four. The conflict is real but narrower than “bucket” (revised, §4): a typedModelStatepassed whole violates C3 (call sites stop naming their inputs), not the string-lookup prohibition. Its Phase 6 explicitly depends on the protocol question this document leaves to its owner. - restart is a
consumer, not a competitor — see §9. One conversation (
restart:metadata) should precede its 2-week execution. - What no other proposal covers: allocation, scope/lifetime, conditional allocation, labels, halo intent, time-level rates, the vocabulary — and all of them assume CF standard names work, which they do not.
On the protocol question itself, the Working Principles are blunt: conceptual integrity comes
from one empowered designer, not committee negotiation — four drafts merged by consensus will
not produce a coherent protocol. The productive question is who decides, not whose design
is best. And events have partially decided it: PR 1301’s merge made the dict-based protocol
the incumbent on main. If the team wants a different protocol, that is now a migration, not
a green-field choice.
12. Open questions
- Is the Fortran-embedded path permanent? (C2). Currently held as a hard requirement; it shapes the entire adopt-don’t-own allocation design. Re-validate explicitly — if it falls, the design simplifies materially.
- Which temperature is the model temperature? When IO and physics disagree (E3), which one is written to output is a science decision nobody has signed off — today the published one is dry.
- Has standalone-driver tracer advection ever produced validated results? While E1 was live it ran on identically-zero mass fluxes; PR 1404 fixed the wiring, but who signs off that post-fix results are scientifically valid?
- Who owns the component protocol? Four signatures were open; PR 1301’s merge made the dict protocol the incumbent while the typed alternatives remain drafts. Someone with design authority must either ratify or supersede it — merging the drafts by negotiation will not converge.
- What per-timestep Python overhead is acceptable, as a number? Still open; with the datapoint that the entire stake measured ~6.8 % on a real model, whatever is built should not be justified on this axis.
- Do we commit to a controlled name vocabulary, and which domain scientist owns it? The
shape is settled (CF or
icon:-prefixed, two-way invariant, enforced at registration); the ownership is not, M4 is gated on it — and it is CCPP’s documented #1 regret. - Exact or scientific restart? The highest-leverage question in the list: it forces a per-field decision on every field in the model (G5), and the restart work’s 2-week appetite means its answer will be set de facto very soon (§9).
Bitwise reproducibility across the refactor?Answered: yes, and it is achievable (ICON-sc, 288 steps); adopt as a release blocker. It rules out lazy derivation, not setup-time derivation.A hard declare→bind→freeze→run lifecycle?Answered: no — a staleness guard beats a freeze (M6-structural); mutation stays legal and stale wiring raises.IsAnswered by PR 1404 as science: half levels (mass_flx_icon half or full levels?nlev+1), both branches of the fix now allocate accordingly — but the answer lives in an allocation call and a comment, not in any declaration; the type still cannot express it (§8), so the knowledge remains one refactor away from being lost.
Appendix: numbering maps for inbound references
Other documents cite the v1 requirement and open-question numbers (e.g. the layered-architecture refactor cites “model-state’s R11” and “its open question 3”). The maps, so those references survive if the originals are retired:
- Requirements → constraints/goals: R1→G2 · R2→G1 · R3→G3 · R4→C2 · R5→C1 · R6→G6 · R7→G7 · R8→G4 · R9→C3 · R10→C4 · R11→G5
- v1 open questions → here: 1→Q1 (Fortran path) · 2→Q3 (tracer validation) · 3→Q4
(protocol) · 4→Q10 (
mass_flx_ic) · 5→Q2 (temperature) · 6→Q8 (bitwise) · 7→Q5 (overhead) · 8→Q9 (freeze) · 9→Q6 (vocabulary) · 10→Q7 (restart)
Appendix: unrelated icon4py defects found along the way
Found while ICON-sc hosted icon4py granules; none is about state containers, all are actionable independently, and none has been filed upstream.
| ID | Finding | Status |
|---|---|---|
| U1 | Graupel cold-glaciation water-budget leak: supercooled qc at T ≲ 233 K near the moist-domain top gains total water, +1.59e-4 kg/m² per Δt=30 s, suppressed by any coexisting ice-phase seed | has a runnable, wrapper-free reproducer on public APIs |
| U9 | is_surface index bug in the graupel scan: k_lev is a carry relative to kstart_moist, compared against an absolute ground_level, so the surface clamps only fire when kstart_moist == 0 | verified independently; one line |
| U2 | wgtfacq_c/wgtfacq_e emitted on shifted K-domain [nlev−3, nlev), visible only in the factory registration | cost ICON-sc ~2 work units; same class as E8 |
| U3–U5 | Grid-factory: mean_cell_area off 4e-5 relative; RBF pentagon divide warnings; keep_skip_values=False + _replace_skip_values makes the RBF matrix exactly singular for file-sourced grids | latent trap for grid-from-file |
| U6 | SPECIFIC_HEAT_CAPACITY_ICE = 2108.0 vs ICON’s 2106.0_wp; live only in a non-default branch | latent, covered by no verification data |
| U7 | satad: ICON silently caps at maxiter; icon4py raises ConvergenceError | bites the first non-default configuration |
| U8 | The multi-substep dycore test is MCH-only (# why is this not run for APE?) | test-coverage gap |
| U11 | total_precipitation_flux computed only under do_latent_heat_nudging=True, else exact zeros | would mislead as a diagnostic |
| U12 | Every solve_nonhydro/diffusion integration test xfails on embedded; the diffusion granule cannot be constructed there | rules out “embedded as reference tier” for a wiring-equivalence harness |
U1 and U9 are filable today with evidence attached; U3–U5 file as one issue; U6/U7/U11 are one-liners; U8 is a test-coverage PR.
Related: Revive components, Physics driver and component design, restart, Layered architecture refactor, Working Principles.