Profile reference¶
The complete schema of a profile file. A profile tells runhealth what a code
prints, and the core never hard-codes a model: everything code-specific is
defined here. For a minimal example and for how profiles are selected, start
with Profiles.
The file name is irrelevant; the name: field identifies the profile. The
bundled profiles in src/runhealth/profiles/ are worth reading as worked
examples: slurm.yaml, icon.yaml, cray-mpich.yaml.
Composition¶
Several profiles can apply to one log. Each contributes its own rules and
settings, in order of priority (lowest first, so that a later profile can
override a setting). The slurm profile has always: true and applies to
every log; every other profile is selected by its detect patterns, or pinned
with --profile a,b.
Concerns should be kept separate. The fabric counters printed by a Cray machine
are unrelated to the model running on it, which is why they are defined in
cray-mpich.yaml rather than in icon.yaml.
Skeleton¶
name: mymodel
description: "What this profile covers"
priority: 10 # applied after lower numbers
always: false # true only for a profile that must apply to every log
detect: [ ... ] # any match selects this profile
attempt_boundary: [ ] # patterns marking a new job attempt in the same file
settings: { ... } # names and column mappings the analysis uses
thresholds: { ... } # numbers the checks compare against
fields: { ... } # one value per run
keyvalues: { ... } # many key/value pairs into one dictionary
series: { ... } # repeated events, in order
markers: { ... } # phase boundaries on the timeline
groups: { ... } # very frequent messages, counted rather than stored
tables: { ... } # structured blocks
outcome: [ ... ] # the run's verdict
Rules¶
Every entry in fields, keyvalues, series, markers, groups and
outcome is a rule. A rule is either a bare regex string or a mapping:
model_version:
re: '^\s*version: (\S+)' # required (tables use `start:` instead)
contains: 'version:' # literal prefilter, see below
ignorecase: false
preamble: false # see "The job script preamble"
Patterns are Python regular expressions, matched with re.search against the
line after the timestamp and rank prefix have been removed. Anchor a pattern
with ^ to refer to the start of the message.
contains: prefiltering with a literal¶
Rules are applied to every line of the file, and a file can hold a million
lines. Before running a regular expression, runhealth checks whether a
literal substring is present, which is much cheaper. It derives that literal
from the pattern itself and falls back to always running the regular expression
whenever the pattern contains an alternation or a quantifier that could make
the literal optional.
Specify contains: explicitly when the derivation fails but a literal is known
to appear. For a pattern such as 'Constructing the .* coupling frame', the
hint coupling frame reduces the parsing time considerably. The literal must
appear in every line the pattern matches, otherwise those lines are
silently skipped.
Detection¶
detect:
- 'master_control: start model initialization'
- 'Timer report, ranks'
- '# ICON run script'
Any match within the first 8 MB of the log selects the profile. Include a pattern that survives an early crash: a job that dies during MPI startup prints none of the model’s own output, but the run script echoed at the top of the file is still present.
The job script preamble¶
Many run scripts echo a copy of themselves before their output begins to be
timestamped. That copy contains every string the script can ever print,
including both its success and its failure message, so rules must not match
there. In a timestamped log, runhealth treats everything before the first
timestamp as preamble and applies only rules marked preamble: true to it,
which is how #SBATCH directives are read.
Two consequences for profile authors:
Anchor an
outcomepattern with^so that it matches the printed line and not theecho "..."statement that produced it.Mark a rule
preamble: truewhen it should match only within the job script.
Sections¶
fields: one value per run¶
fields:
model_version: {re: '^\s*version: (\S+)', contains: 'version:'}
node_count: {re: '^SLURM_JOB_NUM_NODES=(\d+)', cast: int}
final_status: {re: '^Status: (.+)', keep: last}
group picks the capture group (default 1), cast is int or float, and
keep is first (default) or last. A float also accepts what Fortran
prints: a D exponent, an exponent of three digits without its E
(0.22+279), and a field of asterisks, which is read as infinity.
One name is conventional: a field called <component>_ranks, cast to int,
declares how many ranks a component of a coupled model was given. The report
lists them, and the coupling check uses them to assign each timer table to a
component, assuming that the components occupy consecutive blocks of ranks in
the order in which the log announces them.
keyvalues: a dictionary¶
The pattern must capture two groups: the key and the value.
keyvalues:
sbatch:
re: '^\s*#SBATCH\s+--([A-Za-z0-9_-]+)=?\s*(.*)$'
contains: '#SBATCH'
preamble: true
slurm_env:
re: '^(SLURM[A-Z_0-9]*)=(.*)$'
contains: SLURM
The sbatch entry has a special role: its time value is read as the
requested wall-clock limit, on which both the wall-time check and the
attempt-boundary detection rely.
With block: true, the pattern only marks the first line of a block of
key: value lines, which are all collected from there on. A key without a
value heads the lines indented below it, and the nested keys are joined with
/. The block ends at the first line of another shape, at a line indented
less than the first one, or at a key containing an underscore, so a message
such as master_control: ... is not taken for part of it. Lines from other
ranks do not interrupt it. Only the first block of a run is kept.
keyvalues:
build:
re: '^\s*executable: '
contains: 'executable:'
block: true
turns
executable: /path/to/bin/icon
revision: icon-2026.04-46-gfe40584-dirty
model components:
ICON-Land:
revision: icon-land-2026.04-4-g61aadeb
into executable, revision and model components/ICON-Land/revision.
series: repeated events¶
series:
timestep:
re: 'Time step:\s+(\d+) model time (\d\S* \S+)'
contains: 'Time step:'
fields: [step, model_time]
cast: {step: int}
role: progress
fields names the capture groups. Each event also carries the wall clock of its
line.
role marks what the series is for:
progress: the throughput signal. Wall time between consecutive events gives the progress rate; if a captured field named bysettings.model_time_fieldparses as a date, the simulated-time rate is computed as well.io: output and checkpoint events, marked on the timeline. The wall-clock gap between successive events of the same series is also tracked: an outlier gap raises a cadence check, and the “Output cadence” figure plots every series over the run.stability: values a model reports about its own numerical state, such as wind maxima or CFL numbers. They feed the “Numerical stability” check and one figure perfiguregroup.
A stability series takes four more keys:
series:
max_w:
re: 'MAXABS VN, W .* +(\S+?) at level +(\d+),$'
fields: [w, level]
cast: {w: float, level: int}
role: stability
peak: w # the field that is judged and plotted
label: vertical wind |w|
unit: m/s
figure: Maximum wind speed
With peak, the series is judged against the thresholds <series>_warn and
<series>_fail (here max_w_warn, max_w_fail). A value that is not finite
always fails. The other fields of the record, such as the level, are quoted as
evidence. A series with peak is never cut off: once it holds 20,000 records,
neighboring pairs are merged and each keeps its larger value, so a series
reported on every substep still covers the whole run. The first record at or
beyond the _fail limit and the record before it are kept exactly.
Without peak, a stability series is counted as events, listed in the check,
and drawn as dots along the bottom of the figure named by its figure.
A series without peak keeps its first 50,000 records and ignores the rest.
A progress series that may run longer takes bin: last: once it holds 20,000
records, neighboring pairs are merged into the later one, which counts the
reports it stands for and keeps the slowest of them. The first report is never
merged, and the step, the model time and the total wall time stay exact, so
the throughput is unaffected. The slow-interval check still finds every bin
with a slow report and names its step, but counts a bin with several slow
reports once.
markers: phases¶
markers:
model_init:
re: 'master_control: start model initialization'
label: model init
first_step:
re: 'Time step:\s+1 model time'
label: time loop
Markers divide the wall clock into phases. The label names the interval that
starts at the marker, not the marker itself, which is why first_step is
labeled time loop. Only the first occurrence is kept unless all: true is
set.
groups: messages that must not be stored¶
A file can contain hundreds of thousands of copies of a single warning.
groups counts them instead of storing them.
groups:
icon_warning:
re: 'WARNING PE:\s+(\d+)\s+(\S+)'
contains: 'WARNING PE:'
key: 2 # capture group to group by (0 = the whole match)
node: 5 # capture group holding a node name, if any
label: ICON warnings
claim: true # keep these lines out of the generic error scan
Counts are kept per key, per node and per minute of the run. claim: true
tells the core that this family is already accounted for, so that the same
lines are not additionally collected as anonymous errors.
tables: structured blocks¶
Two shapes are supported.
ruled: columns delimited by a rule row of dash sequences, with a header
row beside it. The rule row defines exact column spans, so a name containing
spaces never has to be guessed.
tables:
timers:
start: 'Timer report, ranks (\d+)-(\d+)'
contains: 'Timer report, ranks'
kind: ruled
nesting: '^(\s*)L ' # indentation to tree depth
durations: [t_min, t_avg, t_max] # columns holding 16m43s-style values
trailing-numbers: no rule rows; each row is a label followed by a fixed
number of numeric fields, so a label containing spaces still works.
cxi_counters:
start: '^MPICH Slingshot CXI Counter Summary:'
kind: trailing-numbers
cxi_ratios:
start: '^Computed Ratios'
kind: trailing-numbers
include_start: true # the start line is itself the header row
The block ends at a closing rule row or a blank line. A table opened by a rank-labeled line accepts only lines from that same rank, so interleaved output from other ranks does not corrupt it.
outcome: the verdict¶
outcome:
- {re: '^Script FAILED: (.+)', contains: 'Script FAILED', level: fail}
- {re: '^Script run successfully: (.+)', level: ok}
level is ok, fail or info. The last match in the file takes
precedence, since the actual verdict is written at the end. Capture group 1
becomes the text shown in the report.
attempt_boundary¶
attempt_boundary:
- '^end_of_job_script'
A resubmitted job can append to the same file. runhealth already detects this
from a sequence of unstamped lines, or from a pause longer than the scheduler
could have allowed, and analyzes only the last attempt. A profile can add a
precise marker of its own. Boundary rules are ignored until a sufficient part
of an attempt has been seen, so the first preamble never triggers one.
settings¶
Names and column mappings used by the analysis. All entries are optional.
Key |
Meaning |
|---|---|
|
what the progress unit is called, e.g. |
|
what the rate is called, e.g. |
|
which series is the progress signal (otherwise the first with |
|
the field in that series holding a simulated date |
|
which table holds the timers |
|
the label of its overall timer, e.g. |
|
map of |
|
timer labels that are output rather than computation |
|
timer labels that cover the coupler as a whole, reported for context |
|
blocking gets of the time loop, on which a component waits until its partner delivers |
|
coupler setup, including the very first get; reported, but not counted as waiting |
|
pattern matching the first and last rank in the title of a timer table, which assigns each table to a component |
|
the |
|
keys of that block that describe the run rather than the binary, such as its date or host; neither shown nor compared |
|
network counter blocks |
|
map of |
|
counter names worth reporting |
|
message groups that indicate fabric congestion |
|
the keys within those groups that actually indicate congestion |
thresholds¶
Every number a check compares against. The defaults are defined in
slurm.yaml; a profile or a site can override any of them.
Key |
Default |
Meaning |
|---|---|---|
|
300 |
silence inside the main loop that fails a run |
|
1200 |
silence during setup that warns |
|
5 |
in-loop pause, as a multiple of the typical interval |
|
0.9 |
fraction of the requested limit that warns |
|
3 |
progress interval counted as an outlier |
|
1 |
leading progress intervals treated as warm-up, excluded from the steady-state rate, the outlier count and the plot scale |
|
4 |
gap between |
|
0.1 |
|
|
1.25 |
ratio of slowest to fastest rank that is worth reporting |
|
2.0 |
ratio of slowest to fastest rank that is considered severe |
|
0.2 |
slowdown between first and last quarter |
|
0.05 |
ignore timers below this share of the run |
|
0.15 |
share of a component’s run that all of its ranks spent waiting for the partner, worth reporting |
|
2.0 |
how much larger that share must be than the partner’s to say which component waits for which |
|
1000 |
size of a message family that warrants a warning, provided that |
|
0.2 |
fraction of the whole log that such a family must also account for |
|
0.25 |
one node’s share of a family that makes it suspect |
|
10000 |
recovered Slingshot network timeouts that warn; any nonzero count below is reported for information, or warns if the run failed |
|
100000 |
network timeouts that fail a run |
|
none |
limits for a |
Testing a profile¶
runhealth mylog.out --profile mymodel -o /tmp/check --no-plots
runhealth --list-profiles
If a rule never matches, the usual causes are a contains: literal that does
not appear in every matching line, a pattern anchored with ^ that is in fact
indented, or a rule that should have been marked preamble: true.