Usage¶
Recipes¶
Task |
Command |
|---|---|
Analyze everything in a directory |
|
Analyze logs on a remote cluster from a local machine |
|
Analyze a single log and open the report |
|
Read it in a browser from a login node |
|
Only the most recent logs |
|
Follow a job that is running now |
|
Write a Markdown report instead |
|
A fast terminal summary only |
|
See which logs would be read |
|
Logs with an unusual file name |
|
Ignore the model-specific profile |
|
Use your own profile |
|
Everything is written under -o, and a repeated run over the same directory
reuses the cache, so it takes about a second. Every flag is listed in the
command line reference.
Remote logs¶
The recommended way to run runhealth is from a local machine, pointing at
logs that reside on a cluster:
runhealth santis:/scratch/e1000/run -o report/ --open
Any path argument written as host:/path (or user@host:/path, or a path
relative to the remote home directory, host:logs) is treated as remote,
exactly as scp and rsync interpret it. This assumes that ssh host already
works without a prompt, since runhealth calls rsync over that same
connection to copy the matching logs into <outdir>/.remote-cache/ before
reading them. Only files matching the active glob are transferred, and only
from that one directory, not from its subdirectories. The one exception is a
diagnostics directory
diag.*/ beside the logs, which is transferred whole so that its checks
appear in a locally built report as well.
The result is a report on the local disk, so --open displays it immediately,
without port forwarding or a combination of --serve and ssh -L. It also
combines with --watch: each pass synchronizes first, so a local report
continues to follow a job that is still writing its log on the cluster.
A local path and a remote one can be mixed freely in the same invocation:
runhealth santis:/scratch/e1000/run ./local-logs -o report/
Each run page still names the original host:/path, not the local copy under
.remote-cache/ that was actually parsed, so a locally built report still
states exactly where on the cluster a log resides.
Running jobs¶
A log with no final status is reported as RUNNING while it is still being written, and as INCOMPLETE once it has gone quiet. The silence check then says where it stopped.
When squeue is available, runhealth asks it for the actual state, so a
queued job appears as QUEUED rather than as a broken run, and a job that
the scheduler still believes to be running while its log has gone quiet is
reported as STALLED, the one case in which intervention can still help.
--no-squeue disables this.
When sacct is available, runhealth also reads the SLURM accounting record
of every job, in one call per machine: locally for logs read on this machine,
and through ssh host sacct for logs synced from host. The record
settles the verdict of a log that stops without one (TIMEOUT, NODE_FAIL,
OUT_OF_MEMORY, CANCELLED by <user>), supplies the start and the duration
of a log without timestamps, and reveals a job that printed its success
message but still ended with a nonzero exit code. squeue, in contrast, is
only asked about logs read on this machine.
--no-sacct disables this, and --no-squeue disables both queries.
runhealth /path/to/logs --watch 60 -o report/
re-renders at a fixed interval. Each pass resumes where the previous one stopped, so following a growing 100 MB log costs no more than the new lines.
Output formats¶
--format htmlThe default.
index.htmlplus a page per run, with light and dark themes, interactive figures and a real print stylesheet. The browser’s Print to PDF produces a well-formatted document with sensible page breaks. The figures are written into the pages, so a single.htmlfile is a complete report that can be attached to an email.--format mdreport.mdin GitHub-flavored Markdown. Markdown cannot hold an inline figure, so the figures are written toimages/*.svgbeside it.--format pdfUses WeasyPrint if it is installed (
uv sync --extra pdf). Without it,runhealthwrites the HTML and asks for it to be printed instead of failing. WeasyPrint is not a hard dependency because it is large and printing from a browser is equally suitable.
Performance¶
Reading is strictly streaming, line by line, using counters and bounded shortlists instead of stored lines. A 152 MB log with 820k lines is parsed in about 12 seconds; 443 MB across 31 logs takes about 30 seconds in parallel, with a peak resident memory in the tens of megabytes.
Results are cached under <outdir>/.cache/, keyed on file size and
modification time, so a second pass over the same directory takes well under a
second. --no-cache disables the cache.
Following a long report¶
A pass over several hundred megabytes takes a while, so every stage reports what it is doing. On a terminal the three stages that run once per log draw a bar; parsing counts the megabytes read, and therefore advances within a single large file rather than only when the file is finished:
runhealth: scanning /scratch/e1000/run
runhealth: 11 log(s) to read, 443.2 MB in total
runhealth: parsing on 8 core(s) ━━━━━━━━━━━━╸──────────── 194/443 MB 13s, 17s left LOG.831673.o
When the output is not a terminal, for example in a job script or a pipe, the bars are left out and each stage prints one line as it finishes, so the log stays readable.