Skip to content

What a host uses beyond its workloads - #13

Merged
bencode merged 1 commit into
mainfrom
other
Oct 5, 2026
Merged

bencode merged 1 commit into
mainfrom
other

Conversation

@bencode

@bencode bencode commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Answers "is the load in the workloads or outside them?" from data skym already has. This replaces iteration 3's per-process top, which was dropped: no new collection, and agents are unchanged. It is most useful to an AI agent working through the API, which usually can't run top on the host.

What

HostOverview gains two fields:

  • other_cpu_cores: busy CPUs (cpu_percent / 100 × cpu_count) less the cpu_cores of the host's running workloads.
  • other_memory_bytes: used memory less their memory_used_bytes.

Rules:

  • Both are clamped at zero, because the readings are taken moments apart.
  • Only workloads in the host's latest report count. A workload that is gone keeps its last state and readings until it is archived after 7 days, and would otherwise understate the figure. A workload is counted when its last_seen is not before the host's.
  • A workload without a reading (a container's first minute) counts as none.
  • CPU is null until the host has a rate (the agent's first pass).
  • Memory is a lower bound. Workloads' memory includes page cache that the host's MemTotal − MemAvailable treats as available; for systemd units, all of it. The docs and the skill say so.

The view shows other 0.80 cores · 1.2G beyond the workloads after top cpu on a host's preview and page.

From production data (current API, same arithmetic)

Host Busy Workloads Other
y 0.59 cores 0.07 0.53: most of y's CPU is outside its workloads
i 0.36 0.15 0.21
yi1 0.11 0.07 0.04

Other memory is 0.5–1.7 G per host, which is the kernel/dockerd baseline. On x, a workload that no longer exists but is still listed would have blanked the figure under the first version of the rule. The review then showed that its stale readings must be left out entirely, and they are.

Verification

  • 267 tests pass; clippy and fmt are clean. A server test covers:
    • running versus stopped workloads;
    • another host's apps;
    • gone versus still-listed workloads;
    • no reading;
    • the host's first pass;
    • clamping;
    • a host that never reported.
  • cargo-mutants on the diff: 32 mutants, none missed.
  • The independent review found the stale-workload gap (fixed as above) and the memory bias (documented as a lower bound); CPU is described as "other processes, the kernel".

Upgrade

Server and skym-view only.

HostOverview gains other_cpu_cores and other_memory_bytes: the host's busy
CPUs and used memory less what the running workloads of its latest report
use, never below zero. A workload without a reading (its first minute)
counts as none; one gone that the server still lists until it is archived is
left out, since it keeps its last readings. CPU is null until the host has a
rate. Memory is a lower bound: workloads' page cache is not in the host's
used memory.

The host preview and page show it after the largest CPU users; the skill
tells an agent to point at the host when it is most of the use.
@bencode
bencode merged commit d014206 into main Oct 5, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant