Skip to content

Host CPU and network rates from kernel counters - #9

Merged
bencode merged 1 commit into
mainfrom
collect
Oct 4, 2026
Merged

bencode merged 1 commit into
mainfrom
collect

Conversation

@bencode

@bencode bencode commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Round 2 of richer collection, iteration 1 of 3: host CPU and network. These come only from kernel counters, under the constraint that collection must stay lightweight and must never affect the host.

What the agent reads

File For
/proc/stat (the cpu line, first 8 fields) busy (user+nice+system+irq+softirq), iowait and steal, as shares of all CPUs
/proc/net/dev bytes received and sent at the interfaces backed by a device (/sys/class/net/<if>/device), so bonds, bridges and container interfaces are not counted twice
/proc/uptime the clock for rates; it never steps
  • The agent reads nothing else: no child processes, no threads or sampling loops, no docker stats, no exec.

Safety

  • No panics on input. Every sum and product saturates, division happens only behind a non-zero guard, there is no indexing, and nothing is unwrapped from file contents. All parsers are pure and unit-tested, including malformed input.
  • Rates.
    • Readings are kept with their uptime. A reading that failed carries the previous one, so the next rate spans both intervals.
    • A rate is null on the first pass, across more than ten minutes, after a reboot (uptime went back), or when no interface is present in both readings.
    • A counter that went back (iowait, a NIC reset, a 32-bit wrap) counts as no time or is skipped. It never blanks the rate.
  • Isolation.
    • The new reads are separate from host::read, so load, memory and disks are unaffected.
    • A counter that cannot be read or parsed is logged (warning: host: /proc/…) and left null.
    • Such a failure is not added to Report.errors, which would make the server treat the host's endpoints and log hygiene as unobserved. This changes the design, following the review: a purely informational metric must not affect incident lifecycles.
  • Backstop. deploy/skym.service now has RestartSec=60. A restart is a full first pass (every container, an hour of events), so even a crash loop costs no more than normal passes.

Protocol

These fields are additive with serde(default), and there is no migration:

  • HostState: cpu_percent, iowait_percent, steal_percent, net_rx_bytes_per_s, net_tx_bytes_per_s.
  • HostOverview carries the same fields.
  • Older agents leave them null.

View

  • Hosts tab. The CPU column shows 4 · 23%; with no rate yet, only the core count.
  • Host preview and page.
    cpu      23% busy · 4% iowait · load 0.93   memory 10.3 / 16.4 GB
    net      ↓ 1.2MB/s  ↑ 540B/s
    
    steal is shown only from 1%.

Verification

  • 257 tests pass; clippy and fmt are clean.

  • cargo-mutants on the diff: 93 mutants, none missed. The real-system reads (counters, physical_interfaces, read_parsed) are excluded.

  • skym report --dry-run was run read-only, from /tmp, as the skym user, on sp (1 CPU), i (the busiest) and h129 (CentOS 8, cgroup v1). It measures over the second before the pass.

    Host top busy skym cpu_percent One pass
    sp 1.0% 1.0% ~0 s CPU, 4.7 MB max RSS
    i 2.4% 2.8% 0.04 s CPU, 7.7 MB
    h129 0% 1.5% ~0 s CPU, 6.6 MB

    Network values were of the order a manual read gives (different seconds). The test binaries were removed afterwards.

  • An independent review found no panic path. Fixed from it:

    • "no shared interface" now reads as unknown, not 0 B/s;
    • counter failures are logged instead of reported (above);
    • parse failures are logged too;
    • byte-level rates are readable;
    • the CPU column fits 128 cores;
    • docs wording.

Rollout

  1. The server on x.
  2. The agent and unit on sp first. Record CPUUsageNSec and MemoryCurrent before and after, and watch for 30 minutes.
  3. Then yi1, y, x, i and h129, one at a time, keeping skym.prev for rollback.
  4. skym-view.

Next: iteration 2 (container CPU on cgroup v2 and v1, plus container memory on v1), then iteration 3 (addresses, listening sockets, top processes), each with its own design and review.

The agent reads /proc/stat, /proc/net/dev and /proc/uptime once a pass and
reports, since the previous pass, how busy the CPUs were (busy, iowait and
steal as shares of all CPUs) and the bytes per second at the physical
interfaces. Rates span readings taken by uptime, so they hold across clock
steps and missed passes, and are null on a first pass, across more than ten
minutes, or without a physical interface. Nothing else is read; a counter
that cannot be read is logged and left null, without making the host's
subjects unobserved. Parsing never panics on input.

The unit restarts the agent at most once a minute: a restart is a full first
pass. skym report --dry-run measures the rates over the second before it.

The hosts tab shows CPUs and busy %; a host's preview and page add cpu
(busy, iowait, steal from 1%, load) and net lines.
@bencode
bencode merged commit 658570f into main Oct 4, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant