Files
moonweb-site/infra/monitoring/index.njk
T

58 lines
2.7 KiB
Plaintext

---
title: "Monitoring"
section: "infra"
tags: "infra"
parent: "/infra/"
description: "Prometheus and Grafana monitoring stack for hosts, containers, and hardware metrics."
layout: base.njk
---
<h1>Monitoring</h1>
<div class="detail-content">
<p>Prometheus and Grafana form the central monitoring stack for the whole
homelab, running as Docker containers on the NAS. Every machine in the
flat, including the Proxmox host, Docker VM, Synology NAS, and Raspberry
Pi devices, feeds metrics into this central instance.</p>
<h2>What's monitored</h2>
<ul>
<li><strong>Proxmox host metrics</strong>, a PVE exporter reports node-level and per-VM/LXC CPU, memory, and disk metrics into a dedicated Grafana dashboard.</li>
<li><strong>Host metrics everywhere</strong>, a node exporter runs on the NAS, the Docker host, and every Raspberry Pi, feeding basic CPU/RAM/disk/network stats.</li>
<li><strong>Container metrics</strong>, cAdvisor runs on both the Docker host and the NAS, breaking down resource usage per container.</li>
<li><strong>Network hardware via SNMP</strong>, an SNMP exporter polls a network-attached printer for toner level, page count, and status, proving the same pattern works for any SNMP-capable device.</li>
</ul>
<h2>Dashboards</h2>
<table>
<tr><th>Dashboard</th><th>Focus</th></tr>
<tr><td>Proxmox monitoring (extended)</td><td>Node-level and per-guest VM/LXC metrics</td></tr>
<tr><td>Docker container monitoring</td><td>Per-container CPU/memory/network/disk</td></tr>
<tr><td>Network printer</td><td>Toner level, page count, online/offline status via SNMP</td></tr>
</table>
<h2>How it's deployed</h2>
<p>Prometheus runs as a Docker container on the Synology NAS with a
bind-mounted data directory for persistence. Grafana is deployed as a
sibling container, connected to Prometheus as its primary datasource.
Both containers are defined in a single Docker Compose file and managed
via Portainer. The Prometheus configuration file defines scrape targets,
scrape intervals (15 seconds for hosts, 60 seconds for network devices),
and retention policies.</p>
<h2>Alerting</h2>
<p>Basic alerting is configured through Prometheus Alertmanager, which
sends notifications to a dedicated Telegram channel for critical events
like host-down or disk-full conditions. The dead-man's-switch pattern
ensures that if Prometheus itself stops scraping, an alert fires within
minutes. More sophisticated alerting rules (disk space prediction,
temperature thresholds) are planned but not yet implemented.</p>
<h2>Open items</h2>
<ul>
<li>Alerting rules are not yet configured for most services.</li>
<li>The secondary (test) Proxmox host isn't monitored yet, it's usually powered off.</li>
<li>Retention policy for Prometheus data hasn't been tuned.</li>
</ul>
</div>