{"id":4895,"date":"2026-10-01T11:02:12","date_gmt":"2026-10-01T05:32:12","guid":{"rendered":"https:\/\/cpanelfree.com\/blog\/setting-up-prometheus-and-grafana-for-server-monitoring\/"},"modified":"2026-10-01T11:02:12","modified_gmt":"2026-10-01T05:32:12","slug":"setting-up-prometheus-and-grafana-for-server-monitoring","status":"publish","type":"post","link":"https:\/\/cpanelfree.com\/blog\/setting-up-prometheus-and-grafana-for-server-monitoring\/","title":{"rendered":"Setting Up Prometheus and Grafana for Server Monitoring"},"content":{"rendered":"<p>Operating enterprise Linux fleets without high-fidelity observability turns minor kernel page cache bottlenecks and transient I\/O saturation into catastrophic system downtime. Whether staging container workloads on <a href=\"https:\/\/cpanelfree.com\">CpanelFree<\/a> or managing multi-node bare-metal infrastructure, relying on passive logs or intermittent polling fails to capture sub-second operational anomalies. Implementing a dedicated pull-based Prometheus time-series database coupled with dynamic Grafana dashboards establishes immutable telemetry and real-time visibility across CPU, memory, filesystem, and network sub-systems.<\/p>\n<p><!-- more --><\/p>\n<h2>What Is the Prometheus and Grafana Monitoring Stack?<\/h2>\n<div class=\"wp-block-group\" style=\"background:#f9f9f9;border:1px solid #e7e7e7;border-left:4px solid #001b41;border-radius:4px;padding:16px 20px;margin:20px 0\">\n<p style=\"margin:0;font-size:15px;line-height:1.6;color:#333\"><strong>Quick Summary:<\/strong> A Prometheus and Grafana server monitoring setup pairs Prometheus\u2014an open-source time-series database utilizing a pull-based HTTP scrape architecture\u2014with Node Exporter for Linux kernel metrics and Grafana for real-time visualization. This decoupled architecture ingests raw host telemetry, evaluates PromQL threshold rules, and delivers actionable alerts with minimal system overhead.<\/p>\n<\/div>\n<p>In modern systems administration, infrastructure observability is divided into three distinct pillars: metrics, logs, and traces. While logs provide post-mortem context, time-series metrics deliver proactive state verification. The combination of Prometheus, Node Exporter, and Grafana represents the gold standard for metric-driven monitoring due to its operational simplicity, pull-based scrape model, and multi-dimensional data model identified by metric names and key-value label pairs.<\/p>\n<h3>Core Architecture Components<\/h3>\n<ul style=\"color:#444;line-height:1.8;font-size:15px\">\n<li><strong>Prometheus Server:<\/strong> Acts as the scraping engine, TSDB (Time Series Database) storage engine, and query processing hub via PromQL. It periodically pulls HTTP endpoints exposed by target exporters.<\/li>\n<li><strong>Node Exporter:<\/strong> A lightweight binary written in Go that runs as a system daemon on target Linux hosts, collecting kernel statistics from <code>\/proc<\/code> and <code>\/sys<\/code>.<\/li>\n<li><strong>Alertmanager:<\/strong> Handles alerts emitted by Prometheus, deduping, grouping, and routing notifications to channels like PagerDuty, Slack, or webhook endpoints.<\/li>\n<li><strong>Grafana:<\/strong> The presentation layer that queries Prometheus via PromQL and converts multi-dimensional matrices into dynamic visual panels, heatmaps, and executive dashboards.<\/li>\n<\/ul>\n<blockquote class=\"wp-block-quote\" style=\"background:#f9f9f9;border-left:4px solid #001b41;padding:16px 20px;margin:24px 0\">\n<p><strong style=\"color:#001b41\">Architecture Note:<\/strong> Unlike push-based agents (e.g., traditional Zabbix or legacy Graphite) that bombard monitoring hosts with unthrottled packets, Prometheus controls collection cadence through scheduled pull cycles. If an edge node experiences network degradation or CPU starvation, the central Prometheus server regulates its own scrape rate, preventing cascading ingestion collapse.<\/p>\n<\/blockquote>\n<h2>Production vs Default Architecture Benchmarks<\/h2>\n<p>Deploying Prometheus and Grafana using stock repository configurations often leads to runaway disk space consumption, excessive kernel context switching from extraneous collectors, and memory exhaustion during large query evaluations. Tuning TSDB retention flags, memory bounds, and collector parameters achieves maximum efficiency.<\/p>\n<figure class=\"wp-block-table is-style-regular\">\n<table style=\"width:100%;border-collapse:collapse;margin:24px 0;font-size:15px;text-align:left\">\n<thead style=\"background:#001b41;color:#ffffff\">\n<tr>\n<th style=\"padding:12px 16px;border-bottom:2px solid #001b41\">Feature \/ Metric<\/th>\n<th style=\"padding:12px 16px;border-bottom:2px solid #001b41\">Standard \/ Default<\/th>\n<th style=\"padding:12px 16px;border-bottom:2px solid #001b41\">Tuned \/ Production<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Scrape Interval &amp; Resolution<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">15s uniform scrape interval<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">Tiered (15s core \/ 60s storage &amp; low priority)<\/td>\n<\/tr>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">TSDB Retention Strategy<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">15 days, unconstrained disk size<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">30d time cap + 85% storage volume max-bytes cap<\/td>\n<\/tr>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Prometheus Memory Footprint<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Unbounded heap allocation<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">GOMEMLIMIT capped at 80% RAM + memory limits<\/td>\n<\/tr>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Node Exporter Collector Overhead<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">All collectors enabled (ARP, bcache, infiniband)<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">Filtered collectors, &lt;0.8% single core overhead<\/td>\n<\/tr>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Alert Evaluation Latency<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Unscheduled polling \/ 1m jitter<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">Strict 15s evaluation cycle with duration timers<\/td>\n<\/tr>\n<tr>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Grafana Query Caching<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7\">Direct TSDB query pass-through<\/td>\n<td style=\"padding:12px 16px;border-bottom:1px solid #e7e7e7;color:#20B038;font-weight:600\">Query caching + recorded rule series<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<h2>Step 1: Installing and Hardening Node Exporter on Target Hosts<\/h2>\n<p>To extract granular hardware and operating system metrics from Linux hosts, Node Exporter must run as an isolated, unprivileged system daemon. Running monitoring agents under root creates an unnecessary security vector. We isolate Node Exporter into its own dedicated system user without shell access.<\/p>\n<p>Download the latest stable Node Exporter release from the official Prometheus repositories and install the binary into <code>\/usr\/local\/bin\/<\/code>:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code># Create unprivileged system service user\nsudo useradd --no-create-home --shell \/bin\/false node_exporter\n\n# Download and unpack official binary\nNODE_VERSION=\"1.8.2\"\ncd \/tmp\ncurl -LO \"https:\/\/github.com\/prometheus\/node_exporter\/releases\/download\/v${NODE_VERSION}\/node_exporter-${NODE_VERSION}.linux-amd64.tar.gz\"\ntar -xvf \"node_exporter-${NODE_VERSION}.linux-amd64.tar.gz\"\nsudo cp \"node_exporter-${NODE_VERSION}.linux-amd64\/node_exporter\" \/usr\/local\/bin\/\nsudo chown node_exporter:node_exporter \/usr\/local\/bin\/node_exporter\nrm -rf node_exporter*<\/code><\/pre>\n<p>Next, define a hardened systemd unit file at <code>\/etc\/systemd\/system\/node_exporter.service<\/code>. In production, disable noisy or unneeded collectors (such as bcache, fibrechannel, infiniband, and xfs) to preserve kernel cycles and prevent metric explosion.<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code>[Unit]\nDescription=Node Exporter Hardware and OS Metrics Daemon\nDocumentation=https:\/\/prometheus.io\/docs\/guides\/node-exporter\/\nAfter=network.target\n\n[Service]\nUser=node_exporter\nGroup=node_exporter\nType=simple\nExecStart=\/usr\/local\/bin\/node_exporter \\\n  --collector.disable-defaults \\\n  --collector.cpu \\\n  --collector.diskstats \\\n  --collector.filesystem \\\n  --collector.loadavg \\\n  --collector.meminfo \\\n  --collector.netdev \\\n  --collector.stat \\\n  --collector.time \\\n  --collector.uname \\\n  --collector.vmstat \\\n  --collector.filesystem.mount-points-exclude=\"^\/(sys|proc|dev|host|etc)($|\/)\" \\\n  --web.listen-address=\"0.0.0.0:9100\"\n\n# Security Hardening Flags\nProtectSystem=strict\nProtectHome=true\nNoNewPrivileges=true\nPrivateTmp=true\nProtectKernelTunables=true\nProtectControlGroups=true\nCapabilityBoundingSet=\n\nRestart=always\nRestartSec=5s\n\n[Install]\nWantedBy=multi-user.target<\/code><\/pre>\n<p>Enable and start the service, then verify that the raw metrics endpoint is serving data:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code>sudo systemctl daemon-reload\nsudo systemctl enable --now node_exporter\ncurl -s http:\/\/localhost:9100\/metrics | head -n 20<\/code><\/pre>\n<blockquote class=\"wp-block-quote\" style=\"background:#f9f9f9;border-left:4px solid #001b41;padding:16px 20px;margin:24px 0\">\n<p><strong style=\"color:#001b41\">Security Hardening:<\/strong> Never expose port 9100 directly to the public internet without mutual TLS (mTLS) or network firewall isolation. Restrict inbound traffic on port 9100 via <code>iptables<\/code>, <code>nftables<\/code>, or cloud security groups strictly to the IP address of your central Prometheus server.<\/p>\n<\/blockquote>\n<h2>Step 2: Installing and Configuring Prometheus Core<\/h2>\n<p>The Prometheus core server handles metric ingestion, TSDB chunk storage, and evaluation loops. Create a dedicated user, establish required directory hierarchies, and assign restrictive ownership permissions.<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code># Create unprivileged system user and directories\nsudo useradd --no-create-home --shell \/bin\/false prometheus\nsudo mkdir -p \/etc\/prometheus \/etc\/prometheus\/rules \/var\/lib\/prometheus\n\n# Download and install Prometheus binary\nPROM_VERSION=\"2.54.1\"\ncd \/tmp\ncurl -LO \"https:\/\/github.com\/prometheus\/prometheus\/releases\/download\/v${PROM_VERSION}\/prometheus-${PROM_VERSION}.linux-amd64.tar.gz\"\ntar -xvf \"prometheus-${PROM_VERSION}.linux-amd64.tar.gz\"\ncd \"prometheus-${PROM_VERSION}.linux-amd64\"\nsudo cp prometheus promtool \/usr\/local\/bin\/\nsudo cp -r consoles console_libraries \/etc\/prometheus\/\nsudo chown -R prometheus:prometheus \/etc\/prometheus \/var\/lib\/prometheus\nsudo chown prometheus:prometheus \/usr\/local\/bin\/prometheus \/usr\/local\/bin\/promtool\nrm -rf \/tmp\/prometheus*<\/code><\/pre>\n<h3>Production Prometheus Configuration (`\/etc\/prometheus\/prometheus.yml`)<\/h3>\n<p>The primary configuration file governs scrape intervals, rule evaluation intervals, alerting rules, and scrape targets. Create <code>\/etc\/prometheus\/prometheus.yml<\/code> with the following production-hardened specification:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code>global:\n  scrape_interval: 15s\n  evaluation_interval: 15s\n  scrape_timeout: 10s\n  external_labels:\n    cluster: 'production-primary'\n    datacenter: 'in-west-01'\n\nrule_files:\n  - \"\/etc\/prometheus\/rules\/*.yml\"\n\nalerting:\n  alertmanagers:\n    - static_configs:\n        - targets:\n            - '127.0.0.1:9093'\n\nscrape_configs:\n  # Internal Prometheus self-monitoring\n  - job_name: 'prometheus'\n    metrics_path: '\/metrics'\n    static_configs:\n      - targets: ['127.0.0.1:9090']\n\n  # Fleet Linux Nodes (Node Exporter)\n  - job_name: 'node_exporter'\n    scrape_interval: 15s\n    static_configs:\n      - targets:\n          - '127.0.0.1:9100'\n          - '10.0.1.15:9100'\n          - '10.0.1.16:9100'\n    relabel_configs:\n      - source_labels: [__address__]\n        regex: '(.*):9100'\n        target_label: instance\n        replacement: '${1}'<\/code><\/pre>\n<h3>Systemd Service Unit with Memory Sandboxing<\/h3>\n<p>To prevent Prometheus from triggering the Linux Out-Of-Memory (OOM) killer during heavy range queries, tune Go runtime memory allocation with <code>GOMEMLIMIT<\/code> and enforce systemd resource limits at <code>\/etc\/systemd\/system\/prometheus.service<\/code>:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code>[Unit]\nDescription=Prometheus Time Series Monitoring Engine\nDocumentation=https:\/\/prometheus.io\/docs\/introduction\/overview\/\nAfter=network.target\n\n[Service]\nUser=prometheus\nGroup=prometheus\nType=simple\nEnvironment=\"GOMEMLIMIT=3200MiB\"\nExecStart=\/usr\/local\/bin\/prometheus \\\n  --config.file=\/etc\/prometheus\/prometheus.yml \\\n  --storage.tsdb.path=\/var\/lib\/prometheus \\\n  --storage.tsdb.retention.time=30d \\\n  --storage.tsdb.retention.size=40GB \\\n  --storage.tsdb.min-block-duration=2h \\\n  --storage.tsdb.max-block-duration=2h \\\n  --web.listen-address=\"127.0.0.1:9090\" \\\n  --web.enable-lifecycle\n\n# Sandboxing and Kernel Protection\nProtectSystem=full\nProtectHome=true\nNoNewPrivileges=true\nLimitNOFILE=65536\nMemoryMax=4G\nRestart=on-failure\nRestartSec=5s\n\n[Install]\nWantedBy=multi-user.target<\/code><\/pre>\n<p>Validate the configuration syntax using <code>promtool<\/code> before starting the daemon:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code>sudo promtool check config \/etc\/prometheus\/prometheus.yml\nsudo systemctl daemon-reload\nsudo systemctl enable --now prometheus\nsudo systemctl status prometheus<\/code><\/pre>\n<h2>Step 3: Defining Production PromQL Alerting Rules<\/h2>\n<p>Monitoring without alerting requires constant human screen inspection. Write actionable alerting rules in <code>\/etc\/prometheus\/rules\/host_alerts.yml<\/code> to catch resource saturation before kernel deadlocks manifest:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code>groups:\n  - name: host_saturation_alerts\n    rules:\n      - alert: HostHighCpuLoad\n        expr: (100 - (avg by (instance) (rate(node_cpu_seconds_total{mode=\"idle\"}[5m])) * 100)) &gt; 85\n        for: 5m\n        labels:\n          severity: critical\n        annotations:\n          summary: \"High CPU load detected on instance {{ $labels.instance }}\"\n          description: \"CPU utilization has exceeded 85% for more than 5 minutes (current value: {{ $value | printf \"%.2f\" }}%).\"\n\n      - alert: HostOutOfMemory\n        expr: ((node_memory_MemAvailable_bytes \/ node_memory_MemTotal_bytes) * 100) &lt; 10\n        for: 3m\n        labels:\n          severity: critical\n        annotations:\n          summary: &quot;Host out of memory on {{ $labels.instance }}&quot;\n          description: &quot;Node memory available is below 10% (current: {{ $value | printf &quot;%.2f&quot; }}%).&quot;\n\n      - alert: HostDiskFillingUp\n        expr: (node_filesystem_free_bytes{mountpoint=&quot;\/&quot;}\/node_filesystem_size_bytes{mountpoint=&quot;\/&quot;} * 100) &lt; 15\n        for: 10m\n        labels:\n          severity: warning\n        annotations:\n          summary: &quot;Root filesystem running out of disk space on {{ $labels.instance }}&quot;\n          description: &quot;Root partition free space is below 15% (current: {{ $value | printf &quot;%.2f&quot; }}%).&quot;<\/code><\/pre>\n<p>Validate the rule syntax using <code>promtool check rules \/etc\/prometheus\/rules\/host_alerts.yml<\/code> and trigger a live Prometheus configuration reload via HTTP: <code>curl -X POST http:\/\/127.0.0.1:9090\/-\/reload<\/code>.<\/p>\n<h2>Step 4: Installing and Securing Grafana<\/h2>\n<p>Grafana transforms raw time-series data into actionable dashboards. For enterprise Linux deployments (Ubuntu\/Debian or RHEL\/Rocky Linux), install Grafana from the official package repository to maintain seamless security patch updates.<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code># Install Grafana APT repository and key (Debian\/Ubuntu)\nsudo apt-get install -y apt-transport-https software-properties-common wget\nsudo mkdir -p \/etc\/apt\/keyrings\/\nwget -q -O - https:\/\/apt.grafana.com\/gpg.key | gpg --dearmor | sudo tee \/etc\/apt\/keyrings\/grafana.gpg &gt; \/dev\/null\necho \"deb [signed-by=\/etc\/apt\/keyrings\/grafana.gpg] https:\/\/apt.grafana.com stable main\" | sudo tee -a \/etc\/apt\/sources.list.d\/grafana.list\n\nsudo apt-get update\nsudo apt-get install -y grafana\nsudo systemctl daemon-reload\nsudo systemctl enable --now grafana-server<\/code><\/pre>\n<h3>Production Grafana Hardening (`\/etc\/grafana\/grafana.ini`)<\/h3>\n<p>By default, Grafana binds to <code>0.0.0.0:3000<\/code> and allows user registration. Modify <code>\/etc\/grafana\/grafana.ini<\/code> to bind exclusively to localhost (serving traffic behind an NGINX reverse proxy with TLS), disable open registration, and enforce secure cookies:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code>[server]\nhttp_addr = 127.0.0.1\nhttp_port = 3000\ndomain = monitor.yourdomain.com\nroot_url = https:\/\/monitor.yourdomain.com\/\nenforce_domain = true\n\n[security]\nadmin_user = sysadmin\ncookie_secure = true\ndisable_gravatar = true\nhide_version = true\n\n[users]\nallow_sign_up = false\nauto_assign_org_role = Viewer\n\n[analytics]\nreporting_enabled = false\ncheck_for_updates = false<\/code><\/pre>\n<p>Restart Grafana to enforce the security posture: <code>sudo systemctl restart grafana-server<\/code>.<\/p>\n<h3>Automating Prometheus as a Grafana Data Source via Provisioning<\/h3>\n<p>Eliminate manual GUI clicks by configuring declarative provisioning. Create <code>\/etc\/grafana\/provisioning\/datasources\/prometheus.yaml<\/code>:<\/p>\n<pre class=\"wp-block-code\" style=\"background:#f3f3f3;color:#333;padding:16px;border-left:4px solid #001b41;font-family:monospace;font-size:13px\"><code>apiVersion: 1\n\ndatasources:\n  - name: Prometheus\n    type: prometheus\n    access: proxy\n    url: http:\/\/127.0.0.1:9090\n    isDefault: true\n    jsonData:\n      timeInterval: 15s\n      httpMethod: POST\n    editable: false<\/code><\/pre>\n<p>Upon restarting Grafana, the Prometheus data source connects automatically and is ready to power pre-built community dashboards such as the renowned <strong>Node Exporter Full (Dashboard ID: 1860)<\/strong>.<\/p>\n<h2>Step 5: Production Operational Checklist &amp; Infrastructure Sizing<\/h2>\n<p>Observability infrastructure requires deliberate capacity planning. As your monitored fleet grows from 5 to 500 nodes, write traffic to the Prometheus TSDB scales linearly with active time series.<\/p>\n<ul style=\"color:#444;line-height:1.8;font-size:15px\">\n<li><strong>TSDB Disk Sizing Formula:<\/strong> <code>Disk Required = (Scrapes\/sec) \u00d7 (Retention in Seconds) \u00d7 (Bytes\/Sample)<\/code>. Prometheus averages 1.3 to 2 bytes per sample due to Gorilla compression. At 10,000 samples\/sec with 30-day retention, allocate at least 45GB of dedicated NVMe storage.<\/li>\n<li><strong>Filesystem Tuning:<\/strong> Mount the TSDB storage directory on an <code>ext4<\/code> or <code>XFS<\/code> filesystem formatted with <code>noatime<\/code> to eliminate redundant write overhead for every head-block read.<\/li>\n<li><strong>Continuous Backup:<\/strong> Enable the Prometheus lifecycle API (<code>--web.enable-lifecycle<\/code>) and take transactional snapshots via <code>curl -X POST http:\/\/localhost:9090\/api\/v1\/snapshot<\/code> before performing kernel upgrades.<\/li>\n<\/ul>\n<p>For mission-critical production systems that cannot afford observability downtime, hosting your monitoring tier and application clusters on <a href=\"https:\/\/merahost.org\" target=\"_blank\" rel=\"noopener\">MeraHost Enterprise Cloud<\/a> guarantees dedicated Enterprise NVMe disk I\/O, unmetered network bandwidth, and hardened LiteSpeed acceleration with predictable pricing.<\/p>\n<h2>Frequently Asked Questions<\/h2>\n<details class=\"wp-block-group\" style=\"background:#f9f9f9;border:1px solid #e7e7e7;border-radius:4px;padding:14px;margin-bottom:12px\">\n<summary style=\"cursor:pointer;font-weight:600;color:#001b41\">How much disk space and RAM does Prometheus require in production?<\/summary>\n<p style=\"margin-top:10px;color:#444\">For a typical infrastructure with 10 to 50 Linux servers scraping 1,000 metrics per host at 15-second intervals, Prometheus requires approximately 4 GB to 8 GB of RAM and 40 GB to 80 GB of fast NVMe storage for a 30-day retention window. RAM consumption is primarily driven by active series in the TSDB Head block and query complexity.<\/p>\n<\/details>\n<details class=\"wp-block-group\" style=\"background:#f9f9f9;border:1px solid #e7e7e7;border-radius:4px;padding:14px;margin-bottom:12px\">\n<summary style=\"cursor:pointer;font-weight:600;color:#001b41\">Why use Prometheus pull architecture instead of push-based agents like Telegraf or Zabbix?<\/summary>\n<p style=\"margin-top:10px;color:#444\">The pull architecture ensures that the central monitoring server controls the ingestion rate and network socket allocation. If target systems experience resource spikes or high latency, Prometheus prevents network saturation. Furthermore, pull architectures make health detection immediate: if an HTTP scrape target fails to respond, Prometheus immediately records an <code>up == 0<\/code> metric.<\/p>\n<\/details>\n<details class=\"wp-block-group\" style=\"background:#f9f9f9;border:1px solid #e7e7e7;border-radius:4px;padding:14px;margin-bottom:12px\">\n<summary style=\"cursor:pointer;font-weight:600;color:#001b41\">Can Prometheus and Grafana be installed on the same server?<\/summary>\n<p style=\"margin-top:10px;color:#444\">Yes, for small to medium environments (under 100 monitored nodes), hosting Prometheus and Grafana on the same dedicated virtual or physical server is common and cost-effective. However, strict memory boundaries (via systemd <code>MemoryMax<\/code> and <code>GOMEMLIMIT<\/code>) must be configured to ensure a large Grafana dashboard query does not trigger the OOM killer on the Prometheus TSDB process.<\/p>\n<\/details>\n<details class=\"wp-block-group\" style=\"background:#f9f9f9;border:1px solid #e7e7e7;border-radius:4px;padding:14px;margin-bottom:12px\">\n<summary style=\"cursor:pointer;font-weight:600;color:#001b41\">How do I secure Node Exporter metrics from unauthorized access?<\/summary>\n<p style=\"margin-top:10px;color:#444\">You should never expose port 9100 publicly. Secure Node Exporter by binding it to a private internal network interface (e.g., WireGuard, Tailscale, or VPC subnet), enforcing host firewall rules (<code>iptables<\/code>\/<code>ufw<\/code>) to permit traffic only from the Prometheus server IP, or implementing basic authentication and TLS using Node Exporter web configuration files.<\/p>\n<\/details>\n<div class=\"wp-block-group has-background\" style=\"background:#f9f9f9;border:1px solid #e7e7e7;border-radius:8px;padding:32px;margin:40px 0;text-align:center\">\n<h3 style=\"color:#001b41;margin-top:0;font-size:24px;font-weight:700\">Deploy Enterprise-Grade Production Infrastructure<\/h3>\n<p style=\"color:#444;font-size:16px;line-height:1.6;max-width:680px;margin:12px auto 24px auto\">Need guaranteed performance with zero price hikes? Host mission-critical workloads on <strong style=\"color:#001b41\">MeraHost<\/strong> with pure Enterprise NVMe, LiteSpeed Web Server, and Same Renewal Price, Always (starting at \u20b999\/mo).<\/p>\n<div class=\"wp-block-buttons\" style=\"display:flex;gap:16px;justify-content:center;flex-wrap:wrap\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link\" href=\"https:\/\/merahost.org\" style=\"background:#001b41;color:#ffffff;font-weight:700;padding:12px 28px;border-radius:4px;text-decoration:none;display:inline-block;font-size:15px\" target=\"_blank\" rel=\"noopener\">Explore MeraHost NVMe Cloud &rarr;<\/a><\/div>\n<div class=\"wp-block-button is-style-outline\"><a class=\"wp-block-button__link\" href=\"https:\/\/cpanelfree.com\" style=\"background:transparent;color:#001b41;font-weight:600;padding:12px 24px;border:2px solid #001b41;border-radius:4px;text-decoration:none;display:inline-block;font-size:15px\">Deploy Free Staging on CpanelFree<\/a><\/div>\n<\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Deploy an enterprise Prometheus and Grafana server monitoring stack. Learn production TSDB sizing, Node Exporter hardening, and PromQL alerting.<\/p>\n","protected":false},"author":1,"featured_media":4894,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[213],"tags":[57,177,87,214,101],"class_list":["post-4895","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-monitoring","tag-almalinux","tag-databases-performance","tag-devops","tag-monitoring","tag-sysadmin"],"_links":{"self":[{"href":"https:\/\/cpanelfree.com\/blog\/wp-json\/wp\/v2\/posts\/4895","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cpanelfree.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cpanelfree.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/cpanelfree.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/cpanelfree.com\/blog\/wp-json\/wp\/v2\/comments?post=4895"}],"version-history":[{"count":0,"href":"https:\/\/cpanelfree.com\/blog\/wp-json\/wp\/v2\/posts\/4895\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/cpanelfree.com\/blog\/wp-json\/wp\/v2\/media\/4894"}],"wp:attachment":[{"href":"https:\/\/cpanelfree.com\/blog\/wp-json\/wp\/v2\/media?parent=4895"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cpanelfree.com\/blog\/wp-json\/wp\/v2\/categories?post=4895"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cpanelfree.com\/blog\/wp-json\/wp\/v2\/tags?post=4895"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}