Why Cloud Monitoring Is More Than a Simple Dashboard?
When you move your infrastructure to a cloud platform, there are no more network cables and physical servers in the server room; but the complexities have not only not decreased, they have also expanded to virtualization layers, software-defined networking, and managed services. In such an environment, cloud monitoring is the only way you can be sure your application is actually healthy, not just that the server is responding.
Many technical teams think that installing an open-source tool and viewing a few charts means monitoring is done. But real monitoring is a complete cycle: data collection, defining key metrics, setting smart thresholds, and finally sending alerts to the right channel at the right time. In this article, we will examine all four steps with practical examples and real code.
Key Metrics: What Should You Monitor?
The first common mistake is monitoring everything. If you define an alert for every metric, your team will quickly suffer from "alert fatigue" and ignore real alerts. You need to focus on metrics that directly affect user experience and service health.
Infrastructure Metrics
This category includes the basic resources that can usually be collected through tools like Prometheus, Grafana, or the native services of each cloud provider:
- CPU Usage: Average processor usage over 5- and 15-minute intervals. A rate of 80% for 10 minutes is usually a warning sign.
- Memory Usage: Percentage of consumed memory and swap usage. On Linux, combine the
free -hcommand with cron scripts. - Disk I/O and Disk Space: Free disk space and the number of read/write operations per second (IOPS). Disk filling up is one of the main reasons databases go down.
- Network Throughput: Incoming and outgoing bandwidth. This metric is the first sign during DDoS attacks or abnormal traffic.
Application Metrics
Infrastructure monitoring tells you that the machine is alive; but it doesn't tell you that your application is working correctly. For this, you also need to collect application-level metrics:
- Latency: Response time to requests. Monitor the 95th percentile (p95), not the average; the average can hide anomalies.
- Error Rate: Percentage of requests that respond with a 5xx error code. Suggested threshold: more than 1% over a 5-minute interval.
- Throughput: Number of successful requests per second (RPS). A sudden drop in RPS can indicate a service outage or a problem in the message queue.
- Queue Depth: Depth of message queues (such as RabbitMQ or Kafka). If the queue keeps filling up, the consumer is slower than the producer.
Business Metrics
In professional cloud monitoring, we don't only look at technical metrics. Metrics like conversion rate, number of successful transactions, or number of online users can also be added to the dashboard. This helps you measure the impact of a technical change on the business.
Setting Thresholds: From Unnecessary Alerts to Late Alerts
The threshold is the heart of the alerting system. If you set the threshold too low, your team will be flooded with notifications. If you set it too high, you'll find out about the problem when users have already complained. The solution is to use dynamic and multi-level thresholds.
Static Threshold vs. Dynamic Threshold
A static threshold like "alert if CPU exceeds 80%" is sufficient for small environments. But in large environments, traffic patterns differ. For example, a news website might have high CPU usage in the morning hours and be almost idle at night. In this case, a dynamic threshold calculated based on a 7-day moving average works more accurately.
In Prometheus, you can define dynamic thresholds using record rules. The following example shows a rule for alerting based on deviation from the average:
groups:
- name: dynamic-thresholds
rules:
- record: job:cpu_usage:avg_7d
expr: avg_over_time(instance:cpu_usage:rate5m[7d])
- alert: HighCpuDynamic
expr: |
instance:cpu_usage:rate5m > job:cpu_usage:avg_7d * 1.5
for: 15m
labels:
severity: warning
annotations:
summary: "CPU usage is 50% above 7-day average"
Multi-Level Thresholds
Instead of one alert state, define three levels:
- Info: For example, disk usage has reached 70%. This alert goes to the team's public channel and doesn't require immediate action.
- Warning: Disk usage has reached 85%. This alert is sent via SMS or notification to the person in charge.
- Critical: Disk usage has reached 95%. This alert is communicated via phone call or sent to an emergency notification channel.
This approach allows your team to prioritize and only be woken up in critical cases.
Common Mistake: Ignoring the "For Duration"
One of the most common mistakes in cloud monitoring is alerting based on a single sample. If CPU reaches 90% for 30 seconds, it might just be a short spike that resolves on its own. Always use the for parameter so that the alert only fires after the problem persists for a specific period (e.g., 10 minutes). In the example above, for: 15m means the problem must continue for 15 minutes before an alert is triggered.
Notification Channels: The Right Message, in the Right Channel
A good alert is one that gets seen. If an alert is sent to a channel that no one watches, it's practically useless. The choice of notification channel should be based on the severity of the alert and the required response time.
Channel and Alert Level Matrix
- Email: Suitable for Info-level alerts and periodic reports. Don't use email for critical cases; it might not be seen for hours.
- SMS: Suitable for Warning level. If your team works in shifts, send the SMS to the person on duty.
- Team Messengers (Slack, Telegram, Rocket.Chat): The best option for Warning and Critical level alerts. Using Webhooks, you can send alerts to specific channels.
- Phone Call: Only for Critical level. Services like PagerDuty or Opsgenie provide this capability.
Practical Example: Sending Alerts to Telegram via Webhook
Suppose you're using Alertmanager and want to send Critical alerts to a Telegram bot. First, create a bot and get its token. Then, in the Alertmanager configuration file, define a receiver:
receivers:
- name: 'telegram-critical'
webhook_configs:
- url: 'https://api.telegram.org/bot<YOUR_BOT_TOKEN>/sendMessage'
send_resolved: true
http_config:
headers:
Content-Type: application/json
body: |
{
"chat_id": "<YOUR_CHAT_ID>",
"text": "{{ range .Alerts }}{{ .Annotations.summary }}\n{{ .Annotations.description }}{{ end }}",
"parse_mode": "HTML"
}
Important note: In the Telegram Webhook, you must place the sendMessage method directly in the URL and send the request body as JSON. You can also get the chat_id value through the @userinfobot bot on Telegram.
Common Mistake: Sending All Alerts to One Channel
If all alerts (from Info to Critical) are sent to a single Telegram channel, team members will quickly mute the notifications. Be sure to separate channels: a public channel for Info and Warning alerts, and a restricted channel (or even a separate group) for Critical alerts that only the responsible people are members of.
Conclusion: Cloud Monitoring Is a Process, Not a Tool
Effective cloud monitoring is a combination of choosing the right metrics, setting smart thresholds, and sending alerts to the right channels. By implementing the cycle described in this article, you can move out of a reactive state and proactively identify problems before they impact users.
If your team is just starting with cloud infrastructure, I suggest starting with a simple tool like Prometheus and Grafana and gradually adding complexity. Remember that the ultimate goal isn't having a beautiful dashboard; the goal is reducing mean time to detect and mean time to recover (MTTD and MTTR). Cloud hosting services like what ServerNet provides usually give you basic monitoring tools; but fine-tuning alerts and metrics is always the responsibility of your technical team.
Finally, one practical recommendation: do an "alert review" every month. Remove or modify alerts that haven't led to any practical action in the past month. This will keep your alerting system always organized, efficient, and reliable.