On this page
What server monitoring covers
Monitoring is how you find out about a problem before your customers tell you. It watches whether your site responds, whether the server has enough CPU, memory and disk, whether key services are running and whether certificates are about to expire. When something crosses a threshold, it sends an alert to email, Telegram, Slack or another channel you check.
Good monitoring is quiet most of the time. Too many alerts and people stop reading them; too few and problems go unnoticed. The work is in choosing what to watch and when to wake someone up.
Who needs it
- Businesses that learned their site was down from a customer message.
- SaaS products with paying users who expect the app to be available.
- E-commerce stores where downtime during a campaign means lost orders.
- Agencies responsible for many client sites on several servers.
- Anyone running background jobs, bots or scheduled tasks that fail silently.
What’s included
- External uptime checks for websites, APIs and key pages, from outside your network.
- Resource monitoring for CPU, memory, disk space, disk I/O and network.
- Service checks for web server, database, queues, containers and scheduled jobs.
- SSL certificate and domain expiry alerts with plenty of notice.
- Alert routing to the channels your team actually reads, with sensible escalation.
- Dashboards showing history, useful for capacity planning.
- Optional public or private status page.
How we set it up
- List what matters most: which pages, APIs, services and jobs would hurt the business if they stopped.
- Install monitoring agents or exporters and configure external checks.
- Set initial thresholds based on normal behaviour, not generic defaults.
- Route alerts and test them by triggering a real failure in a controlled way.
- Review alert history after a few weeks and tune out noise.
Tools we use
For simple uptime and certificate checks, Uptime Kuma is lightweight and easy to self-host. For detailed metrics and dashboards, we use Prometheus with node and service exporters and Grafana for visualisation and alerting. Cloud-native options such as AWS CloudWatch, Google Cloud Monitoring and Azure Monitor fit well when your infrastructure is already there. Alerts can go to email, Telegram, Slack or similar tools.
What affects timeline and cost
Basic uptime and resource monitoring for a few servers is quick to set up. More time is needed for many servers, application-level metrics, log-based alerts, on-call escalation, custom dashboards or integrating with existing tools. Hosted monitoring services have their own pricing; self-hosted tools need a small server of their own.
Where monitoring already exists but nobody trusts it, part of the job is cleaning it up: removing checks for services that no longer exist, fixing broken notification channels and adjusting thresholds so the alert history starts to mean something again.
If you would rather not run a separate monitoring server, hosted options are available, and we can set those up instead. The trade-off is ongoing subscription cost versus the small effort of maintaining a self-hosted tool, and we will talk it through with you.
What deserves an alert
Not everything that can be measured should wake someone up. We split signals into those that need immediate action, those that need attention soon, and those that are simply useful history:
| Level | Examples | How it is delivered |
|---|---|---|
| Urgent | Site or API down, database stopped, disk almost full | Immediate alert to the on-call channel |
| Soon | Certificate expiring in the coming weeks, backup failed, memory trending up | Alert during working hours |
| History only | CPU usage, traffic, response times | Dashboards for review and capacity planning |
Thresholds are based on how your server normally behaves. A database server that routinely uses most of its memory for caching should not alert on memory use, while a small web server that suddenly doubles its usage probably should. We watch the first few weeks of data and adjust, so that by the end of setup an alert is something your team trusts and acts on rather than skims past.
Common mistakes
- Alerting on every brief CPU spike, until alerts are muted and a real outage is missed.
- Monitoring only the homepage while the checkout, login or API is broken.
- Sending alerts to a shared inbox nobody watches out of hours.
- Never testing alerts, then discovering the notification channel stopped working.
Related: Monitoring & Logging and Linux Server Administration. See Cloud & Server Management, the cloud and server guide, or contact us.
Frequently asked questions
What is the difference between uptime monitoring and server monitoring?
Uptime monitoring checks from outside whether your site responds. Server monitoring watches the machine itself: resources, services and trends. You usually want both.
Can alerts come to Telegram or WhatsApp?
Telegram, email and Slack are straightforward with common tools. Other channels depend on what integrations are available.
Will monitoring slow down my server?
Monitoring agents use very little resource. We keep collection intervals reasonable.
Do you respond to the alerts as well?
We can, as part of ongoing server management, or alerts can go to your own team.
Can we get a status page for our customers?
Yes. Tools such as Uptime Kuma can publish a simple status page showing whether your key services are up, which reduces support questions during an incident.
Talk to us about server monitoring
Uptime, CPU, memory, disk and service checks that alert you before users notice.