Service Level Agreements (SLAs) are the contracts you make with your customers about uptime and availability. A 99.9% SLA sounds straightforward — but without systematic monitoring, incident logging, and reporting, proving compliance (or identifying breaches) becomes an exercise in guesswork.
This guide explains SLA fundamentals, the math behind uptime percentages, and how to use Vigilmon to monitor, document, and prove SLA compliance.
What Is an SLA?
A Service Level Agreement is a formal commitment — often contractual — that specifies the minimum level of service a provider guarantees. In the context of web services and APIs, SLAs typically define:
- Uptime percentage: The proportion of time the service must be available
- Response time: Maximum acceptable latency for API calls or page loads
- Incident response time: How quickly the provider acknowledges and begins resolving an outage
- Credits and remedies: What the customer receives if the SLA is breached (service credits, refunds)
Understanding Uptime Percentages
Not all nines are created equal. Here's what common SLA commitments actually mean in terms of allowed downtime per month:
| SLA | Monthly downtime allowed | Annual downtime allowed | |---|---|---| | 99.0% | ~7 hours 12 minutes | ~3 days 15 hours | | 99.5% | ~3 hours 36 minutes | ~1 day 20 hours | | 99.9% ("three nines") | ~43 minutes | ~8 hours 45 minutes | | 99.95% | ~21 minutes | ~4 hours 22 minutes | | 99.99% ("four nines") | ~4 minutes 19 seconds | ~52 minutes | | 99.999% ("five nines") | ~26 seconds | ~5 minutes |
Moving from 99.9% to 99.99% reduces your monthly tolerance from 43 minutes to just over 4 minutes. Each additional nine requires substantially better infrastructure, monitoring, and incident response — and commands substantially higher customer trust.
Calculating Uptime
The formula is straightforward:
Uptime % = ((Total minutes - Downtime minutes) / Total minutes) × 100
For a 30-day month:
- Total minutes = 43,200
- If you had 15 minutes of downtime: Uptime = ((43,200 - 15) / 43,200) × 100 = 99.965%
The challenge is accurate downtime tracking. Without external monitoring, you're relying on internal metrics — which may miss partial outages, CDN failures, or regional routing issues that affect real users but don't crash your servers.
Why External Monitoring Is Essential for SLA Compliance
Internal metrics (server CPU, pod health, application logs) tell you what's happening inside your system. But SLAs are promises to your customers — and customers access your service from the outside.
External monitoring from Vigilmon checks your service the same way customers do: making real HTTP requests from global probe locations and validating the response. This catches:
- CDN edge failures — your origin is up, but users in specific regions can't reach you
- DNS failures — your servers are running but DNS resolution is broken
- SSL certificate errors — HTTPS connections are rejected due to certificate issues
- Load balancer misconfigurations — some backend nodes are returning errors
- Partial outages — specific endpoints or API routes are failing while others work
All of these affect your SLA commitment to customers, even if your internal dashboards look green.
Setting Up Vigilmon for SLA Monitoring
Step 1: Map Your SLA Scope
Before creating monitors, define exactly what your SLA covers. Common SLA-scoped resources include:
- Primary web application (homepage, login, core user flows)
- Public API endpoints
- Authentication and authorisation services
- Key customer-facing integrations
Create a Vigilmon monitor for each resource included in your SLA.
Step 2: Configure Check Frequency
Your check interval determines the granularity of your downtime detection. For SLA purposes:
- 1-minute checks: Recommended for 99.99%+ SLAs where every 4-minute window matters
- 2-minute checks: Suitable for 99.9% SLAs — misses at most 2 minutes before detection
- 5-minute checks: Acceptable for 99.5% SLAs with more generous downtime tolerance
A 5-minute check interval means an outage could run for up to 5 minutes before detection. If your monthly tolerance is only 43 minutes (99.9%), that's over 10% of your budget consumed before you even know there's a problem.
Step 3: Enable Multi-Location Verification
Configure Vigilmon to verify failures from multiple probe locations before firing an alert. This eliminates false positives (a single probe location having a transient network issue) while ensuring you catch real outages affecting your users.
For SLA compliance, multi-location checks also give you geographic incident data — useful when a customer claims they experienced downtime in a specific region.
Step 4: Set Up Alerting for Rapid Response
SLA breach risk grows every minute you don't respond. Configure:
- Immediate alerts to your on-call engineer via Slack, PagerDuty, or Teams
- Escalation alerts if the incident is not acknowledged within 5 minutes
- Customer-facing status page updates so customers know you're aware and working on it
Step 5: Log All Incidents
Vigilmon automatically logs incident start times, end times, and duration. For each incident:
- Review the incident timeline in your Vigilmon dashboard
- Add internal notes documenting root cause and resolution steps
- Export the incident data for SLA reporting
Using Status Pages to Prove SLA to Customers
A status page transforms your monitoring data into customer-visible evidence of your reliability.
Vigilmon's status page feature lets you:
- Publish real-time status for each component or service covered by your SLA
- Display historical uptime — customers can see your 30-day, 90-day, and annual uptime record
- Post incident updates as events unfold, demonstrating active response
- Archive resolved incidents with root cause summaries
When a customer questions your SLA compliance, point them to your public status page. A transparent uptime history — even one that shows incidents — builds trust. Customers understand outages happen; what they don't forgive is silence and opacity.
SLA Credit Management and Incident Logging
When you breach an SLA, your contract likely requires issuing service credits. To manage this fairly and accurately:
- Export Vigilmon incident reports with exact downtime duration in minutes
- Apply the uptime formula to calculate the month's actual uptime percentage
- Compare against SLA thresholds to determine whether credits are owed
- Document the incident with root cause, contributing factors, and remediation steps
Vigilmon's response time history and incident logs give you the raw data for this calculation. A well-documented incident record — including what was monitored, when the alert fired, and when service was restored — is your best defence against disputed SLA claims.
SLA Reporting Templates
For monthly SLA reporting to customers, include:
Monthly SLA Report — [Month Year]
| Metric | Value | |---|---| | Reporting period | [Start date] – [End date] | | Total minutes | 43,200 | | Downtime minutes | [X] | | Achieved uptime | [X.XX%] | | SLA commitment | 99.9% | | SLA status | ✅ Met / ⚠️ Breached | | Incidents | [Number] | | Credits owed | [None / $X] |
Include links to your Vigilmon status page and any incident post-mortems.
Getting Started
SLA compliance starts with accurate, continuous, external monitoring. Without it, you're either over-crediting customers for downtime you didn't have, or under-acknowledging real outages that erode trust.
Set up Vigilmon to monitor every endpoint in your SLA scope, configure multi-location checks, and publish a status page that shows your customers you take availability seriously.
Start your free account at vigilmon.online — have your first SLA monitors running in under five minutes.