Founder narrative
Anya Petrova9 min read6 views

The month an SLA breach cost me $240 at $47K MRR: a founder diary (2026)

A composite founder diary (2026): at $47K MRR an enterprise customer's procurement team invoked section 8.2 and claimed a service credit for 47 minutes of downtime. Why their number and my number were different lengths and theirs was the one that counted, why the same 42 minute outage breaches in February and clears in July, why the $240 credit was a seventeenth of what the month cost, and the three clauses the large published SLAs carry that mine did not.

Minimalist editorial illustration: a long thin charcoal horizontal line with terracotta square end caps and a single small gap near its centre, above three stacked horizontal bars of increasing length in pale grey, warm sand and solid terracotta, joined to the gap by one hairline vertical rule.
Minimalist editorial illustration: a long thin charcoal horizontal line with terracotta square end caps and a single small gap near its centre, above three stacked horizontal bars of increasing length in pale grey, warm sand and solid terracotta, joined to the gap by one hairline vertical rule.
In this story
Our monitoring recorded 47 minutes of unavailability on the 14th. Per section 8.2 we are requesting the service credit.

Two sentences, from a procurement mailbox I had never seen before, at 8:52 on a Tuesday morning. Section 8.2 lived in a contract I had signed eleven months earlier and had not opened since. I was at $47,000 MRR in 2026, this one customer was $2,400 of that, and I spent the first twenty minutes of the day looking for my own signature page.

Quick answer (2026): This is a composite founder diary about an SLA breach from the vendor side of the table, as a small SaaS company rather than as an IT help desk. Three things surprised me. My outage and their outage were different lengths, and because they had external monitoring and I had application logs, their number was the one that counted. The allowance I had promised, 99.9% of a month, turned out to depend on which month it was: the same 42 minute outage breaches in February and clears in July. And the credit came to $240, roughly a seventeenth of what the month actually cost me. The clause I needed was never the credit table. It was the two sentences the large vendors print underneath theirs.

The number in the email was not my number

Their monitoring said 47 minutes. My application logs said 31.

Both were honest. Their figure came from an external probe hitting a health endpoint on a fixed interval from outside my infrastructure. Mine came from the application writing about itself. Those two instruments cannot agree during an outage, and the reason is the whole problem: an outage that stops requests reaching your application does not appear in your application's logs. My 31 minutes was the window in which my service was up and unhappy. Their 47 minutes included the sixteen minutes in which it was not answering at all, which is precisely the part I could not see.

I looked for something to put beside their number and found nothing. No external probe. No status page. No incident timeline written down while it was happening, because I had been busy fixing it. I could not dispute 47 minutes. I could only decline to agree with it, and those are not the same thing.

That is the first lesson and it has nothing to do with law. An SLA is enforced by whoever holds a measurement. If only one party is measuring, the contract effectively says whatever that party's dashboard says.

What 99.9% actually allows, and which month you are in

I had never calculated it. I had written 99.9% into a contract because 99.9% is the number that goes into contracts.

A downtime calculator such as uptime.is puts 99.9% monthly at 43 minutes and 50 seconds. That figure uses an average month. Real agreements do not. Google defines its Monthly Uptime Percentage as the total number of minutes in a month, minus downtime minutes, divided by the total number of minutes in a month, per the Compute Engine SLA. A calendar month, whichever one you are in.

So the allowance moves:

Scroll to see more

Month lengthTotal minutes0.1% allowance
February, 28 days40,32040 min 19 s
30 day month43,20043 min 12 s
31 day month44,64044 min 38 s

A 42 minute outage is a breach in February and is not a breach in July. Same failure, same duration, same contract, opposite answer. Nobody writing about SLAs had mentioned this to me, and it is not a technicality: it is a 4 minute 19 second swing on a 43 minute budget, which is ten percent of the entire allowance.

My outage fell in a 31 day month. Their 47 minutes was over by 2 minutes and 22 seconds. My 31 minutes was under by 13 minutes and 38 seconds. The whole month turned on 142 seconds that only one of us could measure.

Nobody had defined what "down" meant

The part that took longest to accept is that both numbers were wrong about the customer's actual experience.

Their probe hit a health endpoint. My health endpoint answered as soon as the web tier came back, so their clock stopped there. But the background job queue was forty minutes behind by then, and everything this customer actually used my product for ran through that queue. So from their users' point of view the product was unusable for well over an hour, on both sides of a window neither instrument was measuring.

They never claimed for that, because they could not. My contract, like almost every uptime SLA I have since read, defined downtime as total unavailability of the service. Degraded, slow, backlogged and wrong-but-responding are all fully compensable experiences for the customer and are worth exactly nothing under the agreement.

That cuts both ways, and it is worth saying plainly. Customers who feel genuinely let down frequently have no claim, and vendors who feel unfairly claimed against frequently have no defence. Neither situation is anybody acting in bad faith. It is that a single availability percentage is a very crude instrument for a question that is actually about whether the product did its job, and everyone signs it anyway because it is the instrument that exists.

The honest version of my own position: I owed them $240 for the thing that was measurable, and I owed them something much larger and unwritten for the thing that was not.

The credit was $240

The contract said 10% of the monthly fee for a month below 99.9%. The monthly fee was $2,400. The credit was $240.

Here is what the month cost, as best I can reconstruct it:

  • Two hours of a lawyer reading section 8.2 and explaining what it did not say: about $700
  • Roughly three working days of my own attention, spread over two weeks, which is the expensive line and the one nobody invoices you for
  • Setting up external monitoring, which is a bill in the low tens of dollars a month and was overdue anyway
  • Building and populating a status page
  • A renegotiation I had not planned to open

Call it $4,000 in cash and time against a $240 credit. Seventeen to one. And the $240 was the only figure anyone had written down in advance.

This is the part that reads backwards from the outside. Service credits are not the risk. They are capped, small and predictable, which is exactly why the big providers are relaxed about publishing them. The risk is the unbudgeted fortnight, the renegotiation you enter from a weak position, and the customer quietly deciding at renewal that you are not ready for them.

What the large SLAs say, and mine did not

I went and read three real published agreements properly, for the first time, from companies whose scale I do not have but whose lawyers I would like to borrow.

Google Cloud logo Atlassian logo

They agree on the shape. Credits are a percentage of the bill, tiered by how far you missed. Amazon's EC2 SLA commits to 99.99% at region level and pays 10% below that but at or above 99.0%, 30% below 99.0% down to 95.0%, and 100% below 95.0%. Google's tiers are 10%, 25% and 100%. Atlassian's cloud SLA commits to 99.9% on its Premium plan and 99.95% on Enterprise, with credits from 5% up to 50% depending on plan and shortfall, "calculated as a percentage of the monthly fees attributed to the affected Eligible Cloud Product".

My contract had a version of that table. What it did not have were the three sentences that sit around it, and every one of the three cost me something.

One. A claim window. Amazon requires that "your credit request must be received by us by the end of the second billing cycle after which the incident occurred". Atlassian requires a ticket "within fifteen (15) days after the end of the calendar month". Mine required nothing. A credit claim against me was open forever.

Two. Sole remedy. Google's agreement states plainly that "this SLA states Customer's sole and exclusive remedy for any failure by Google to meet the SLO". That single sentence is what turns a credit table into a ceiling. Without it, my credit table was not a cap on my exposure. It was a floor with a number printed on it.

Three. A maintenance exclusion. Atlassian excludes downtime caused by "routine scheduled maintenance or reasonable emergency maintenance". I had no exclusion at all. Which means the twenty minute migration window I had announced by email two weeks in advance, at 3am on a Sunday, with nobody logged in, counted against my own uptime number. I had spent a year carefully scheduling maintenance and then contractually agreed to be penalised for it.

The second claim, for a month I had already closed

Three weeks later, the same mailbox claimed for an earlier month.

They were entitled to. Nothing in my contract stopped them. I went back through logs I had never kept for this purpose, reconstructed a month I had mentally filed as finished, and eventually got to 99.94%, comfortably inside the commitment. No credit owed. It took most of a day.

That is when the claim window stopped looking like a way to avoid paying. A claim window is not a shield. It is a deadline that keeps the question answerable while the evidence still exists. Fifteen days after month end, everyone still has the logs and everyone still remembers. Four months later, one party has a dashboard and the other party has an afternoon of archaeology.

What I changed

In this order, which matters:

  1. Measurement before contract. Better Stack style external probes from two regions, on a short interval, before I touched a single clause. I did not want to renegotiate anything until I could answer a claim with a number of my own. Grafana Cloud or a self hosted probe would have done the same job.

Better Stack logo Grafana logo Statuspage logo

  1. A public status page with a written incident log. Every incident gets a timestamped entry while it is happening, not afterwards.
  2. The three missing clauses, added at renewal: a 30 day claim window, an explicit sole and exclusive remedy sentence, and a maintenance exclusion with 48 hours notice.
  3. I removed the uptime commitment from the standard plan entirely. Not lowered. Removed. An unmeasured 99.9% on a self serve plan is a liability I was carrying for no revenue and no competitive reason. The commitment now exists only where it is negotiated, priced and measured.
  4. I stopped signing numbers I was not measuring. This is the same lesson the security questionnaire taught me at $34K MRR and I had not generalised it. A commitment you cannot evidence is not a commitment, it is an exposure.

The uncomfortable part: the enterprise contract that contained section 8.2 was the first one I ever closed, at $21K MRR, when I was so relieved to be signing anything that I read the commercial terms carefully and the operational ones not at all. Every clause that hurt me at $47K was one I had agreed to at $21K.

What actually happened

I paid the $240 as a credit on their next invoice, without arguing about the 142 seconds. Arguing would have cost more than $240 in my own time on day one, and it would have cost the relationship.

They renewed, at a higher price, on the rewritten agreement. MRR went from about $47K to roughly $48K across the following two months, which had nothing to do with the SLA and everything to do with the fact that the renewal conversation happened at all.

The thing I did not expect: months later, their procurement contact told me the status page was what changed their mind. They had never doubted that the outage happened. They had doubted whether I would tell them about the next one.

The one thing I would tell you

Open your own customer agreement and find the uptime number. Then do two things that cost nothing.

Compute the allowance for the shortest month of the year, not the average one. If your number is 99.9%, that is 40 minutes and 19 seconds in February, and it is the tightest budget you have agreed to all year.

Then ask what measurement you own that could answer a claim about last month. If the answer is your application logs, you do not have a measurement. You have a record of the requests that arrived, which is the one thing an outage guarantees will be incomplete.

I found both of those out on a Tuesday morning, from a procurement mailbox, eleven months after signing. It is a much cheaper hour when nobody is waiting for your response.

A

Written by

Anya Petrova

Anya Petrova writes first-person founder diaries for OperatorBook, reconstructed as composites from interviews with bootstrapped SaaS founders. She focuses on the months that do not make the highlight reel: the pricing changes, the churn scares, and the quiet operational decisions that move MRR.

Frequently asked questions

Is this a real founder's diary?

It is a composite. The founder is a blend of several bootstrapped SaaS operators who received uptime SLA credit claims from enterprise customers in 2026. The MRR figures (about $47K rising to roughly $48K), the $2,400 monthly contract value, the 47 versus 31 minute measurement gap, the $240 credit, the roughly $4,000 in time and fees and the two claims are self-reported and lightly rounded. No single named customer, company, contract or procurement contact is described. The downtime arithmetic and every quoted clause are real: the published Amazon EC2, Google Cloud Compute Engine and Atlassian cloud service level agreements are cited directly and were read in full, and the month length calculations are shown so you can check them.

What counts as an SLA breach?

For an uptime SLA, a breach is a Monthly Uptime Percentage below the committed figure, measured over a defined period against a defined definition of downtime. Two details do most of the work. First, the period is normally a calendar month rather than an average month, so the allowance changes with month length: 99.9% is 40 minutes 19 seconds in February and 44 minutes 38 seconds in a 31 day month. Second, the contract decides what counts as downtime at all, which is why exclusions for scheduled and emergency maintenance appear in every well drafted agreement. A breach is not the same thing as an outage.

How much downtime does 99.9% uptime allow per month?

It depends on the month. A 30 day month contains 43,200 minutes, so 0.1% is 43 minutes 12 seconds. A 31 day month contains 44,640 minutes, giving 44 minutes 38 seconds. February, at 28 days and 40,320 minutes, allows only 40 minutes 19 seconds. Calculators that quote 43 minutes 50 seconds are using an average month of about 30.44 days, which no calendar month actually is. If your agreement measures per calendar month, the shortest month of the year is the tightest commitment you have made, and it is the one worth checking first.

Do I have to pay a service credit if the customer never asks for it?

In the published agreements of the large providers, no, because the credit is claim based. Amazon states that a credit request must be received by the end of the second billing cycle after the incident, and Atlassian requires a ticket within fifteen days after the end of the calendar month in which the outage occurred. If a claim window is absent from your own contract, as it was from mine, a claim can arrive at any time, including for a month whose logs you no longer hold. A claim window is less about avoiding payment and more about keeping the question answerable while the evidence still exists.

What does sole and exclusive remedy mean in an SLA?

It is the sentence that turns a credit table into a ceiling. Google's Compute Engine SLA states that the SLA is the customer's sole and exclusive remedy for a failure to meet the service level objective, which means the published credit is the whole of the customer's contractual recourse for that failure. Without such a sentence, a credit table sets out what you will definitely pay without limiting what else might be claimed. If your agreement lists credit percentages but never says they are the exclusive remedy, you have written down a floor rather than a cap.

Should a small SaaS company offer an uptime SLA at all?

Only where it is negotiated, priced and measured. The failure mode described here is a self serve plan carrying a 99.9% commitment that nobody had ever calculated, and which therefore produced liability with no matching revenue. The order that worked was to instrument first and contract second: put external probes and a public status page in place, so that a claim can be answered with a number you own, and only then negotiate the commitment. An uptime number you cannot evidence is not a promise to your customer, it is an exposure on your own balance sheet.

Founder narrative

The month I closed my first enterprise deal at $21K MRR

The month my SaaS hit $21,000 MRR I signed my first enterprise deal: one $40,000-a-year contract worth 16% of my revenue. A first-person 2026 diary on the 280-question security review, the SSO build I pulled off the roadmap, net-60 cash delays, concentration risk, and the three rules I wrote afterward.

9 min read105
Founder narrative

The month I lost a day of customer data at $40K MRR: a data-loss diary (2026)

In 2026, the fastest way to lose your biggest customer is a backup you never tested. A composite founder diary about the night a botched production migration erased roughly 22 hours of customer data at $40K MRR, when the newer backup had been failing silently for months and the freshest clean copy was almost a day old. It covers what was lost, the email that had to go out, the agency that churned, and the restore-drill and point-in-time-recovery changes that cut the recovery point objective from about a day to five minutes.

8 min read91
Founder narrative

The month a website accessibility lawsuit threat landed at $45K MRR: a founder diary (2026)

A composite founder diary (2026): at $45K MRR a twelve page ADA demand letter arrived by certified mail alleging eleven barriers and giving fourteen days to respond. Why a demand letter is not a website accessibility lawsuit and what that changes, why the published filing counts of 3,117 and 4,928 describe a category I was not in, the overlay widget I nearly bought at 11pm, the nine defects behind the eleven allegations, and what the whole month actually cost.

12 min read33