Cates Works
All capabilities
Care

Monitoring & Reliability

A site that’s down or slow quietly loses you money. As part of a Care plan I set up uptime and performance monitoring, error reporting, and health checks, and keep an eye on them for you — so issues get caught and resolved before they turn into missed calls or abandoned carts.

“I only find out my site is broken when a customer tells me — if they bother to.”

What's included

  • Uptime & performance monitoring
  • Error & crash reporting
  • Health-check endpoints
  • Security & dependency alerts
  • Proactive fixes before they cost you

Related reading

The empty account that loaded forever

Green builds, passing health checks, every API returning 200 — and the first person to sign up got a loading skeleton that never resolved. The cause was an empty list that no screen had been designed to render.

The API said 200 OK. Three times, nothing had happened.

A value I could write but never read back, a PATCH that returned success and quietly persisted nothing, and a login test that could only pass by defeating the security control I had just installed. Three ways a success response lied during one production deploy.

Your deploy checklist is a hypothesis, not a fact

I brought a battle-tested deploy checklist to a repo that was shaped differently underneath. Three times in one session it asked me to build something the app already did another way. Then the deploy went out clean and broke twice in an hour — and a HAR file turned both investigations into two-minute fixes.

I Almost Shipped a Production Outage Twice — Because the Dashboard Said Everything Was Fine

Every deploy tool on the market will tell you “Ready,” “Healthy,” “Success.” During one production rollout, all three were true and the site was still broken or serving stale code — twice, for two unrelated reasons. Here’s the verification discipline that actually catches what a green dashboard can’t, and why I run it on every deploy, not just the ones that go wrong.

The bug report cited a fix. The fix was for a different bug.

A detailed security report landed on one of my own systems, complete with a specific commit as proof this exact class of bug had already been caught once before. The commit was real. It fixed something else entirely. Here’s what checking the citation — instead of just the argument — turned up.

I stopped trusting my own review, so I made two AI agents argue about it

The agent that writes a fix is the worst-positioned reviewer of that fix — it already believes the design is right, or it wouldn’t have built it that way. Here’s what changed when I stopped asking one agent to check its own work, and started asking a second one to try to break it instead.

It wasn’t DNS — but everyone had already decided it was

The site went down hours after a DNS change, so the DNS change was obviously the cause. It wasn’t. Here’s the detail that ruled it out in thirty seconds, and the two cheap decisions that actually caused the outage.

The 200 OK that means “page not found”

Your error page can look perfect and still be lying to Google about whether it exists. A tour of the bugs that pass every automated check and only show up in a real browser.

A green deploy is not a working site

A quarter of my own site was returning server errors while every build was green and every deploy said “Ready.” Here’s the failure class that no build step can catch, and the check that now runs after every release.

Permission bugs don’t throw errors. They just leak.

A crash tells you something broke. A broken permission check quietly returns someone else’s data with a 200 OK. Here’s what a hard audit of my own platform turned up, and the rule that closed all of it.

Speed is a feature: why every 100ms costs you customers

Performance isn’t a technical vanity metric — it’s the first impression, the conversion lever, and the SEO signal most sites quietly fail. Here’s how I think about making sites genuinely fast.

Let’s talk

Have a project, or a product that could work harder?

Most projects begin with a short, no-pressure discovery call.