Building to last

Software that fails gracefully

Every system breaks somewhere. The good ones keep working when a piece fails. Four habits we build in, and where they’ve saved the day.

2 min read

Every piece of software fails eventually. A network drops, a provider has an outage, a person clicks the wrong thing, a disk fills up. Nobody can stop that. What you can decide is what happens next.

Software that fails gracefully loses a little and carries on. Software that doesn’t loses everything at once, usually at the worst moment. The difference is almost never clever code. It’s a handful of habits, built in from the start.

1. Keep the important part out of the fragile part

The fragile part of a system is whatever you don’t control: an outside service, an AI model, someone else’s API. The important part is the answer your users depend on.

Our fantasy football draft assistant had one deadline that couldn’t move: draft night. It uses an AI advisor to explain its suggestions in plain words, but every number comes from ordinary arithmetic. If the AI is slow, wrong or down, the dollar values still update with every pick. The explanation is a bonus, not a dependency.

2. Start from what you already have

A system that waits for the server before it shows anything is only as good as the connection. Web Messaging opens instantly from a copy kept in the browser, then fetches only what’s changed. On a bad connection it’s a little behind, not blank.

3. Make actions undoable

People make mistakes, and so do voice commands. Our self-hosted voice assistant can undo what it just did, and each of its tools is marked with whether it can be undone at all. “That wasn’t what I meant” is a normal thing to say to a computer, and it should have a normal answer.

The same idea works anywhere: soft deletes instead of hard ones, a history of changes, a confirmation only for the things that truly can’t come back.

4. Prove the backups

A backup that’s never been restored is a hope, not a backup. CenterSuite takes nightly encrypted backups, and once a month there’s a restore drill: the backup is actually restored and checked. The drill is what turns “we have backups” into “we can recover”.

The test we use

For any system we build, we ask: if this one piece disappeared for an hour, what would people see? If the answer is “everything stops”, we change the design before writing more code.

It’s not glamorous work, and when it’s done well nobody notices. That’s rather the point.

Next step

What’s the job you’d like to stop doing by hand?

Describe the problem in a few sentences. Within one working day you’ll hear whether software would fix it, roughly what it would take, and what we’d look at first.