One Live, One Standing By: How Flux Keeps Your App Online
Fluxers! Every machine on Flux is run by an independent operator, somewhere in the world, on their own hardware. Machines get rebooted, upgraded, moved and occasionally switched off. So how does an application with real data — a game world, a database, a WordPress site — keep running when the node under it goes away?
For a large share of the network, the answer is two letters at the start of a storage path: g:. It turns on Flux’s primary/standby mode, and 290 public applications use it today, alongside many more private ones. This article explains how it works.
Three kinds of storage
When you describe a component on Flux, you tell it where its data lives. The prefix on that path decides how the network treats it:
- No prefix — ordinary local storage on each instance. Fast and simple, ideal for stateless services where every copy is interchangeable.
- r: — replicated storage. Data is synchronised across every instance, and synchronisation starts only once all of them are confirmed running.
- g: — primary/standby. One instance is the primary: it runs the application and owns the data, read and write. The others are standbys: they keep a continuously synchronised copy of that data, ready to take over.
The g: mode is built for exactly the workloads where two writers at once would be a disaster — a game world, a database, a site with uploads. There is only ever one primary.
Keeping the copies in step
Under the hood, FluxOS uses Syncthing to replicate the data volume from the primary to its standbys. On the primary, the folder sends and receives. On a standby, it only receives. Changes flow one way — out from the live copy — so a standby is never at risk of writing back stale data.
Electing a primary
The nodes hosting an application decide among themselves which one is primary. Only the elected primary starts the component; on every standby, it stays stopped while its data keeps syncing.
The election is built with a strong bias towards caution, because the worst possible outcome is two primaries, or a primary running on an empty disk:
- No election before the disk is proven safe. After a reboot, a node waits for its storage monitor to complete a full successful pass before it will elect anything. That pass checks every synchronised volume is genuinely mounted, and switches any that are not to receive-only. A primary is never started on a folder that is not really there.
- “I don’t know” is never treated as “no”. If a node cannot tell who the primary is, it does not assume there is none and promote itself. It holds.
- A consistent view per pass. Every application elected in the same pass shares one view of the network, so two apps on the same node can never reach opposite conclusions about whether a quiet peer is down or the node itself is.
- A clean handover. A node that is the elected primary stops its component before it hands the application back, so the old primary is never still writing when the new one starts.
Sending visitors to the right place
Electing a primary is half the job. The other half is making sure your visitors and players reach it. That is the work of the Flux Domain Manager, which routes your application’s domain to the instance that is actually serving.
It remembers the current primary and checks that node first on every pass, rather than probing every instance each time. When the primary moves, it follows. It is patient by design, too: a slow check never cuts short the confirmations a primary gets, and an application is only forgotten after it has been gone for a full day — not after a single missed answer.
What it looks like from outside
For a player or a visitor, all of this is invisible, which is the point. The Flux game servers are a good example: a Dragonwilds server on Flux runs as two copies on separate nodes, one live and one on standby with the world synced, so the machine under your world is never a single point of failure.
If the node running your primary disappears, a standby that already holds your data becomes the primary, the Domain Manager points your address at it, and the world comes back — on a different machine, run by a different operator, possibly in a different city.
Cheaper, too
Because standbys hold data rather than running a full live copy of the application, primary/standby storage carries a 20% discount on Flux. You pay less for the configuration that keeps you online.
Where it is going
The FluxOS v9 specification on the roadmap makes this a first-class part of how an application is described: replicas and placement declared directly, including both active-standby and active-active layouts, alongside load balancing with health checks, sticky sessions, retries and connection draining.
Use it
To turn it on, prefix your component’s data path with g: — for example g:/data. The Flux documentation covers all three storage modes in the components guide, and the deploy editor in the new FluxCloud preview has a structured mount builder that sets it up for you.
Thousands of independent machines, any one of which can come and go — and an application that simply keeps running. That is what a decentralized cloud is supposed to feel like.
Posted in Education
by RunonFlux
Tags:
Comments
Leave a Reply
You must be logged in to post a comment.
