Designing FitSynQ for Real-World Reliability
Designing FitSynQ for Real-World Reliability
How FitSynQ keeps Ghanaian gyms running through everyday power and connectivity interruptions
Most SaaS (software-as-a-service) products are designed around two assumptions that rarely get stated out loud: the lights stay on, and the internet is there when you reach for it. In Ghana, those assumptions are usually fine. Until they are not.
Ghana's grid and internet infrastructure are good enough to run modern businesses every day. However, power cuts and network cuts occur often enough to warrant deliberate product design considerations. If the product assumes constant power and fast internet during operating hours, routine interruptions quickly become operational friction for gym staff and members.
FitSynQ is a multi-tenant gym management platform, with four main components: a NestJS backend, a Next.js admin dashboard, a member mobile app, and a check-in kiosk.
In this post I want to walk through the design decisions made across the system to keep the product useful even on imperfect days.

The dashed lines are the connections we assume can vanish: the kiosk's link to the backend, and Redis. Everything the kiosk needs to admit a member lives on the gym side of that boundary, on the device itself.
The idea underneath the whole design is that we don't treat offline as a rare edge case to degrade into gracefully. Members turning up to check in is one of the product's most frequent workflows, so the kiosk shouldn't need a live internet connection to do it. The rest of this post is how that plays out.
The check-in cannot depend on the network
A member scans their phone at the kiosk. If that needs a round trip to our server, then a brief dropped connection can turn into a queue of irritated people. So the kiosk is built to admit members with no network at all.
Three things get primed on startup and refreshed on a timer while there's a connection. They're independent, so one failing doesn't take the others down with it:
// Scanner.tsx: non-blocking, fault-tolerant priming
Promise.allSettled([
fetchAndCachePublicKey(), // RSA public key → localStorage
refreshOfflinePassIfNeeded(), // short-lived kiosk authority token
refreshRoster(), // branch member snapshot → IndexedDB
]);
With those cached, the whole check-in flow runs without the network.
Token validation happens locally. Member QR codes are RS256-signed JSON Web Tokens (JWTs). The kiosk fetches our RSA public key once, caches it, and verifies scanned tokens in-browser with jose. It carries a 300-second clock skew tolerance on purpose: a kiosk that's been offline, restarted, or running on a machine with a slightly drifting clock should not reject a member who is standing right there with a valid pass. Validation also tells expired apart from invalid_signature and no_key, so the screen can say "open your app to refresh" instead of a useless "invalid."
The kiosk also keeps an offline pass, a short-lived (4h) token it fetches while it's authenticated and refreshes once it's within thirty minutes of expiry. Before it will bank check-ins on its own, it confirms a non-expired pass is present. That check is a local expiry gate, not a signature check, and the pass never gets sent to the server. Its job is narrow: you can only get a pass by authenticating, so holding a fresh one means the device signed in within the last few hours, which stops a long-dormant kiosk from quietly admitting members. The real server-side authority lives elsewhere. When the queue syncs, every event rides in on the kiosk's device-token auth, which the backend verifies. The QR codes are the opposite case: verified cryptographically right there on the device, as above. The pass gates behaviour; the QR proves identity.
The check-in gets written down before anything else happens. Accepted scans go into IndexedDB first:
// offline_checkins object store: literally commented "write-ahead log"
{ id, clientEventId, qrToken, kioskId, occurredAt,
status: 'pending', retryCount: 0, createdAt }
The member sees "You're checked in!" the moment that row lands. It's optimistic UI, but only in the display sense, since the durable record already exists on the device. Duplicates get caught at two levels. The clientEventId has a unique index, so a sync that retries the same event can't create a second row. Two separate scans of the same member are a different problem: the kiosk collapses repeat scans within a short window before they're ever queued, and the server independently treats a scan for a member who already has an open session as a duplicate at sync time. The server is the authority there; the local window just spares the member a confusing second confirmation while offline.
When connectivity comes back, a one-shot justCameOnline signal (true for exactly one render, then cleared) kicks off the sync manager. It batches pending rows into one POST /sessions/sync-offline with a 30-second timeout and applies the server's verdict per event: created or duplicate marks the row synced, error bumps its retry count, ceiling of five.
The decision I care about most in this whole flow is four lines of nothing:
} catch {
// Transport failure: leave everything 'pending'.
// Do NOT consume retry budget. Flaky connectivity is not bad data.
}
A network blip leaves the queue alone. Retry budget is only ever spent on genuine server-side rejections. On a connection that is having a bad hour, spending a check-in's five retries on transient failures would quietly delete valid data, so the code refuses to do it.
An offline check-in has to count exactly once
Storing a check-in offline is the easy part. The hard part is making sure it lands exactly once when the connection comes back, no matter how many times the sync retries or a request half-succeeds before the network drops again.
The kiosk handles that by giving every scan its own event id the moment it happens, before any server is involved. That's the clientEventId in the write-ahead record: a crypto.randomUUID() stored with the queued row. When the queue syncs, the server dedupes on it (findFirst({ where: { clientEventId } })) and skips anything it has already recorded. A sync that times out after the server already committed, then retries, can't create a second session. The kiosk owns the identity of the event; the server owns the record it becomes.
There's a smaller, related habit underneath this. Every database record's id is generated in application code as a UUID v7 (universally unique identifier), not by Postgres, so there isn't one @default(uuid()) in the schema. Generating ids in app code lets the server build a record and its relationships in one transaction without a round trip, and UUID v7 is time-ordered (the first 48 bits are a timestamp), so you get that without the index fragmentation random UUIDs cause. Those ids are globally unique with no central sequence too, which comes back later when we talk about sharding.
Redis: optional as a cache, load-bearing as a queue
Caching layers have a habit of turning into hard dependencies when nobody's looking. If the app starts throwing 500s when Redis hiccups, congratulations, you now have two things that have to be up instead of one.
At boot, the cache module doesn't take the connection object's word for it. It runs a real set/get/del round trip and compares the value that comes back. If any of that fails it returns a config with no store attached, and NestJS falls back to an in-memory cache. A Redis outage at startup makes the API slower; it doesn't make it dead. The socket is tuned to fail fast rather than hang (connectTimeout: 10s, commandTimeout: 5s, disableOfflineQueue: true) with a reconnect strategy that gives up after three attempts instead of retrying forever.
Per operation, every get swallows its error and returns undefined, which callers already treat as a cache miss and fall through to Postgres for. There's also an application-level 5-second Promise.race sitting behind that as a backstop, with a comment admitting it exists "in case Redis-level timeouts don't fire." The rule holds throughout: a cache failure, invalidation included, gets logged and ignored, never thrown. Losing the fast path should never block the correct one.
The cache story ends there, but Redis wears a second hat that isn't optional. It's also the BullMQ backend and the distributed-lock layer. If Redis is genuinely down, the worker won't start, and an incoming Paystack webhook can't be enqueued, so the controller deliberately returns a 500 instead of a 200. That makes Paystack redeliver within its 72-hour window rather than dropping the event silently. It's the correct behaviour, but it isn't graceful degradation: for the queue, Redis is a hard dependency, and the safety net is Paystack's retries, not a local fallback. So the two roles get treated differently on purpose. As a cache, Redis is optional and degrades to memory. As the job queue and lock, it's load-bearing, and we lean on the payment provider's redelivery to cover the gap.
Payments need replay safety
Payments are where a temporary network problem stops being annoying and starts being a way to lose or duplicate real money. We run both cards and mobile money through Paystack, which processes the two rails behind one integration. Cards behave the way most payment code assumes; mobile money doesn't, so a fair amount of what follows is about the specific places MoMo breaks those assumptions.
Anything durable goes through BullMQ. A Paystack webhook gets its 200 the moment it's persisted and enqueued, not after it's processed. The job carries attempts: 5 with exponential backoff and a jobId equal to the webhook ID, so a Paystack redelivery can't spawn a second concurrent job. Exhaust the retries and it dead-letters: the row stays processed: false with a marker for monitoring to replay it, rather than disappearing. The worker is a separate process with no HTTP server at all, which means it can crash, restart, or move independently of the API and then pick its jobs back up straight from Redis.
Idempotency is layered by consequence. The event ID is a SHA-256 of signature:eventType:payloadId, so it can't be forged predictably. Payout webhooks go further and gate on state: a "transfer failed" event only re-credits a balance if the payout is still PROCESSING. Re-crediting one that already settled would literally manufacture money on a duplicate delivery.
The rail-specific realities are encoded directly. Money is tracked in pesewas, integer minor units, end to end, with no float ever touching currency, whichever way it was paid. Mobile-money payouts are capped at GHS 10,000 (Ghana cedis) per transfer with the remainder rolled into the next run, because that's a real limit we don't get to argue with. Card subscriptions auto-charge, but mobile money can't, since MoMo mandates aren't reusable the way a saved card is, MoMo subscriptions fall back to MANUAL_RENEWAL with reminders instead of silent charges. And a failed card charge retries on a [1, 3, 7]-day backoff behind a soft grace window that keeps the subscription active with its end date extended across the whole schedule. Nobody should lose gym access because their bank had a bad Tuesday.
The scanner keeps reading when the tab isn't focused
The last piece is physical hardware, and it caught me out.
The original scanner integration was a keyboard wedge: a document keydown listener watching for the characteristic burst of a barcode scan. Works fine, right up until you remember it only works while the kiosk tab has focus. On a shared reception PC, staff switch between apps and tabs all day, and scans can start landing in the wrong field or nowhere at all.
So check-in scanning moved to the Web Serial API, reading raw bytes off the USB-COM (serial) port. Serial data reaches the tab no matter which window has focus, so the kiosk keeps admitting members in the background while the same machine does five other jobs.
The hook around it is defensive in ways that only make sense once you've watched real hardware in a real reception setup. It retries the port open through the async-close race you hit on a hot re-pair. It backs off on read errors and declares the link broken after five consecutive failures rather than hot-spinning. It auto-reconnects on replug. And it falls back to the old keyboard-wedge path when the browser has no Web Serial, so a scanner someone left in HID (human-interface-device) mode still works.
Observing the system so we catch problems fast
There's one more piece. This one isn't about surviving failures. It's about seeing them. A robust product also has to be observable. We need eyes on a distributed system (an API, a separate worker, Postgres, Redis, BullMQ) and a way to catch problems before a gym owner calls us about them. So the backend ships real telemetry, not a console.log and hope.
It's OpenTelemetry-first. Instrumentation is imported as the literal first line of main.ts, before NestFactory, so the auto-instrumentation can hook express, pg, and ioredis before those modules load. Traces, logs, and metrics all leave the app over OTLP (the OpenTelemetry Protocol) to a separate OpenTelemetry Collector, which fans out to three places: Sentry for errors (with alerting on new and spiking issues), Loki for logs, and Grafana Cloud and Prometheus for metrics and dashboards. Every log carries the active traceId, the request ID, and the tenant (organizationId, branchId), so an error in Sentry or a slow request on a dashboard is one click from its exact log lines and trace. The global exception filter tags each captured error with that same tenant, user, and request context and maps severity sensibly: 4xx demoted to info, client disconnects (ECONNRESET, EPIPE) dropped as noise, so the alerts that fire are the ones worth waking up for.
The metrics are the usual RED signals (request rate, errors, duration histograms) plus domain counters that show the business actually flowing: check-ins, registrations, active sessions, payouts tracked in pesewas with their own success, failure, and duration series.
Two details keep the dashboards honest. The first is restart behaviour. Prometheus counters live in memory, so a restart (a deploy, an out-of-memory kill, the platform recycling an instance) would otherwise blank today's numbers. So point-in-time values are re-derived from Postgres on boot: the active-session gauge and the day's check-in and registration totals come straight from the database, which means a 3am redeploy doesn't leave a hole in the graphs.
// MetricsService.onModuleInit
await this.syncActiveSessionsFromDatabase(); // re-derive live truth from Postgres
The second is cardinality. URL paths are normalized (/members/:id, not a fresh series per uuid) and dynamic Paystack error strings are collapsed to a fixed allow-list, so a burst of novel error messages can't explode the series count. Per-branch labels do grow with the number of branches, but that's a bounded, known quantity, a few dozen gyms rather than open-ended user input, so we let it grow.
The health checks report more than up or down. /health/redis returns a third state, degraded, when the cache has fallen back to in-memory or is answering slower than a second. That surfaces the exact condition from the Redis section: the system is serving correctly but without its cache, visible on a status check instead of buried in a latency graph.
The durability guarantees from earlier are queryable too, which is what turns them from a claim into something we can monitor. A dead-lettered payment webhook leaves its row processed: false with a DEAD_LETTER marker, so one query surfaces everything that exhausted its retries and needs a replay. If a payout's balance-revert fails mid-way, the row is stamped metadata.reconciliationNeeded: true and logged as CRITICAL, so an inconsistent balance is a row we can find rather than money that quietly went wrong. The worker logs a memory heartbeat every minute and threads an exitReason through every shutdown path, so if it dies we get told why instead of finding a silent gap in the graphs.
Scaling past one of everything
Right now the whole backend is one API server and one worker. That's honest about where the product is: a handful of gyms, not a nationwide chain. The question worth asking is whether the design paints us into a corner when that changes, and mostly it doesn't, for the same reason everything above works.
The API is stateless. Auth is a JWT, tenant context is rebuilt per request, and the one piece of in-process state, the in-memory cache fallback, is something we already treat as disposable. So the first move at scale is the boring one: run several API instances behind a load balancer. Nothing in the request path assumes it's the only copy. The cache fallback going per-instance is the only wrinkle, and since a cache miss just falls through to Postgres, caches that disagree across instances cost latency, not correctness.
Workers are the same story, and this is where BullMQ comes in. BullMQ is built to run many workers against one set of Redis-backed queues, so adding worker processes is a deploy change, not a rewrite. The idempotency we needed for reliability pays off a second time here: because every job is keyed and every handler is safe to run twice, two workers racing on the same queue can't double-charge anyone. The step after "more workers" is splitting queues by job type, putting payments on their own pool and letting notifications and report generation share another, so a Monday-morning flood of scheduled reports can't starve a payment webhook.
It's tempting to read "more load" as "time for Kafka," but that conflates two different tools. We'll keep BullMQ for jobs; Kafka only earns its place if the shape of the problem changes, not the volume. BullMQ is a job queue: do this work, retry it, back off, dead-letter it. Kafka is a durable, replayable event log that many independent consumers read at their own pace. They look similar from across the room but solve different problems up close. You don't reach for Kafka because you have more jobs; a gym platform is orders of magnitude below the throughput where BullMQ on Redis starts to strain. You reach for it when you want an event backbone, where every check-in, payment, and personal record becomes an event that several systems replay on their own: analytics, a warehouse, a materialized read model, maybe ML features later. The day we actually want that, Kafka or Redpanda is the right tool. Until then it's operational weight with no payoff, and adopting it early would be architecture theater dressed up as foresight.
The real ceiling isn't the queue anyway. It's Postgres, the way it usually is for multi-tenant SaaS. The levers there are well-worn, and we'd pull them roughly in order: read replicas so reporting and analytics stop competing with live writes, then time-partitioning on the big append-only tables (check-in sessions, exercise sets, audit logs), and eventually tenant sharding, routing organizations onto separate database clusters, if one primary stops being enough. That last one is less frightening here than it usually is, for a specific reason. IDs are generated in application code as UUID v7, globally unique with no central sequence, so moving a tenant's rows to another cluster can't cause an id collision. The same instinct that keeps identity off a central authority for the offline path shows up here too. Nobody planned that; it's just what generating your own ids gets you.
None of this is a roadmap with dates on it. It's a claim about order: the cheap, boring scaling moves are all available to us before any of the expensive, interesting ones become necessary, and the plan is to keep resisting the interesting ones until the problem actually asks for them.
The through-line
None of these pieces is clever on its own. IndexedDB, RS256 tokens, a Redis-backed job queue, integer money, a serial port. It's ordinary infrastructure, and a lot of teams would reach for the same parts.
What shaped the system was one rule we decided not to bend: the things a gym leans on minute to minute shouldn't stop working when the network does. Once that's non-negotiable, a lot of the architecture stops being a choice. The kiosk has to accept the check-in on its own, so each scan needs an identity the kiosk can assign before the server sees it. That client-owned key only stays safe if the server treats it as idempotent, so the server dedupes on it. Redis can't be allowed to take the API down with it, so in its cache role a miss and an outage look identical to the caller. Each decision is mostly forced by the one before it.
What I didn't expect was how often a decision made for reliability paid for something else too. Generating our own UUID v7 ids in app code was partly about the offline path. It's also what would make tenant sharding safe years from now, since globally unique keys with no central sequence can't collide when a tenant moves clusters. The idempotency that stops a member being double-charged on a redelivered webhook is the same property that lets us run ten workers instead of one. You don't often get to answer today's problem and next year's with one decision. We did here, not by planning it, but because "make every step safe to repeat" turns out to be good advice at every scale.
So the honest summary isn't a feature list. We committed early to one uncomfortable constraint, that the network won't be there when it matters most, and let it settle a lot of arguments we'd otherwise still be having.