Known gaps

Risk Register & Hardening Roadmap

This page is the honest list. Everything below was found by us, is reproducible, and did not get fixed before the delivery date — either because it needed a decision we could not make alone, or because the clock ran out. The quality gates are green because they do not cover these items, so we are listing them explicitly rather than leaving the green build to speak for the whole system.

Where this comes from. A full internal audit of the three repositories (backend 85d4f7b, contracts 640fb44, mobile 83de7f1), run on 12 August 2026 — manual review plus tooling. Severities below are ours: high means it weakens a control that already exists, medium is a real defect with a workaround, low is hygiene. Nothing here is rated critical — no item on this page moves value to the wrong account or exposes a key. Items marked accepted for the demo are deliberate scope decisions, listed with their exposure.

Where the build actually stands

The quality gates that do run are green, and it is worth being precise about what they prove and what they do not.

Repository What passes What is not covered
backend 629 tests / 58 suites, 0 failures; branch coverage 84.32 % against an 84 % gate; tsc, lint and npm audit clean, on every push and pull request No integration tests at all — test/ holds fixtures and not one *.spec.ts, and the testcontainers dependency is installed but unused
contracts 136 tests including invariants, 100 % line/branch/function coverage of src/, slither clean across 93 detectors, gates on every push Nothing at the code level. The exposure is operational — one EOA governs the deployment
mobile 357 tests / 41 suites, 90 % coverage threshold, tsc clean, lint clean (6 warnings) No dependency check in CI, and 13 high-severity advisories sit in the build chain today

Blockers — must be fixed before the system holds real value

Rate limiting counts every client as one, behind the proxy high

ThrottlerGuard is registered globally and the per-route limits are set, but Express is never told to trust the reverse proxy (app.set('trust proxy', 1)). Deployed behind Traefik/dokploy, every request arrives with the proxy's IP, so all clients share a single bucket. The effect is the inverse of the intent: "10 logins per minute" becomes 10 logins per minute for the whole platform, and anyone can lock everyone out of sign-in with ten requests. It also removes the only brake currently standing between an attacker and the address grab above.

Why it is still open: found late, during the audit, after the throttling work was already merged and looked complete. The fix is one line plus a decision on the hop count; a proper fix also moves the throttler store to Redis, since in-memory buckets multiply by the number of replicas.

Demo signer addresses ship in a migration high

The issuer, relayer and guardian addresses are hard-coded in 1786200000000-SeedSigners.ts and seeded into every environment, production included. They belong in configuration with a key ceremony behind them, not in schema history.

Accepted trade-offs

Deliberate scope decisions, not oversights: we know the exact exposure, the exposure is bounded, and closing it needs a product policy call rather than a patch.

Address linking is first-come, ownership is proven at claim time medium accepted for the demo

POST /wallets records whatever address the caller sends; proof of ownership is enforced later, in the claim flow (confirmOwnership(), ERC-1271 / ecrecover against the EIP-712 digest). The important consequence of that ordering is what it is not: registering someone else's address does not let you claim their cashback, because the claim still requires a signature you cannot produce. Value cannot move to the wrong account through this path.

What it does allow is denial of service on the link itself. Register an address first and the legitimate owner gets already linked; there is no release or reassignment endpoint, so their cashback is unreachable until an operator intervenes. The mobile client states the same thing rather than hiding it — walletLinking.ts carries the comment "there is no reset endpoint", compares the linked address, and blocks payment on a mismatch, so the user sees a hard stop instead of a wrong balance. The endpoint is throttled at 10 requests/minute, which raises the cost of doing this at scale without removing it.

Why we accepted it. The griefer needs to know the victim's address ahead of them and gains nothing but the nuisance — no funds, no rewards, no account access. Closing it properly is a signature challenge on linking (the primitive already exists in the claim flow) plus a reassignment path, and the reassignment path is the real work: it needs a policy on who may take an address back, on what evidence, and what happens to coupons already accrued against it. That is a product decision, and the codebase is frozen for delivery — so the honest state is a named, bounded, reversible trade-off rather than a rushed rule we would have to redo. It is the first item in the post-delivery batch below.

Deployment perimeter

The application-level hardening landed — production mode is enforced at startup, dev-only routes are compiled out, helmet and CSP are strict. What did not land is everything around the container.

Gap State today Size of the fix
Swagger is public medium setupSwagger() is called unconditionally, so /docs is served in production One conditional
CORS defaults to * medium CORS_ORIGINS has '*' as its default value, so a missing variable opens the API to every origin rather than failing Remove the default, require the variable
Migrations run on boot medium migrationsRun: true for all five service roles — five processes race to migrate on deploy, and a bad migration takes the deploy with it Separate release step
Container runs as root medium No USER directive in the Dockerfile Two lines
Secrets are not split by role medium .env.issuer.example and friends exist, but docker-compose.yml mounts one shared ./.env into all five containers — so settlement, which needs no keys, receives both signing keys Templates are already written; wiring them into compose is the remaining work
Unsafe flag combinations are still valid medium NODE_ENV=production together with ENABLE_DEV_TEST_TOKEN=true or ALLOW_PLAINTEXT_SIGNING_KEY=true passes validation. Separately, the dev-route flag is read from process.env at module load, before .env is parsed — so setting it in the file has no effect, only a real environment variable does. The direction is fail-closed, but the behaviour is undocumented One .superRefine() in the env schema
UTILITY_TOKEN_CONTRACT_ADDRESS is the zero address low .env.example ships 0x000…000 while the deployed UTL is 0x63dE56C3…. It passes validation and fails later, inside the supply reconciler — breaking exactly the control that detects minting outside the pipeline Paste the real address; reject the zero address in requireAddress()

Backend defects we know about

Claim idempotency is broken by the fix that moved the network call medium

Ownership verification was correctly moved out of the database transaction — but the challenge load moved with it. Because loadChallenge() rejects a consumed challenge, a retry with the same Idempotency-Key now fails with CHALLENGE_INVALID before it ever reaches the idempotency layer. The replay path is unreachable in precisely the scenario it was built for: the client timed out and retried. The app masks it by generating a fresh key per attempt, so it is not visible in normal use — but the guarantee the endpoint advertises is not there.

Secret access is no longer logged medium

SecretsService used to write a security_event on every read, write and delete of an encrypted seed blob — misses included, because a burst of misses is the shape of an enumeration attempt. Those three log lines were removed during a logging cleanup and not replaced. Today /secrets/* leaves no trace beyond the request log. That was the only detective control over the most sensitive resource in the system, and losing it was not deliberate.

A malformed indexer response is indistinguishable from "no payments" medium

The indexer response schema was relaxed to z.unknown() so that one broken network/token pair cannot take down the nine healthy ones — right goal. But the implementation collapses "this pair is not served" and "the response did not parse" into the same answer: no transfers. If the provider's schema drifts, payments are skipped silently and coupons are never accrued. The monitor's sampler catches it eventually via monitor.payment_not_indexed, so there is a net — but it is sampled and delayed, and this is the one failure mode that under-pays users without an error anywhere.

Others medium low

Item Detail
Cooldown counts expired claims nextClaimAt() filters on status <> FAILED, so an EXPIRED coupon still holds the user in cooldown
State lives in process memory Nonce manager, throttler store, balance-refresh flags, alert deduplication and the live-price cache are all in-process — each of them behaves differently the moment there is a second replica
Beta dependencies are not pinned Three WDK packages on the money path sit on ^1.0.0-beta.*, so a beta release can land in a build unreviewed
Unbounded growth refresh_tokens rows are never pruned; the live-price cache never evicts and its key is caller-supplied
Local npm test needs exported env vars Several suites bootstrap config through validateEnv(process.env), so the pre-push hook fails on a clean checkout and gets bypassed
Address normalisation bypassed in two places Raw toLowerCase() instead of canonicalAddress(); the Tron branch of ownership verification still recovers with an Ethereum prefix
CLAUDE.md is stale Still describes the scaffold phase, ~50 files and a 20 % coverage threshold

Mobile and the client/server seam

Gap Why it matters
Sign-out does not call POST /auth/logout medium The backend gained refresh-token family revocation and a logout endpoint; the client never wired it up. signOut() clears the keychain locally, and the refresh token stays valid server-side for up to seven days — so on a shared or lost device, signing out does not end the session
Session is stored without an accessibility constraint medium The backup key is written with WHEN_UNLOCKED_THIS_DEVICE_ONLY; the session tokens are written with no options, so they are eligible for device backup and restore elsewhere. The refresh token opens the encrypted seed blobs — it deserves at least the protection the key it guards has
13 high advisories in the build chain medium All transitive through metro (image-size parser DoS), so they affect the bundler and not shipped app code — but npm audit fix --force wants to downgrade React Native to 0.72.17, so there is no clean upgrade today. The mobile CI runs lint, typecheck and test, and no dependency check at all
TronGrid credentials are bundled into the app low TRON_API_KEY/TRON_API_SECRET are inlined by react-native-dotenv and extractable from the APK/IPA. For a paid key that is a credential leak, not configuration — it should be proxied through the backend
No build-time check that the release API URL is HTTPS low The default is http://. Both platforms fail closed (Android cleartext disabled, iOS ATS), so it cannot ship insecurely — it can just ship broken, and a build-time assertion is cheaper than debugging it in a release
Leftover demo config low export const anySecret = ANY_SECRET and the matching .env.example entry are dead code with a name that reads like a finding in any grep for "secret"

Smart contracts

The contracts are the strongest part of the delivery — 100 % coverage, clean static analysis, invariants under test. Both gaps here are about operations and edges, not the core logic.

The deployment is governed by a single EOA low by design, for the demo

governor resolves to the deployer: 0x934d57BC…, a plain external account. That key holds DEFAULT_ADMIN_ROLE on both CouponClaim and UTL, plus owner and the LayerZero delegate. With it, one could grant itself issuer and relayer roles, set the threshold to 1, raise the caps and mint arbitrary UTL. The whole defence-in-depth story — threshold signatures, per-claim caps, epoch caps — reduces to the safety of that one private key.

This is a deliberate demo choice, flagged in the deployment config and printed by the deploy script itself, which states that production wants a multisig plus a timelock. We are listing it anyway because for a contract deployed to a public network it is the single largest on-chain risk, and it should be named rather than inferred. Related: UTL is a LayerZero OFT, so owner can call setPeer at any time, after which the receive path mints without MINTER_ROLE. The deploy script verifies no peers exist — but only at deploy time. Nothing monitors RoleGranted, CapsUpdated or PeerSet today; the watcher listens only for Claimed.

Two edges in CouponClaim low

Deploy.validate() rejects a configuration where perClaimCap > epochCap — correctly, since the epoch cap binds first and the per-claim value is then dead config. The live setCaps() setter accepts that same combination; it only checks for a zero length. An operator can make the mistake on a live contract exactly as easily as in a config file.

Separately, claim() accepts amount == 0. The call passes every check, permanently burns the paymentRef in nullifierUsed and emits Claimed without minting. The backend never sends zero — but paymentRef is the anti-replay primitive, and spending one for nothing should not be reachable.

The threshold is real, but K = N = 1

Worth saying plainly, since "K-of-N threshold attestation" appears throughout this documentation: the mechanism is built, tested and correct — sorted signers, deduplication, 1 ≤ K ≤ N held across every role change, the threshold refusing to drop below the signer count. The deployed configuration runs K = 1, N = 1. So today the guarantee in practice is one issuer key, with the relayer as the only second barrier. Raising it is a configuration change and a key ceremony, not code — but until that happens, the demo's real property is 1-of-1 and we would rather state the numbers than the shape.

The order we would fix them in

# What Why first
1 trust proxy One line, and it makes the rate limiting that already exists mean something
2 Claim idempotency, secret access logging, signers out of the migration, then ownership proof on linking with its reassignment policy The set that has to be closed before the system holds value belonging to anyone but us. The linking change comes last in this batch because it is the only one waiting on a product decision rather than on code
3 Deployment perimeter Swagger, CORS, migrations as a release step, non-root, the real UTL address, per-role env files. Mostly configuration, and the templates are already written
4 Client and seam Wire up logout, constrain the keychain entry, add a dependency check to mobile CI
5 Contracts Multisig and timelock on the governor before anything of value is minted; the two setter edges alongside it
6 Operability Integration tests, shared state into Redis, pinned beta packages, indexer failure distinguished from an empty result

What we did not do, and would not claim. There has been no dynamic or load testing, no live exercise against Sepolia beyond reading the recorded deployment artefacts, and no source-level audit of the WDK packages themselves. The list on this page is what static review of our own three repositories found — it is not a certificate that nothing else is there.