Usage Billing Review

Credit Wallet Data Models for AI SaaS Products

Model credit wallets as immutable ledgers, not mutable balances.

Staff Writer · · 12 min read
Cover illustration for “Credit Wallet Data Models for AI SaaS Products”
Credit Wallets and Prepaid Billing · September 7, 2026 · 12 min read · 2,742 words

A wallet looks like a balance field from the outside: a number that rises when someone buys credits and falls when someone spends them. Underneath, it's a graph of related entities, each with its own constraints and its own lifecycle. Model it as a single mutable number, and six months later nobody can explain a disputed charge; by then the schema is load-bearing enough that fixing it means a migration, not a patch.

The wallet itself is the top-level container. It belongs to a customer, a team, or an org, depending on how the product scopes access, and it carries metadata: status, currency, allocation rules. Below that sits the grant, a discrete credit issuance event with a source (purchase, trial, refund, promotional), an amount, an issue timestamp, and an expiration timestamp. Below the grant sits the ledger entry, an immutable record of every debit or credit against the wallet, and that entry is the append-only record everything else derives from.

Balance is a derived, queryable view computed from the ledger. It gets recalculated or cached, but it is never ground truth in its own right. Then there's the entitlement check, the read path that answers a specific question at request time: given this action, does the wallet have enough balance, and which grants will it draw from? Finally, an expiration schedule governs when grants lapse, in what order they get consumed, and what happens to unspent credits when a billing cycle closes.

The relationships between these entities carry as much weight as the entities themselves. A wallet can hold several active grants at once with different expiration dates, so every ledger entry has to reference the specific grant it draws down, not just the wallet in aggregate. A customer can hold multiple wallets scoped to different products or tiers, so an entitlement check has to know which wallet it's even querying before it can answer anything.

Most first-pass implementations collapse this graph into one or two tables: a wallet with a balance column, maybe a transactions table bolted on afterward. That design looks fine in a demo and breaks under concurrency, under audits, under any scenario involving expiration or multiple grant sources, which is every scenario an AI product eventually faces. The debt comes due the first time finance asks for a reconciliation report, and it comes due at the worst possible moment.

The ledger as the only reliable source of truth

The instinct, especially early on, is to store balance as a mutable integer and update it in place each time a transaction happens. It feels simple, and the trouble doesn't surface until the system is under load or under dispute, which is exactly the worst time to discover it.

A mutable balance column creates three problems with no clean workaround. Concurrent writes race: two API calls hit at nearly the same instant, both read the same starting balance, both pass the sufficiency check, both debit, and the wallet goes negative even though the code never intended to allow overdraft. There's no audit trail, so when a customer disputes a charge, or an internal team tries to reconcile a discrepancy, nothing records what happened, in what order, or why the number is what it is. And there's no replayability: if a downstream billing system fails partway through processing, nothing exists to reconstruct the correct state, because the log of individual events was never kept in the first place.

The fix is to treat every debit and credit as an immutable append to a ledger table, with balance computed, or cached with careful invalidation, from that log rather than stored as a fact in itself. Ledger entries are never updated or deleted; a correction is a new entry, a reversal record, not an edit to history. Each entry carries a wallet ID, a grant ID indicating which grant it draws down, an amount, a direction, a timestamp, an idempotency key, and a reference to whatever event triggered it. The idempotency key matters more than it looks: most upstream systems deliver usage events with at-least-once semantics, meaning the same event can arrive twice, and the ledger has to deduplicate on write or risk double-billing a single action.

This is the same principle behind double-entry bookkeeping, a discipline that predates computing by centuries and exists precisely because financial state has to be reconstructable and disputable. Treating balance as a view over the ledger rather than a stored fact is what makes audits possible: the balance at any past point in time can be recomputed on demand, which is exactly what a dispute or a compliance review requires.

Grant design: why not all credits are the same row

A grant is a bounded, discrete issuance of credits, and it needs its own object, not a fold into a generic balance top-up. Treating every credit addition the same way, as just a number added to a pool, throws away information the business needs later. This is the mistake that shows up hardest at renewal time, when nobody can say which credits were promotional and which were paid, and finance is left guessing at deferred revenue.

Different grants expire on different schedules. A promotional grant might lapse in 30 days while a purchased grant holds for a year, and a flat balance has no way to represent that distinction. Consumption order matters too, both operationally and sometimes contractually. FIFO, consuming the oldest-expiring credits first, is the friendliest default for most customers, but some enterprise contracts specify LIFO or a pro-rata draw instead, so the model needs a configurable priority per wallet rather than a hardcoded rule. Rollover policy attaches to the grant, not the wallet, since some grants carry unused credits into the next period and others forfeit them at cycle end. And revenue recognition ties to grant consumption, not grant issuance; prepaid credits sit on the books as a liability until they're actually spent, a distinction that matters to finance even if it's invisible to the customer.

A workable grant schema needs, at minimum, a grant ID, wallet ID, source, amount issued, amount remaining (a cached view, never the authoritative figure), issue timestamp, expiration timestamp, rollover policy, and priority. With that in place, a product can support scenarios a flat balance simply cannot handle: trial credits that expire in two weeks, clearly separated from paid credits good for a year, promotional top-ups configured to burn down before purchased credits touch, enterprise committed-spend credits that roll over annually under whatever the contract specifies.

Skip first-class grants, and these scenarios get bolted on in the application layer instead, as conditional logic scattered across the codebase. That patchwork is fragile, hard to test, and invisible to anyone in finance trying to reconcile what actually happened. It's also, plainly, the single most common reason teams end up rebuilding their billing layer eighteen months after shipping the first version. Nobody plans for that rebuild. It just arrives, usually right when the company can least afford the distraction.

Expiration logic and the edge cases that break naive implementations

Expiration sounds like a one-line rule: credits past their date are gone. In practice it forces a handful of decisions that have to be made explicitly, before the first line of code, or the system will make them inconsistently by accident. Inconsistency here is what turns into a support ticket, and eventually a churned account.

Consider a request that lands at 23:59:59 on a grant's expiration date, with processing that takes 200 milliseconds. Is that grant still eligible when the request lands, or only if it finishes before the clock turns over? The defensible answer, and the only one worth building, is to evaluate expiry at request receipt, not at ledger write, and to enforce that rule consistently everywhere the check happens, not just in the common path. Anything less consistent is a bug waiting for an angry customer to find it.

Partial consumption at the expiry boundary raises a related problem. A grant might have credits remaining exactly when it expires, mid-debit, so the ledger needs to close it out cleanly with an expiration entry that zeroes its balance and carries a timestamp and a reason code, rather than leaving a dangling remainder that never gets accounted for.

Expiration also has a UX dimension the data model has to support directly. Telling a customer "you have thousands of credits expiring in 9 days" requires expiration timestamps that are queryable per grant, not buried inside some opaque blob column. Products that surface this kind of expiration state as a visible UI element have made it a baseline customer expectation, and any system that can't query expiration per grant will struggle to meet it.

Rollover introduces one more fork, and it's the one most teams get wrong by taking the easy path. When a grant rolls over, does the system create a brand-new grant with its own ID and a fresh expiry clock, linked back to the original through a parent grant reference? Or does it just extend the existing grant's expiration date? Extending is simpler to build, but it destroys the period boundary and muddies the audit trail; creating a new, linked grant preserves full lineage, at the cost of slightly more bookkeeping, and holds up under scrutiny later. The extend-in-place shortcut is the wrong call almost every time. Anyone tempted by it should picture explaining, a year in, why last January's grant never actually closed.

Expiration itself should run as a scheduled background job, separate from the request path entirely. Billing logic has no business slowing down the operation a customer is actually waiting on.

Entitlement checks at millisecond speed

An entitlement check answers one question before an AI action runs: does this wallet have enough credit to allow it? For AI products specifically, this check sits directly on the critical path, ahead of the LLM call itself, and LLM calls are already expensive and often already slow. Tacking hundreds of milliseconds of billing overhead onto that is not a tradeoff most products can absorb.

The naive approach, querying the full ledger, summing every non-expired grant, and comparing that sum against the requested amount, is correct and too slow under real load. Correctness alone doesn't earn a place in the request path; if it did, nobody would need caching at all. The pattern that actually holds up in production separates the check from the settlement into two phases. First, pre-authorization: atomically reserve the estimated credit amount the moment the request starts, marking those credits as tentatively held without debiting them outright. Second, settlement: once the operation finishes, debit the actual amount consumed, release any excess reservation, or release the entire hold if the operation failed. It's the same shape as authorization and capture in payment processing, applied to credits instead of dollars.

The atomicity of that reservation step is not optional, and treating it as an implementation detail is how wallets go negative in production. Without it, two concurrent requests can both check the balance, both see enough headroom, both proceed, and the wallet ends up negative even though every individual check passed. The reservation needs to be atomic at the database level, through a row-level lock or an optimistic concurrency check with retry, not through an application-level flag that two processes can both flip at once.

Speed comes from caching, not from skipping the check. A cached view of available balance, updated on every ledger write and served from a fast read store, gets the entitlement check down to sub-millisecond response times. The catch is that cache invalidation has to happen synchronously with the ledger write. If the cache and the ledger fall out of sync even briefly, the entitlement check approves requests it shouldn't.

Streaming and other variable-cost operations complicate this further, since the actual cost of an LLM response isn't known until the response finishes generating. The workable pattern reserves a configurable maximum at the start of the operation and settles to the true cost at the end. If actual usage exceeds that reserved maximum, the system either blocks the stream mid-flight or allows the overage and flags it for review, depending on account policy.

Multi-wallet and scoping: when one wallet per customer isn't enough

Early-stage products tend to give each customer a single wallet, one balance, one meter. It works fine at small scale and stops working the moment an enterprise customer's structure gets more complicated than a single account, which for most B2B products happens sooner than the team expects, usually right around the first six-figure contract.

A customer might buy credits for two separate products, and those credit pools should not be interchangeable unless the contract explicitly says so. An enterprise parent account might hold a shared credit pool while its subsidiary accounts each get their own wallet drawing against a budget allocated from that parent. Trial credits for a new feature need to stay isolated from purchased credits for the core product, so a customer testing something new doesn't accidentally burn through paid balance, or vice versa.

A single shared credit pool across all of a platform's AI capabilities is worth noting as a deliberate product decision rather than something the underlying technology forces. Building that kind of universal pool still requires the billing model to enforce it explicitly; it isn't the default that falls out of having only one wallet type available. Teams that assume a single wallet is the simple option and multi-wallet is the complex one have it backwards: a single wallet is just a multi-wallet system that hasn't hit its first enterprise contract yet.

Supporting this properly means the schema needs a wallet scope, indicating which product, feature, or billing metric a given wallet applies to, and a wallet priority, determining which wallet gets drawn from first when a request could plausibly pull from more than one. Parent and child wallet relationships need to exist for org hierarchies, where a child wallet draws down its own balance and, depending on configuration, optionally overflows into the parent's pool once exhausted.

Entitlement checks in this kind of multi-wallet setup have an extra step before they can even look at balance: resolving which wallets are eligible for the request in the first place. That resolution is policy, not data, and it needs to be configurable without requiring a schema migration every time a new scoping rule shows up. Auto-top-up rules follow the same logic, scoping to individual wallets rather than to the customer's account broadly, so a top-up triggered by wallet A falling below its threshold issues a new grant to wallet A specifically, not to some account-level bucket that might not even be the one running low.

Consistency guarantees the model requires and how to enforce them

A credit wallet is a financial system in miniature, and the consistency bar for it is the same bar applied to any ledger controlling access to paid-for resources. There's no lighter version of correctness available just because the unit is called a credit instead of a dollar. Teams that treat it as a lesser standard, on the theory that "it's not real money," are the ones who eventually have to explain a negative balance to an auditor. "The cache was stale for eleven seconds" is not an explanation an auditor accepts.

The first non-negotiable guarantee is no double spend: concurrent debits, taken together, must never exceed what's actually available. Enforcing this means serializable isolation on the reservation step, or optimistic concurrency with conflict detection and retry when serializable isolation isn't practical at scale. Advisory locks scoped per wallet ID are a reasonable middle ground for databases that can't sustain full serializable isolation under heavy concurrent load.

The second is idempotency on write. Usage events arriving from distributed systems carry at-least-once delivery guarantees almost by default, which means the same event can and will arrive more than once eventually. Every ledger write path needs to accept an idempotency key and deduplicate against it, so inserting the same event twice produces exactly the same ledger state as inserting it once. That key has to live in a durable index, not in application memory, or a service restart quietly reopens the door to double-counting.

Neither guarantee is exotic, and neither should be treated as optional scope for a later release. Both are standard requirements for any ledger-backed system handling money or its equivalent, and an AI product's credit wallet is exactly that kind of system, whether or not the product team building it originally thought of it in those terms.

Sources

  1. kinde.com
  2. dodopayments.com
  3. docs.tigerbeetle.com
  4. flexprice.io
  5. flexprice.io

More in Credit Wallets and Prepaid Billing