Usage Billing Review

Backfill and Late-Arriving Event Handling in Metering Systems

Duplicate events and late-arriving data require structural safeguards to prevent billing errors.

Senior Contributing Editor · · 12 min read · Updated
Cover illustration for “Backfill and Late-Arriving Event Handling in Metering Systems”
Usage Event Metering and Aggregation · August 21, 2026 · 12 min read · 2,631 words

The hardest problem in production metering is not throughput, it is correctness. Duplicate events, late arrivals, and retroactive corrections each cause billing errors that no amount of ingestion capacity will fix. An engineer tasked with building a metering pipeline will often reach for the pattern that already feels familiar: fire usage events to a queue, aggregate them overnight, push the totals to a payment processor. That approach resembles a logging pipeline more than a financial system, and it tends to survive staging only to break once real clocks, real networks, and real retries get involved.

The failure modes are specific and compounding. A duplicate event produces a double charge. A clock that drifts even slightly can push an event into the wrong billing window. A retroactive correction, if there's no model for handling it, leaves the ledger wrong permanently rather than wrong temporarily. None of these are edge cases in the statistical sense, they are structural risks built into any system that bills customers based on measured usage rather than a flat fee.

The distinction that matters is between an analytics pipeline and a financial one. An approximate count, an eventually-consistent total, or a silent retry are all acceptable tradeoffs when the output feeds a dashboard. None of them are acceptable when the output feeds an invoice. Usage metering has to clear a categorically higher bar than the data infrastructure patterns it superficially resembles: the cost of an error isn't a wrong number on a chart but a wrong charge on a customer's card.

The Cursor incident from June and July 2025 is a compact illustration of what happens when this bar isn't met. A developer on an annually-billed plan with uncapped usage received a four-figure invoice, a direct consequence of a plan structure with no spend controls built in. Cursor's own CEO, Michael Truell, described the episode as a failure to communicate the pricing change clearly, and the company offered refunds for unexpected charges incurred in that window. When spend controls and billing logic are left unmodeled, the gap becomes visible at the worst possible moment, in a customer's inbox, as a bill they didn't expect.

The schema decision that determines whether late events can be handled at all

Diagram: Event Time vs. Receipt Time: The Schema Split Everything Depends On. Visualizes: Show the contrast between two timestamp fields — occurredAt (client-side wall clock, when the billable action happened) and receivedAt (server-side, when the…

Before idempotency, before aggregation windows, before any backfill procedure, there is one schema decision that any of those mechanisms depend on: the separation of event time from receipt time. Conflating event time and receipt time into a single timestamp field is a mistake teams typically discover only when they try to handle their first late-event edge case, at which point fixing it requires an expensive backfill rather than a configuration change. One is occurredAt, the client-side wall clock reading of the moment the billable action actually happened. The other is receivedAt, the server-side timestamp set the instant the system observes the event at the ingestion boundary. These are not interchangeable, and treating them as if they were is the single most consequential mistake a metering system can make early in its design.

Billing window assignment has to be computed from occurredAt, not receivedAt. If a pipeline assigns billing periods based on when an event was received rather than when it happened, a late-arriving event gets silently filed into the wrong billing period. This produces an undercharge with no error, no alert, and no visible symptom, until someone runs a reconciliation and finds the numbers don't add up.

Teams that conflate the two fields rarely discover the problem by design review. They discover it the first time they need to handle a late-arriving event and realize the system has no way to do so without rewriting historical records. At that point, a straightforward configuration change becomes an expensive backfill operation, performed under pressure, against production data, because the schema never made the distinction in the first place. Every mechanism described in the rest of this piece, idempotency, windowed aggregation, bounded backfill, and signed corrections, assumes this separation already exists in the schema. None of them can be retrofitted cleanly onto a system that merged the two timestamps from the start.

Idempotent event IDs and duplicate billing prevention at the ingestion layer

With event time and receipt time properly separated, the next correctness problem to solve is duplication, and it has to be solved structurally rather than through convention. Distributed systems do not stay correct because engineers agree to be careful. They stay correct because the architecture makes incorrectness impossible, and that means idempotency enforced through deterministic, client-generated keys with a true unique constraint in the event store.

A common but insufficient fix is to have the server assign a UUID to each incoming event. If the client retries after a network timeout, the server has no way to know the retry represents the same logical event as the original request. It assigns a new UUID to what is logically the same event, and both records persist.

The pattern that actually works inverts where identity gets assigned. The client generates the idempotency key deterministically, for example by computing SHA-256 over the tenant ID, the resource, the timestamp, and a nonce, so that the same logical event always produces the same key regardless of how many times it's retried. The event store enforces a unique constraint on that key. When a duplicate write arrives, the store rejects it silently and returns an HTTP 200 with the original result, instead of an error that would prompt the client to retry and risk a retry storm. The division of labor here is load-bearing: the client owns the identity of the event, and the server owns the enforcement of uniqueness. Neither side can do the other's job.

This same requirement doesn't stop at the ingestion boundary. It extends to the sync layer between the internal billing system and any external payment processor. Every usage record written across that boundary needs an idempotency key tied to the internal aggregate ID, because a retry that drops or varies the key can write the same usage record twice to the processor and produce a double charge on the customer's statement. Idempotent ingestion, applied consistently at both boundaries, also absorbs a class of operational failures that would otherwise cause real damage: upstream services that retry aggressively under load, deployment rollouts that accidentally replay an event queue, and client-side bugs that send the same event multiple times all become harmless, because the store simply discards the duplicates before they reach the ledger.

Windowed aggregation and the bounded late-arrival window

Idempotent ingestion solves duplication at the level of a single event. The next layer of the architecture has to decide how individual events become the totals that actually appear on an invoice, and that decision is where the system either stays correctable or becomes brittle.

Aggregating raw events at query time, computing the bill fresh from the full event history every time a billing run executes, is a dangerous default. It means every billing run is a full recompute over a growing and effectively unbounded dataset. It means any schema change invalidates every prior result, since there's no stable intermediate artifact. And it means there's nothing for an auditor or an engineer to point to and check, because the "aggregate" only ever exists transiently, inside a single query.

The pattern that holds up in production is pre-aggregation into immutable time-window buckets. An aggregation worker runs on a schedule, queries events by occurredAt for a defined window, computes totals per metric, and upserts the result into a separate aggregate store. The upsert is what makes this safe: re-running the aggregation over a window that's already closed produces exactly the same result each time, which is what makes the aggregation itself idempotent.

A design still has to decide how long a window stays open before it's treated as final. Rather than freezing a billing partition the instant its period ends, a pipeline can keep a bounded recent window open, for example the last three days, and recompute that window on every run to absorb any events that trickle in slightly late. A formal backfill is reserved for corrections that land after that window has already closed. The policy that governs the boundary of that window should have a concrete threshold attached to it: events arriving more than 24 hours late relative to receivedAt should trigger a manual review flag rather than silent acceptance into the aggregate. That 24-hour threshold is the direct policy expression of the occurredAt/receivedAt separation established earlier, it gives the schema decision an operational consequence rather than leaving it as an abstract design preference. In practice, mature pipelines adopt a deliberate asymmetry: recent days are allowed to change, older days are frozen, and a frozen day is reopened only when the correction it requires is large enough to justify the operational cost of doing so.

Backfill versus replay

Once a late correction falls outside the bounded recent window, the operation required to fix it is a backfill, and conflating event time and receipt time is what turns that fix from a configuration change into an expensive backfill operation.

Replay re-emits historical events exactly as they were, without any transformation applied. Backfill is controlled reprocessing: it applies a transform, a corrected schema, or a pricing fix to historical data, and it does so with verification and a bounded scope around the correction. Backfill is defined as the controlled process of reprocessing or filling missing data, events, or state into a system after a gap, a delay, or a schema change, and the word that carries the most weight in that definition is "controlled." A backfill is reproducible, observable, and auditable. It is not an ad hoc manual fix applied under pressure and left undocumented.

Several situations call for a backfill rather than a simple replay: pipeline downtime that left a gap in the record, a transformation bug that corrupted values on the way into the aggregate store, a schema change that leaves older records incomplete relative to the new structure, late arrivals that missed the rolling window entirely, and targeted corrections scoped to a specific date range for business reasons. Historical data carries dependencies that a naive replay ignores: the code version that originally processed it, the upstream contracts in effect at the time, the partition logic that determined where it landed, and assumptions baked into the original load that may no longer hold. A safe backfill accounts for all of that context. A replay, by definition, does not.

A team that finds itself running frequent backfills isn't just exercising a recovery tool, it's revealing an observability gap.

The four-phase framework for running a backfill safely in production

Diagram: Four Phases of a Safe Production Backfill. Visualizes: Visualize the four sequential phases for running a backfill safely: Phase 1 — Isolation & Scoping (determine damage extent, scope creep is the most common cause of hours becoming a…

Running a backfill against a live production system carries real risk, because that backfill will compete with ordinary production pipelines for the same compute and the same storage unless it's deliberately isolated from them. A backfill with no isolation, no rate limiting, and no staged validation will either degrade the production system it's running alongside or fail outright once it hits resource contention. Four phases keep that risk bounded.

The first phase is isolation and scoping. Before any code runs, the team has to determine how far back the damage extends, because a bug that looks localized at first glance often turns out to have touched more partitions, more downstream consumers, and more unstated assumptions than the initial assessment suggested. Scope creep discovered mid-execution is the most common reason a backfill that should have taken hours turns into a multi-day incident.

The second phase is development and testing. The correction logic, whether it's a new schema, a pricing fix, or an adjustment delta, gets applied against a bounded historical slice in a non-production environment first. Before any of it touches real data, the team verifies that re-running the backfill produces the same result each time, an idempotency test performed specifically on the correction logic itself.

The third phase is execution and monitoring. The backfill runs under rate limiting and resource quotas so it doesn't crowd out the live pipelines running alongside it, and it emits its own metrics, reprocessed event counts, latency, and reconciliation deltas, so the operation stays observable rather than opaque. A rollback condition is defined before the run starts, not improvised after something goes wrong.

The fourth phase is validation and swap. The backfilled output gets compared against the expected state using row counts, aggregate totals, and reconciliation deltas, and only once that validation passes do downstream consumers, the billing engine, the invoice generator, customer-facing dashboards, get pointed at the corrected data.

The PostHog case is a concrete demonstration of what happens when this infrastructure simply doesn't exist. The flag-serving gate for behavioral cohort evaluation required a column called last_backfill_events_at to be set before it would open, but no production process ever wrote that column. The gate stayed permanently closed, and cohort evaluation was broken in production, silently, with no error surfaced anywhere in the system. The four phases above aren't bureaucratic overhead layered onto a simple rerun; they're what resolves a backfill cleanly instead of leaving a gate closed indefinitely because nobody built the process that was supposed to satisfy it.

Modeling retroactive corrections as signed adjustment records, not in-place edits

A metering bug discovered after invoices have already gone out requires a correction, and the instinct to simply patch the historical aggregate so the number looks right again is understandable, since it's the fastest way to make the dashboard match reality. That instinct is also the wrong one, because a patched aggregate is indistinguishable from an aggregate that was correct from the start. There's no record that a correction happened, no mechanism to reverse it if the correction itself turns out to be wrong, and no answer available when an auditor asks why an invoice changed between one reporting period and the next.

The correct model treats a correction as its own event: a signed adjustment record appended to an immutable ledger, never an in-place rewrite of history. A correction event carries the original aggregate ID being corrected, a signed delta that's negative to reduce a charge or positive to add one, a reason string explaining why the correction was made, the identity of whoever authorized it, and the timestamp it was applied. Every one of those fields is required for the record to hold up under audit, not optional metadata attached for convenience.

The append-only structure has a second benefit beyond the audit trail: it's reversible. Corrections made under time pressure are sometimes themselves wrong, and when that happens, the reversal is just another signed adjustment event, not a second rewrite of the same row. The full sequence, original record, correction, reversal, stays intact and inspectable in order, which is exactly the property an in-place edit destroys the moment it overwrites the original value.

This isn't only an engineering preference for clean data modeling. It connects directly to how usage-based revenue has to be recognized and audited. The applicable revenue recognition standard treats usage overages as variable consideration, which has to be estimated and updated at each reporting period rather than recognized once and left alone. Auditors reviewing a SaaS company's usage-based pricing examine the historical accuracy of those variable consideration estimates and check whether the underlying billing system records actually support the numbers finance reported. A billing history built on in-place edits cannot produce that support, because the evidence of what changed and why has already been overwritten. Signed adjustment records on an append-only ledger are what make that audit possible, and the practice sits among the top audit focus areas for SaaS companies running usage-based pricing today. Flexprice, a metered billing infrastructure platform built for AI and API products that charge by tokens, credits, or compute, handles this class of billing correctness problem at the engine level rather than pushing it onto the product team.

Sources

  1. Metering pipelines for usage billing: idempotency, skew, and Stripe sync — MVP Factory
  2. What is Backfill? Meaning, Architecture, Examples, Use Cases, and How to Measure It (2026 Guide)
  3. Backfill for streaming behavioral cohorts · Issue #70574 · PostHog/posthog

More in Usage Event Metering and Aggregation