Usage Billing Review

open-source billing engines for AI-native SaaS: a comparison of self-hostable options

Which self-hosted billing engine handles real-time usage at API speed.

Features Editor · · 11 min read · Updated
Cover illustration for “open-source billing engines for AI-native SaaS: a comparison of self-hostable options”
Rating Engines and Pricing Rules · August 31, 2026 · 11 min read · 2,578 words

The shift from seat-based to consumption-based pricing didn't creep up on anyone paying attention. Usage-based adoption among SaaS companies climbed from roughly 30% in 2019 to around 85% by 2024, and that speed is the whole story here. Most billing infrastructure on the market, open-source or otherwise, got built for a world where pricing moved on quarterly cycles. It shows up as friction the moment usage starts moving on a per-second basis instead.

AI workloads expose the gap fast. A flat monthly seat can absorb a lot of variance in how someone uses a product, but it can't absorb the swing between a routine API call and a token-heavy agent run that costs orders of magnitude more to serve. Agentic AI pushes this further: outcome-based billing is emerging as a new pricing category, and subscription-era billing engines were never built for it. Credit-based pricing is the industry's answer so far. In the PricingSaaS 500 Index, 79 companies now run credit models, up from 35 at the end of 2024. Credits give customers a budget they can actually predict, without forcing vendors back into flat-rate seats. Hybrid pricing, subscription plus usage in the same plan, now shows up at 43% of companies and is projected to hit 61% by the end of 2026. Any billing engine worth putting in production in 2026 has to run subscription, usage, and credit logic at once, for the same customer, often on the same invoice line.

This piece looks at the open-source, self-hostable billing engines that AI-native teams actually put on their shortlist. It asks a narrower question than most comparisons bother with: which one holds up when usage moves at API speed instead of monthly-cycle speed.

Diagram: Usage-Based Pricing Adoption: 2019 → 2024. Visualizes: Show the acceleration of usage-based pricing adoption among SaaS companies from roughly 30% in 2019 to around 85% by 2024.

The four criteria that actually separate these engines for AI workloads

Most open-source billing comparisons run down a feature checklist: invoicing, dunning, tax handling, payment gateway hookups. Those matter, but four other things do the real separating when an AI product scales.

Real-time metering granularity. A multi-step agent workflow can throw off 3,000 events a minute. If the metering layer can't keep pace, enforcement ends up checking a number that's already stale. The failure mode is specific and it's ugly: a customer blows past their credit limit, the metering layer catches up an hour later, and by then the overage already got served. Somebody eats that cost, either the vendor or the customer in a billing dispute nobody wants to have. Real-time metering means tracking consumption as it happens, a separate function from real-time invoicing, which still settles on a cycle. What changes is whether the system knows, right now, what somebody has actually used.

Credit-balance enforcement speed. People conflate this with metering constantly, but it's a distinct problem. Enforcement is the gate that stops a request the second a balance hits zero. At the millisecond speeds a modern API runs at, enforcement latency stops being a billing detail and turns into a product-integrity question. Slow enforcement either blocks somebody legitimate or lets an over-limit request slip through, and neither outcome is good.

Hybrid model support. A seat fee plus per-token overage plus a prepaid credit wallet is just what one customer's plan looks like at a lot of AI companies right now. The billing engine has to run all three lanes simultaneously as the baseline case, not as a feature on next year's roadmap. Roughly 44% of SaaS companies now charge for AI-powered features. That's increasingly the normal shape of the bill.

Operational burden of self-hosting. Open source trades a license fee for engineering time, infrastructure know-how, and an on-call rotation that never really ends. What actually matters: which dependencies the engine drags in (Kafka and ClickHouse show up constantly), how well the team already knows the language and runtime, how good the docs really are, and whether the engine needs a separate pre-aggregation layer bolted on before it can even touch raw events. Data residency and compliance, SOC 2, on-prem mandates, are often the real reason a team started looking at self-hosted options in the first place. Whether an engine genuinely delivers on those guarantees is frequently what decides it.

Diagram: Hybrid Pricing Is Becoming the Majority Model. Visualizes: Show the rapid rise of hybrid pricing (subscription + usage in one plan) as a share of SaaS companies: 43% today, projected to reach 61% by end of 2026.

How the standard self-hosted metering stack is architected

Diagram: The Standard Self-Hosted Metering Pipeline. Visualizes: Illustrate the four-stage architecture that most open-source billing pipelines follow for AI workloads: (1) Ingestion — raw telemetry (API calls, tokens burned, storage written, agent…

Most open-source billing pipelines in this space follow the same rough shape: an event queue, usually Kafka, feeds a high-volume OLAP store, usually ClickHouse, which feeds a pricing and rating layer, which feeds invoice generation. Each stage exists to add structure before raw data hits anything that touches money.

Ingestion grabs raw telemetry, API calls, tokens burned, storage written, agent actions completed, and normalizes and dedupes it. Metering turns that stream into aggregated, billable metrics. Rating and enforcement apply the pricing rules, check credit balances, and cut off consumption when a limit hits. Invoice generation closes out the period and produces the record a customer actually sees.

One principle holds the whole thing together: usage collection has to stay decoupled from the main application. A queue or a dedicated ingestion service absorbs burst traffic without blocking the product itself, which matters a lot when an agent workflow can spike event volume with zero warning. Idempotency isn't optional here either. Duplicate events at AI-scale volume happen routinely, so the engine has to dedupe without dropping anything real in the process.

Where these engines genuinely split is in what they hand the operator to build. Some demand a pre-aggregation layer before raw events even reach the billing engine, meaning the team designs and runs that layer itself. Others take raw event streams directly and handle aggregation internally, a real operational difference and one of the cleaner ways to tell these systems apart. Coverage of the pipeline varies too. Some engines meter usage but leave invoicing, dunning, credit wallets, and a customer portal to be sourced elsewhere, while others aim to cover the whole path end to end.

One project here has built the largest community of any self-hostable billing engine in the space, a GitHub footprint north of ten thousand stars and hundreds of forks as of mid-2026. The architecture philosophy is modular and developer-assembled: it hands the operator well-designed API primitives for metering and pricing, and the operator wires the pieces together rather than getting a pre-packaged suite. It processes up to a million billing events per second, which covers most AI-native workloads even at real scale, and it counts OpenAI and Cribl among its named customers, a real signal for high-volume API billing specifically. It's licensed AGPL-3.0, which teams need to check against their own product's distribution model before committing to anything. The project raised $15 million in March 2024, and commercial pricing has since moved to fully quote-based for Business and Enterprise tiers, so total cost of ownership is harder to pin down up front than it used to be. Here's the gap for AI-native teams: the modular, build-it-yourself design is a real strength if you've got billing engineers to spare, but credit-wallet enforcement and hybrid-model logic are additional build work, not ready-made primitives waiting to be flipped on. Best fit is a team with dedicated platform engineering capacity that wants fine control over billing logic and is comfortable with what AGPL means for its own codebase.

OpenMeter (now Kong Metering & Billing) — a purpose-built meter, not a full billing engine

OpenMeter became part of Kong in 2025. The cloud product now lives inside Kong Konnect as Kong Metering & Billing, while the open-source core stays on GitHub as its own project.

Its scope is narrow and specific. It ingests high-volume usage events and turns them into billable data, with SDKs for Node.js, Python, and Go, and ClickHouse running underneath for fast aggregation. It's a real capability, well-built, and focused squarely on the metering stage. Invoicing, dunning, subscription management, credit wallets, and customer portal flows all need to come from separate tooling or something built on top.

Self-hosting it means the team needs real Kafka and ClickHouse chops, and even with that in hand, full billing still needs a payment and invoicing layer bolted on. Licensing is Apache 2.0, permissive, no AGPL headaches to untangle. Who it actually serves well: AI, API, and DevOps teams that already have billing infrastructure running and want a best-in-class ingestion and metering layer added on top, not teams starting from nothing. For an AI-native company that needs credit enforcement, hybrid plan logic, and invoice generation out of one self-hosted system, this alone doesn't get you there.

Meteroid — Rust-native architecture designed for raw event throughput

Meteroid's core bet is different from most of the field. It's built in Rust and designed to take raw event streams directly, no pre-aggregation layer required before events reach the engine. That's a real design choice, because it removes an entire piece of infrastructure that teams on other engines have to build and babysit themselves.

It's aimed at IaaS, PaaS, and AI-driven SaaS workloads that need to chew through millions of usage events a second without standing up a separate aggregation pipeline first. Where it's still catching up is track record. It doesn't have the extreme-scale production history that older projects have racked up, and it's missing the deep enterprise tax and accounting modules a legacy engine would offer.

Best fit is a mid-sized SaaS startup or an API-first product that wants modern billing infrastructure fast and doesn't have complicated enterprise accounting to satisfy yet. For AI teams specifically, skipping the pre-aggregation layer is a genuine win: one less system to run and watch at 3am. Still, check the engine's credit-wallet and hybrid-plan support against the actual roadmap before committing. Don't assume those exist just because the throughput story is strong.

Kill Bill and jBilling — when enterprise legacy requirements drive the choice

Kill Bill is the oldest project in open-source billing, full stop. It's Java-based, subscription-first in its architecture, and comes with a genuinely deep catalog and invoicing feature set, extended by a mature plugin ecosystem built up over years. It's self-hostable, SOC 2 compliant, has an active GitHub repo, and it's proven itself at large enterprises with heavy legacy accounting and ERP integration needs.

The gap, for an AI-native workload, is architectural, not some minor missing feature. Kill Bill's subscription-first design means real-time credit enforcement and token-level metering just aren't native. Getting either working takes plugin development or a workaround, and that's real cost to plan for, not a five-minute config tweak.

jBilling sits in similar territory: modular, enterprise-focused, built for complex billing workflows and ERP or CRM integration across high-volume industries. The honest read on both projects is the same. They're legitimate, proven choices when the hard requirement is ERP integration or an accounting workflow already in place that can't get rebuilt from scratch. Their fit weakens sharply when the hard requirement is millisecond-level usage enforcement for AI agents, and no amount of plugin work fully closes that gap.

UniBee — lightweight self-hosting for early-stage billing needs

UniBee positions itself as easy-to-deploy subscription and usage billing for startups whose billing needs are still simple. Managed cloud pricing starts at $99 a month for Cloud Starter and $399 a month for Cloud Business, with Enterprise on custom terms. That makes it one of the few options here with actual public pricing instead of a quote-only wall.

The ceiling here is the same thing as the selling point: it's lightweight by design. Teams that end up needing credit wallets, real hybrid-model complexity, or high-volume real-time enforcement will outgrow it, probably sooner than they'd like. That's simply the scope it's built for. It fits pre-scale startups that need working billing running now and plan to revisit the whole decision once product complexity actually demands something bigger.

The operational cost of self-hosting that comparisons usually skip

Open source trades a license bill for engineering time, infrastructure know-how, and an on-call rotation that doesn't end. Engines built on Kafka and ClickHouse need teams who've actually run both in production, and neither system is easy to keep healthy at scale. That's a real cost, not a hypothetical footnote.

The maintenance work teams tend to underestimate: schema migrations every time pricing models change, idempotency edge cases that only show up under burst traffic nobody load-tested for, reconciliation runs when upstream events arrive late or out of order, and compliance updates every time SOC 2 or data residency rules shift underneath the product. None of it shows up on a feature comparison chart. All of it shows up on an engineering calendar eventually, usually at the worst time. The incidents that surface at the worst moments are rarely about missing features; they're about edge cases in event handling that only appear under real production conditions.

The data residency argument for self-hosting is real. Teams in regulated industries, or with enterprise customers who require on-prem deployment, have a genuine, defensible reason to self-host. It's a contractual reason, not just a philosophical preference. But that reason needs an honest cost attached to it: the engineering hours to build and launch, the infrastructure spend, the on-call rotation, the opportunity cost of billing engineers who aren't spending that time on the actual product. Self-hosting removes a licensing line item, but it hands back infrastructure, compliance, and maintenance work that plenty of teams don't see coming, especially when on-prem or SOC 2 is a hard requirement rather than a nice-to-have. Managed platforms like Flexprice offer real-time metering and credit-enforcement capability comparable to the self-hosted engines described here, while carrying the operational weight, the infrastructure, the documentation gaps, the on-call burden, as part of the service itself. That trades implementation time for billing agility, with less of the permanent engineering tax that comes with running it yourself.

How to match engine to workload given the tradeoffs above

There's no single right answer here, and anyone who tells you otherwise is selling something. The match comes down to engineering capacity, workload complexity, and data residency requirements, and different teams weigh those three differently for reasons that make sense given where they sit.

A team that needs a metering layer only, with invoicing infrastructure already in place, should treat a purpose-built meter as exactly that: a focused tool, paired explicitly with a separate invoicing layer, sized to that one job. A mid-sized team where raw event throughput is the real bottleneck, with no appetite to build a pre-aggregation layer, should look hard at a Rust-native design built for that exact problem, and still check credit-wallet support against the actual roadmap before signing off. An enterprise where ERP or accounting integration is the hard requirement should lean toward the legacy, subscription-first engines and budget for plugin work to cover AI-native metering on top. A pre-scale startup that needs billing running in days, not months, should grab the lightweight option and pair it with a clear plan for the day it gets outgrown.

Underneath all of it sits one real question: is the goal to own the billing infrastructure, or to own the pricing logic? Owning the infrastructure, through a self-hosted engine, buys data control and no vendor lock-in. Owning the pricing logic, through a configurable platform, buys speed to iterate. Teams that mix the two up tend to over-build the infrastructure and under-invest in the logic that actually drives revenue. Companies running hybrid pricing models report the highest median growth rate of any pricing strategy, 21%, and whatever billing engine sits underneath that growth has to flex between subscription, usage, and credit lanes at once, for the same customer, without forcing the team to pick one lane and bend the rest of the business to fit it.

Sources

  1. unibee.dev

More in Rating Engines and Pricing Rules