BubbaBuiltSelf-hosted inference · OpenAI-compatible API
BubbaBuilt is the billing, metering and access layer in front of your GPU servers. Customers get a familiar API. You keep the hardware.
Point the gateway at any OpenAI-compatible server — vLLM, Ollama, llama.cpp — and keep inference on your own GPUs.
Cryptographically random keys shown once, stored as hashes, with prefixes, rotation and per-key usage.
Per-plan request limits and transactional credit accounting that can never go negative.
Every request recorded with tokens, latency, cost and charge — without storing prompts by default.
Row-level security on every customer-owned table, plus server-side authorization on every endpoint.
Publish stable model names and swap the hardware behind them without breaking customer integrations.