As of June 5, 2026, a Portkey AI gateway pricing worksheet should separate total gateway requests from recorded logs, guardrail checks, retention, cache hits, retries, fallback routing, budget controls, provider tokens, self-hosted operations, and Makefun review work. A single model-price row will miss the work that usually creates the overage.
The short answer: use Portkey’s pricing page and AI Gateway docs as the source of truth, then model recorded-log volume separately from provider token spend. Guardrails and observability can be valuable workflow controls, but their cost belongs beside review queues, alert triage, cache misses, retry loops, and publication handoff.
Portkey source snapshot for the calculator
Publication-time checks reached the official Portkey pricing page, AI Gateway docs, Guardrails docs, cost management docs, and logs and analytics docs. Treat those pages as the current source for plan wording, log and metrics retention, gateway routing, guardrail behavior, budget controls, and observability limits.
| Cost row | What to model | Why it matters |
|---|---|---|
| Hosted plan baseline | Current Portkey plan, included recorded logs, retention, and any published overage block. | The gateway fee is different from model tokens and reviewer labor. |
| Recorded logs | Sampling share, 100% recording scenarios, alert routing, export needs, and retention requirements. | Total gateway requests can be much larger than the logs a team chooses to retain. |
| Guardrails | Synchronous blocking checks, asynchronous logging checks, PII or policy checks, and human review queues. | Guardrails affect latency, routing behavior, and escalation work, not only the vendor bill. |
| Cache, retry, fallback | Cache-hit share, cache misses, retry rate, fallback traffic, rate limits, budget limits, and incident review. | Cache and retries can reduce or increase provider-token spend depending on prompt stability and failure rate. |
| Self-hosting and direct APIs | Operations hours, monitoring, upgrades, provider keys, budget alerts, and policy maintenance. | A self-hosted row can avoid some hosted fees but adds engineering and incident response work. |
Formula for Portkey gateway overage planning
monthly_cost = hosted_gateway_fee + recorded_log_overage_blocks + provider_token_spend_after_cache + retry_and_fallback_tokens + guardrail_review_minutes + budget_alert_review + self_host_ops_hours + Makefun_publication_handoff
The formula is intentionally conservative. It keeps provider tokens outside the gateway fee and keeps human review outside the recorded-log row. That makes overage, governance, and source-refresh costs visible before a buyer assumes every request should be fully logged forever.
Worksheet 1: recorded-log overage for Makefun source refresh
A Makefun SEO source-refresh worker sends pricing-page extraction, changelog summaries, metadata enrichment, and article QA prompts through an AI gateway. The worksheet starts with 100,000, 1,000,000, and 3,000,000 monthly gateway requests, then tests recorded-log sampling at 25%, 50%, and 100%.
- Keep included recorded logs and any current overage block separate from provider tokens.
- Model log retention and metrics retention as buyer requirements, not as a generic feature checkbox.
- Add source verification, duplicate checks, editorial QA, media generation, and WordPress handoff outside the gateway row.
- Use current Portkey pricing wording before quoting a budget to a customer.
Worksheet 2: guardrails for support and editorial triage
A support or editorial automation uses guardrails for PII, hallucination, policy, and output-shape checks before returning a draft or logging an issue for review. A useful sensitivity row uses 250,000 monthly requests split between synchronous blocking checks, asynchronous logging checks, and low-risk sampled logs.
- Synchronous guardrails belong in latency and routing rows.
- Asynchronous guardrails belong in log, alert, and reviewer-queue rows.
- Do not treat guardrails as a universal compliance guarantee without current same-scenario evidence.
- Escalation review should be priced beside provider tokens and gateway fees.
Worksheet 3: cache, retries, fallback, and self-hosting
A media metadata or editorial automation routes repeated prompts through cache, retries failed provider calls, falls back across models, and compares hosted Portkey with self-hosted gateway operations. Start with 1,000,000 monthly requests, a 40% cacheable repeated-prompt share, a 3% retry rate, a 2% fallback route rate, and a 10-hour monthly self-hosting maintenance sensitivity row.
- Cache hits lower provider-token spend only when prompts are stable enough to reuse.
- Retries and fallbacks add tokens, monitoring, idempotency checks, and failed-output review.
- Budget limits and rate limits need alert-review rows because someone must respond to the limit.
- Self-hosting should include upgrades, monitoring, policy maintenance, and incident response.
Same-use comparison rows
Use adjacent Makefun calculators as comparison lanes, not same-intent duplicates: Requesty AI gateway cache routing cost calculator, Groq Batch Flex prompt cache cost calculator, Claude Batch API prompt cache cost calculator, Gemini Batch API context cache cost calculator, and Hugging Face Inference Providers routed request cost calculator.
Requesty, Vercel AI Gateway, OpenRouter, LiteLLM, Helicone, LangSmith, Vellum, direct provider APIs, and self-hosted gateway operations should be compared only after refreshing the same request volume, logging policy, provider-token route, cache behavior, guardrail behavior, and review workflow. OpenRouter pricing is a useful marketplace/gateway reference, but it is not a Portkey log-retention clone.
Cost-control checklist
- Choose which requests need recorded logs before modeling overages.
- Separate provider-token spend from Portkey plan and observability rows.
- Split synchronous guardrails from asynchronous guardrail logging.
- Add cache misses, retry loops, fallback traffic, and budget alerts as separate rows.
- Price human review for PII, hallucination, policy, source-refresh, and editorial QA events.
- Refresh official pricing, docs, and same-unit comparison rows immediately before purchase or publication.
Publisher caveats
The Publisher gate checked the target URL, WordPress search, sitemap overlap, official Portkey source reachability, five live internal links, and a unique permanent featured image before creating this post. The article avoids cheapest, safest, fastest, most reliable, compliance, privacy, and universal-winner claims because those require current same-scenario evidence and customer-specific testing.
FAQ
Is every Portkey gateway request a recorded-log overage? No. Model total requests and recorded-log policy separately. Sampling, retention, export needs, and alert review can change the worksheet.
Do guardrails replace human review? No. Guardrails can route, block, log, or flag requests, but the cost model should still include escalation and policy review work.
When does self-hosting make sense? Consider self-hosting only after pricing operations hours, monitoring, upgrades, incident response, provider-key management, and policy maintenance beside hosted gateway fees.



