As of June 4, 2026, Requesty AI Gateway cost planning should start with a worksheet, not a headline savings claim. Separate direct provider token spend from cache-eligible calls, then add Requesty credits, auto top-up, project and API-key spend caps, BYOK responsibility, failover routes, observability value, and the Makefun workflow owner who pays for retries and review.
Publication-time checks used the official Requesty pricing comparison page, Requesty AI Gateway page, spend limits and rate limits docs, BYOK docs, and quickstart docs. Treat those pages as the source of truth because model rows, free-credit language, cache behavior, and gateway terms can change.
Requesty pricing snapshot for cache, credits, routing, BYOK, and spend limits
| Worksheet row | What to record | Planning caveat |
|---|---|---|
| Provider baseline | Current input and output token rows for the models you actually route. | Do not mix fresh research prompts with repeated extraction prompts when estimating cache value. |
| Requesty credits | Free-credit language, prepaid credits, auto top-up threshold, and wallet balance rules from the current pricing page. | Credits and top-up settings control interruption risk, not only monthly average cost. |
| Cache behavior | Cache-eligible request share, cache-hit rate, cached response value, and non-cacheable jobs. | Cache savings are scenario-dependent and should not be sold as guaranteed savings. |
| Spend limits | Project limits, API-key limits, alert path, and what happens when a cap stops a workflow. | A cap without retry, queue, and owner routing can create publication failures. |
| BYOK and failover | Provider key ownership, upstream billing/rate limits, fallback models, and routing policy. | BYOK can shift billing, support, and quota responsibility outside Requesty. |
Cost formula: provider baseline, cache, credits, and Makefun side costs
A useful Requesty calculator starts with this formula: monthly cost equals uncached input tokens times current model input rate, plus uncached output tokens times current model output rate, plus gateway credit or margin effects, minus verified cached-response value, plus retries, source fetch, evidence storage, reviewer minutes, alert handling, and failed-cap recovery. Keep those non-token Makefun rows visible because they often decide whether a gateway workflow is operationally cheaper.
Worksheet 1: SEO source-refresh jobs with repeated extraction prompts
For SEO source refresh, duplicate checks, and structured extraction, model monthly requests, average input/output tokens, cache-eligible share, actual cache hit rate, retry rate, failed source checks, and the editor minutes needed to approve rows. Requesty can help only when the worksheet isolates repeated extraction from fresh research and records which project or API key owns the spend.
Worksheet 2: media metadata and transcript routing with fallback policies
For metadata, transcript classification, moderation-support, and source cleanup jobs, start with a cheap default model, then price the fallback share separately. Failover can protect a workflow from upstream 429s or outages, but it can also change model quality, latency, output review cost, and token price. The worksheet should show default model tokens, fallback model tokens, provider 429 retries, transcript storage, media QA, and alert review as separate rows.
Worksheet 3: support and coding-agent jobs with project and API-key caps
Support classification, issue triage, and coding-agent maintenance need governance rows before scale: project monthly limit, API-key limit, owner tag, environment tag, alert thresholds, retry policy, and interruption cost. A spend cap is useful only when it fails loudly and routes the job to a human owner instead of silently dropping publication or support work.
Same-use comparison table for gateway and observability alternatives
Compare Requesty with alternatives by workflow unit. Vercel AI Gateway and Cloudflare AI Gateway matter when the platform stack already owns deployment, logging, and routing. LangSmith, Braintrust, and Langfuse matter when tracing, evals, prompt experiments, and observability are the commercial center. LiteLLM self-host can fit teams that want infrastructure control. Direct provider APIs can win when one provider, existing discounts, or committed-use terms are enough. Avoid cheapest, best, fastest, safest, reliability, compliance, or universal savings claims without current same-scenario evidence.
Relevant Makefun context includes the Hugging Face Inference Providers routed request cost calculator, Claude Batch API prompt cache cost calculator, Gemini Batch API context cache cost calculator, Brave Search API LLM context answers cost calculator, Pinecone Assistant context token RAG cost calculator, and Botpress AI spend message event knowledge cost calculator.
Cost-control checklist for Requesty gateway workflows
- Separate cacheable repeated calls from fresh research, creative, or high-risk review prompts.
- Record current model rows, free-credit terms, prepaid credit balance, auto top-up settings, and source date.
- Set project and API-key caps before unattended Makefun jobs run on a schedule.
- Define fallback model limits and review quality drift before treating failover as a savings feature.
- Tag requests by workflow, owner, environment, model, and job type so alerts route to the right person.
Risks and Publisher caveats
- Refresh Requesty pricing, model rows, cache behavior, spend-limit docs, BYOK docs, quickstart docs, and competitor rows before quoting exact budgets.
- Do not claim Requesty is the cheapest, best, fastest, safest, most reliable, or compliant option without current same-unit proof.
- Keep provider billing, rate limits, support, privacy, and data-handling responsibility separate when BYOK is used.
- Include retries, source fetch, storage, QA, alert handling, and workflow interruption as real Makefun costs.
- Use topic-specific permanent media only; avoid provider logos, copied UI, temporary URLs, third-party hotlinks, and reused media.
FAQ
Is Requesty always cheaper than direct provider APIs?
No. The answer depends on model mix, cache eligibility, actual cache-hit rate, BYOK terms, provider discounts, retries, and the cost of operating alerts and reviews.
What should a Requesty calculator measure first?
Start with direct provider baseline tokens, then isolate cacheable calls, prepaid credits, auto top-up, project and key caps, fallback share, and human review cost.
Why include Makefun workflow owners in the worksheet?
Gateway limits and alerts only help if they route failures to the team that owns the workflow. Otherwise a spend cap can protect billing while still breaking source refresh, support triage, or publishing work.



