Skip to main content
The Guild platform computes estimated LLM spend in real-time based on industry-standard per-million-token list prices. All calculations are handled linearly and accurately at the individual LLM call level.

Effective-dated pricing

LLM cost calculations are date-aware. Guild prices each usage day at the rate in force on that specific day, so updating a model’s list price does not retroactively alter historical usage dashboard figures. A model can define a single flat rate or a list of price periods, each with an effective_from date (at 00:00 UTC). Spend lookups select the price in force on the day the tokens were consumed. For example, the claude-sonnet-5 promotional rate applies through August 31, 2026, and a higher rate applies from September 1, 2026 onward. Some OpenAI models, such as gpt-5.6-terra and gpt-5.6-luna, also use effective-dated rates: their launch rates apply through August 6, 2026, and reduced rates apply from August 7, 2026 onward.

Serving platform rates

The same model can reach Guild through more than one platform, such as the publisher’s own API, Bedrock, or OpenRouter, and a platform can charge a different rate than the publisher. When the pricing table has a rate for the platform that served a call, Guild prices the call at that platform’s rate. Otherwise it uses the publisher’s list price. A platform rate can also be specific to a region, such as Bedrock in eu-west-2; a call served from a region without its own rate uses the platform’s general rate. In the Models breakdown, a model that a gateway served shows a via [Platform] marker, and its rate panel names the platform whose rates it shows.

Default list prices

Below are the default list prices (in USD per Million Tokens) configured on the platform.
The claude-sonnet-5 launch rate is promotional. It applies through August 31, 2026; from September 1, 2026 the rate rises to 3.00inputand3.00 input and 15.00 output per million tokens.

Default fallback rate

If a model name is not recognized or does not match any of the custom tiers above, the system logs a warning and falls back to the default list rate (matching the Claude Sonnet 4.6 tier). The Models breakdown table displays a warning indicator next to the model name to signal that the spend is an estimate.
  • Input: $3.00 per Million
  • Output: $15.00 per Million
  • Cache Read: $0.30 per Million
  • Cache Write: $3.75 per Million

Prompt caching dynamics

To provide highly accurate cost accounting, prompt caching is calculated uniquely per provider:
  • Anthropic / OpenAI: Prompt cache write tokens are priced at a premium rate (cache_write / cache_create), and subsequent hits are charged at a heavily discounted cache_read rate.
  • Google Gemini: Google structures prompt caching differently, billing cache storage per hour rather than a per-token write rate. As a result, cache_write_tokens are priced at $0.00 (cache_create: 0.0), while cache_read_tokens represent the discounted input rate.
  • Double-charging prevention: The input_tokens count (labeled Prompt tokens in the dashboard) is canonically cache-read inclusive for all providers, so it already contains the cache_read_tokens. To prevent double-billing, the platform’s query layers deduct cache_read_tokens from input_tokens on each LLM call before applying the pricing rates, isolating the billable uncached prompt portion. Thus, the calculation always evaluates as billable_input = max(input_tokens - cache_read_tokens, 0).