Effective-dated pricing
LLM cost calculations are date-aware. Guild prices each usage day at the rate in force on that specific day, so updating a model’s list price does not retroactively alter historical usage dashboard figures. A model can define a single flat rate or a list of price periods, each with aneffective_from date (at 00:00 UTC). Spend lookups select the price in force on the day the tokens were consumed. For example, the claude-sonnet-5 promotional rate applies through August 31, 2026, and a higher rate applies from September 1, 2026 onward. Some OpenAI models, such as gpt-5.6-terra and gpt-5.6-luna, also use effective-dated rates: their launch rates apply through August 6, 2026, and reduced rates apply from August 7, 2026 onward.
Serving platform rates
The same model can reach Guild through more than one platform, such as the publisher’s own API, Bedrock, or OpenRouter, and a platform can charge a different rate than the publisher. When the pricing table has a rate for the platform that served a call, Guild prices the call at that platform’s rate. Otherwise it uses the publisher’s list price. A platform rate can also be specific to a region, such as Bedrock ineu-west-2; a call served from a region without its own rate uses the platform’s general rate.
In the Models breakdown, a model that a gateway served shows a via [Platform] marker, and its rate panel names the platform whose rates it shows.
Default list prices
Below are the default list prices (in USD per Million Tokens) configured on the platform.The
claude-sonnet-5 launch rate is promotional. It applies through August 31, 2026; from September 1, 2026 the rate rises to 15.00 output per million tokens.Default fallback rate
If a model name is not recognized or does not match any of the custom tiers above, the system logs a warning and falls back to the default list rate (matching the Claude Sonnet 4.6 tier). The Models breakdown table displays a warning indicator next to the model name to signal that the spend is an estimate.- Input: $3.00 per Million
- Output: $15.00 per Million
- Cache Read: $0.30 per Million
- Cache Write: $3.75 per Million
Prompt caching dynamics
To provide highly accurate cost accounting, prompt caching is calculated uniquely per provider:- Anthropic / OpenAI: Prompt cache write tokens are priced at a premium rate (
cache_write/cache_create), and subsequent hits are charged at a heavily discountedcache_readrate. - Google Gemini: Google structures prompt caching differently, billing cache storage per hour rather than a per-token write rate. As a result,
cache_write_tokensare priced at $0.00 (cache_create: 0.0), whilecache_read_tokensrepresent the discounted input rate. - Double-charging prevention: The
input_tokenscount (labeled Prompt tokens in the dashboard) is canonically cache-read inclusive for all providers, so it already contains thecache_read_tokens. To prevent double-billing, the platform’s query layers deductcache_read_tokensfrominput_tokenson each LLM call before applying the pricing rates, isolating the billable uncached prompt portion. Thus, the calculation always evaluates asbillable_input = max(input_tokens - cache_read_tokens, 0).