Skip to main content
Usage > Overview gives you detailed, real-time observability into LLM usage, token consumption, and financial spend. It helps you monitor costs, track performance, and drill down into usage at the workspace, agent, provider, and model level.

Accessing the dashboard

Every organization member can open the Usage dashboard, and so can the owner of a personal account. To access it:
1

Choose the account

Use the account switcher at the top of the left nav to pick your organization or personal account.
2

Open Usage

Click Usage in the left nav. It opens on Overview.
Non-admin members get a read-only view scoped to their own usage. Organization-wide figures, and the Users and Triggers breakdowns, require the Admin role.

KPI strip

At the top of the dashboard, the KPI strip displays high-level metrics for the selected time window:
  • Spend: Total estimated spend in USD, based on standard list prices.
  • Tokens: Total token volume (Prompt + Output + Cache Write). Prompt tokens are cache-read inclusive, so they already contain Cache Read tokens.
  • Sessions: Distinct user/agent sessions that recorded usage in the selected period.
  • Cache Rate: The percentage of prompt tokens served from cache (e.g., Anthropic Prompt Caching or OpenAI Cached Input).
  • Output Ratio: Output tokens as a share of total billable (uncached) input plus output tokens. Dividing by billable input prevents cache replays from diluting the ratio and hiding over-generation.
  • Budget: Spend against the workspace, agent, or user spend budget, when one is set. Visible to account admins only.
  • VS Previous Period: Percent change in spend compared to a baseline window. The baseline adapts to the active time window, and the KPI tooltip states the comparison it used. If no baseline data exists, this displays “No prior data”.
    • Rolling presets (7d, 30d, 90d, 1y) compare against the equal-length period immediately preceding the window — the last 30 days against the 30 days before that.
    • Month-aligned calendar ranges — a range starting on the 1st and ending within the same month — compare against the same days of the previous month. Aug 1–4 compares to Jul 1–4, and a full month such as Jul 1–31 compares to all of June. The day of month is clamped for shorter months, so Mar 31 compares to Feb 28 or Feb 29.
    • Any other calendar range compares against the equal-length span immediately before the selection, so Jul 10–12 compares to Jul 7–9.

Story panel

A Story panel sits at the top of the dashboard with a grounded narrative of the selected window’s usage. Guild generates it from your own metrics, so every claim points at a number you can inspect:
  • Verdict — a short read on usage and spend health for the window.
  • Cited findings — the KPIs, chart series, and efficiency findings behind that verdict.
  • Cost lever annotations — callouts on the related KPIs and series showing where spend can come down.

Ask Insights

Use the Ask Insights panel to ask questions about the usage you are viewing without leaving the dashboard. Open it from the Usage dashboard to start a chat grounded in the dashboard state on screen.
  • Seeded context — the panel opens with context from your current view, including the selected time window, active filters, and the metrics on screen, so answers reflect exactly what you are looking at.
  • Starter questions — the panel suggests dynamic starter questions drawn from your usage data, such as which agent is most expensive or how your spend changed over the window. Select a question to send it, or type your own.

Where Ask Insights sessions are saved

Each Ask Insights session is saved in your personal the-smith/<username> workspace, which Guild provisions automatically. You can reopen these sessions from that workspace at any time. To keep the panel from cluttering your general history, Ask Insights sessions are excluded from the Recent chats list in the command palette. They remain accessible from your workspace.

Usage over time

The Usage Over Time chart displays daily metrics in an interactive bar chart.
  • You can toggle between Tokens (daily token volume) and Spend (daily USD spend) using the segmented control.
  • Both views are split and stacked by model, so each bar shows what every model contributed to the daily total.
  • Hovering over any bar on the chart displays a tooltip with the formatted daily value and date.
  • The card is titled Usage over time, or Usage this month when the selected range covers the current month from the 1st.

Spend forecast

The metric control offers a third option, Spend Forecast, which extends the current month’s observed spend to a month-end projection so you can see the expected total before the month closes. It appears only on the top-level dashboard — workspace, agent, and user drill-in scopes show only Tokens and Spend.

Last 365 days (monthly)

Select the 1y (“Last 365 days”) option in the time window range selector to switch the chart into a monthly view. This view buckets the rolling 365-day window by calendar month, so each bar aggregates one month of usage.

Timezone bucketing

The daily bars are bucketed in one of two timezones, depending on how you chose the window.
  • Rolling presets (7d, 30d, 90d) bucket by your local day, detected from your browser.
  • Custom calendar ranges bucket by UTC day, because the range you pick is interpreted as UTC dates. Bucketing those in a local zone west of UTC would date the range’s first hours to the day before it starts — and, for a range starting on the 1st, to the previous month.

Model-stacked spend and interactive legend

In the Spend view, the chart stacks each bar by model so you can see how much each model contributed to spend in a given day. An interactive model legend appears below the chart:
  • Click a model in the legend to isolate that model’s series.
  • Click the same model again to toggle it back and restore the full stacked view.

Bar drill-down

Click any daily bar, in either the Spend or Tokens view, to reveal an attribution drill-down panel below the chart. The selected bar gets a highlight band and the other columns dim. The selection is held in the URL, so it survives switching between the Tokens and Spend metrics.
  • Day drill: Lists the agents that spent on that day. Each agent shows its per-agent spend total and visual model dots indicating which models it used.
The drill-down panel is always USD-denominated, including in the Tokens view: it lists per-agent spend rather than per-agent token counts, because token splits are not part of the drill payload.
Click an agent in the panel to jump directly to its row in the usage breakdown. You can narrow down the metrics shown across the entire page using the filters in the header:
  • Time Window: Choose one of the rolling presets — 7d (last 7 days), 30d (last 30 days), 90d (last 90 days), 1y (last 365 days) — or select a custom range with the date-range calendar picker.
    • The calendar picker defaults to the first of the current month through today, so usage lines up with budget cycles that run on calendar months.
    • Custom ranges can span up to 365 days, up from the previous limit of 90 days.
    • Choosing a calendar range drops the rolling preset, choosing a preset drops the calendar range, and clearing the calendar range returns the window to the 30d preset.
    • The default current-month range is left out of the ?start and ?end URL parameters, so a shared link keeps meaning “this month” instead of pinning to the dates that were current when you copied it.
  • Provider: Filter metrics to a specific LLM provider (e.g., Anthropic, OpenAI, Gemini).
  • Model: Filter metrics to a specific model (e.g., claude-sonnet-4-6, gpt-4o).
  • Platform: Filter metrics to a specific platform that served the calls (e.g., Anthropic, Bedrock, OpenRouter).
Active filters are displayed as removable chips below the header. Clicking the X on a chip clears that specific filter, or click “Clear filters” to remove all active filters.

Usage breakdown

Below the chart, the breakdown panel splits organization spend and token counts across tabs:

Workspaces

Shows spend, token count, session count, and share of spend for each active workspace. Clicking on any workspace row drills down into its specific usage details. Eligible workspaces also show an Optimize button, which starts an optimization run that tests cheaper configurations for the workspace’s agents against their own production history.

Agents

Shows spend, token count, session count, and share of spend attributed to specific versioned agents. Clicking on an agent row drills down into its specific usage details.
Usage from tool tasks or direct assistant tasks without a committed agent version is omitted from the agent breakdown but is captured in the organization/workspace totals.

Users

Available to organization admins. Shows spend, token count, session count, and share of spend for each active user in the organization. Clicking on a user row drills down into their specific usage details. Admins can configure a per-user monthly spend budget that enforces a spending limit for each accountable user. The spend shown in this tab is the same per-user accrual those budgets are measured against.

Triggers

Available to organization admins. Shows spend, token count, session count, and share of spend for each trigger whose sessions recorded usage in the window, so you can see what automated runs cost. Unlike the Workspaces, Agents, and Users rows, a trigger row doesn’t drill down.

Platforms

Shows spend and token counts grouped by the platform that served each model call, such as Anthropic, Bedrock, or OpenRouter. A platform that serves from a specific region shows it after the name, such as Bedrock · us-east-1. A platform is who served the call (for example, Bedrock). A provider is the model’s publisher (for example, Anthropic). Because the same model can reach Guild through more than one platform, this tab shows where your calls were actually served.

Providers

Shows spend and total tokens grouped by the model’s publisher (for example, Anthropic), regardless of which platform served the call.

Models

Shows spend and token counts grouped by LLM model name.
  • Rates & spend breakdown: Click any model row to expand it in-place and reveal the per-million-token list prices (Input, Output, Cache Read, Cache Write) and the exact calculated spend contribution for each token type. For historical views, the rate shown is the one in force on the day the tokens were consumed, so past figures stay stable when a model’s list price changes.
  • Served-via marker: When a gateway rather than the publisher’s own API served a model’s calls, the row shows a via [Platform] marker (for example, via Bedrock). A model served through several platforms lists them, largest spend first. The expanded rate panel then labels its prices Rates for [Platform], naming the platform with the most spend for that model. See Serving platform rates.
  • Estimated pricing indicator: Models that Guild priced using the default list price display a warning indicator next to the model name. Hover the indicator to see a tooltip, or expand the row to read the estimate note in the rate panel.
Active monthly budgets recompute using current list prices if a model’s price changes mid-month, so a cut can reopen a budget and an increase can trip it early. Completed historical months are frozen and are not affected by later price changes.

Drilling down (workspace, agent, and user scopes)

Clicking on a row in the Workspaces, Agents, or Users breakdown navigates to a dedicated drill-in view:
  • The breadcrumb navigation changes to Usage / Workspace · <Name>, Usage / Agent · <Name>, or Usage / User · <Name>.
  • A prominent scope pill badge is displayed below the header indicating the current drill-in filter.
  • The KPI cards, chart, and remaining breakdown tabs are automatically scoped to represent only the selected workspace, agent, or user.

User drill-in

Clicking a row in the Users breakdown scopes the metrics, charts, and breakdowns to reflect the selected user’s spend. User attribution is based on the chat-session initiator, representing interactive spend. Automated trigger and test spend is not user-attributed, so the Users and Triggers tabs are hidden under a user scope. A View user profile header action link opens the user’s Agent Hub profile.

Exporting data

The Export button in the page header downloads the window’s usage as CSV. Selecting it opens a dialog to pick the datasets and confirm the download. It is available here and on the Platform spend page.

Spend budgets

Admins can set and monitor spend budgets from the dashboard. The Budget KPI and the budget column in the usage breakdown are admin-only, so a member who can view Insights but not administer the account does not see budget ceilings, spend against them, or the editor. The budget editor opens as a popover rather than an empty input:
  • Spend so far — how much the scope has spent in the current cycle, so you can size a ceiling against real spend.
  • Suggested budget — a single suggested ceiling derived from the scope’s daily spending pace, taking whichever is higher of the current cycle’s pace and the last full cycle’s, projecting it across the month, and adding a 25% buffer. The buffer is exactly 1 / 0.8, so a steady month finishes at the 80% warning line rather than past it. A scope with no spend in either cycle is suggested $50 instead. Click the suggestion to fill the field.
  • Validation — a budget must be greater than what the scope has already spent this cycle; the editor rejects a lower amount rather than saving a budget that is already exhausted.

Efficiency findings

The findings engine proactively scans LLM usage and surfaces actionable recommendations to reduce wasted spend. Each waste-detection check flags a specific pattern so you can act on it:
  • Prompt cache written but rarely reused (cache_write_without_read): Triggers when cache-write volume is high but the read-to-write reuse ratio is extremely low, meaning you pay to write prompt cache that is barely read back.
  • Output-heavy generation (over_generation): Triggers when output tokens dominate the billable input plus output mix, indicating verbose or costly output.
  • High LLM retry rates (retry_burn): Triggers when LLM calls retry frequently, indicating potential performance or reliability issues that burn spend and latency.

Platform tab

The Platform tab charts bring-your-own-key (BYOK) provider spend month over month, with per-developer attribution synced from your provider admin APIs. Like the rest of Insights, organization-wide figures require the Admin role. See Platform spend for the chart, month drill-down, and six-month outlook.