How It Works
The Vantage callback runs inside your LiteLLM process and emits one usage record per request to a local Vantage collector over a Unix socket. The collector compacts those records into batches, uploads them to a Vantage-managed bucket using short-lived presigned URLs, and Vantage joins that usage to your provider costs during each provider’s cost ingestion. See How LLM Enrichment works for the shared indexing, join, split, and tag pipeline.How Cost Rows Are Split
Vantage splits each matching provider cost row into enriched rows, allocated proportionally by token share. Splits are additive, so the enriched rows always reconcile to the original total: usage your logs do not cover stays on a leftover row, and cost rows with no matching usage pass through unsplit. Enabling enrichment never changes your totals.How Cost Splitting Works
Data Freshness and Backfill
Enrichment runs as part of each provider’s existing cost ingestion, so it follows that provider’s refresh cadence. Recent days are reprocessed within a rolling three-day window so late-arriving batches are picked up. See the provider data refresh documentation for per-provider timing. A billing period is enriched whenever it is processed while an active collector exists, for as long as the matching batches remain available to Vantage. Enrichment begins from when your collector starts uploading usage; it does not backfill spend from before the collector was running. Late-arriving batches for recent days are picked up automatically within the rolling three-day window. Re-enrichment reads the already-normalized cost data, so it does not require a full cost re-import.Prerequisites
Before you begin, make sure:- An active cost integration exists for at least one supported provider: OpenAI, Anthropic, AWS, Google Cloud, Azure, SpaceXAI, or Baseten.
- Your Azure deployments in LiteLLM, if any, use the
azure/(Azure OpenAI) orazure_ai/(Azure AI Foundry) model prefix. Vantage routes usage by the provider LiteLLM reports, so a deployment configured with theopenai/prefix is reported as OpenAI even when its base URL points at Azure, and its usage will not match your Azure costs. - You run a LiteLLM proxy where you can install the Vantage callback and run the collector as a sidecar or adjacent container. The collector needs a persistent filesystem for its spool and outbound HTTPS access to Vantage Core and the presigned Amazon S3 URLs it returns.
- You have a Vantage API token for the collector, in addition to the installation token you create in Step 1.
- You have a Vantage Organization Owner or Integration Owner role. See Role-Based Access Control.
Set Up LiteLLM Enrichment
Setup has three steps: create the source and copy the installation token, install the callback and deploy the collector, and configure which cost integrations to enrich.Step 1: Create the Source and Copy the Installation Token
In Vantage, go to the Integrations page. Under LLM Enrichment, add LiteLLM, then select Connect Source. Vantage provisions a collector integration and shows a Collector Setup screen. On that screen, copy the installation token from Step 1: Copy the Installation Token. It looks something likesrvc_data_intgrtn_9f8e7d6c5b4a3f21. This token authorizes your collector to upload usage to Vantage. You paste it into your collector configuration in the next step.
Step 2: Install the Callback and Deploy the Collector
LiteLLM Enrichment has two runtime pieces that work together:- Vantage callback: A Python package loaded inside your LiteLLM proxy process that records each request’s usage.
- Vantage collector: A sidecar container that receives events from the callback over a Unix socket, batches them, and uploads them to Vantage.
Install and enable the Vantage callback in LiteLLM
vantage-litellm-callback package into the same Python environment as your LiteLLM proxy (the callback requires Python 3.10 or later), then register it so LiteLLM loads it.Install the callback package, pinning the version that matches your collector image (the current release is 0.0.7):pip on PATH. If pip is not found, bootstrap it first, then install into the same interpreter:vantage_callback importable, so the callback path above resolves on its own. Some LiteLLM images only load callbacks from Python files next to config.yaml. If LiteLLM does not pick up the callback after install, place a shim at /app/vantage_callback.py (the same directory as config.yaml) that imports callback_instance.Deploy the Vantage collector sidecar
vX.Y.Z publishes the image tags X.Y.Z, X.Y, X, latest, and an immutable commit-SHA tag. Pin a full version tag (X.Y.Z) rather than latest; the current release is 0.0.7.Give the collector a persistent local spool, a Unix socket shared with LiteLLM, and outbound network access to Vantage. Configure it through environment variables (or your platform’s equivalent, such as a Helm value or a command-line flag):VANTAGE_COLLECTOR_SOCKET_PATH to the same socket path on the LiteLLM container so the callback and collector connect.VANTAGE_COLLECTOR_ID and its own persistent spool so Vantage can keep each batch stream distinct.Restart LiteLLM
Container Deployment Constraints
The release collector image is builtFROM scratch, has no shell, runs as UID/GID 65532, and declares volumes for its spool and socket. These properties shape how you run it on any container platform; the notes below call out what that means, with ECS Fargate as a worked example.
- Health and metrics endpoints: The collector serves
GET /healthzfor readiness and PrometheusGET /metricson its metrics address (:9090by default). Use these for orchestrator health checks and monitoring. These endpoints are unauthenticated, so do not expose the metrics port outside your private network. - Memory: Request at least 256 MiB for the collector container and set a 512 MiB limit unless your own load testing supports a different value. The collector sizes its internal buffers from the container’s memory limit.
- Per-task isolation: Each task or pod needs its own socket and spool volumes; do not share a spool across tasks. Use a stable
VANTAGE_COLLECTOR_IDper instance, such as the task hostname. - Shared socket: Mount the same socket volume into both the LiteLLM and collector containers, and set
VANTAGE_COLLECTOR_SOCKET_PATH(LiteLLM) andVANTAGE_SOCKET(collector) to that socket. - Secrets: Supply
VANTAGE_API_TOKENandVANTAGE_INTEGRATION_TOKENthrough a secret store (for example, AWS Secrets Manager), not as plaintext values in a task definition or Compose file. - Volume ownership (ECS Fargate): Because the collector runs as UID
65532, and Fargate’s ephemeral volumes are root-owned, the collector may be unable to write its socket and spool by default. Override the collector task’suserto"0", or wrap the image with an entrypoint thatchowns the mounts before dropping privileges.
systemd service instead of a sidecar: install the binary, create a non-login vantage user, and provide the same environment variables through a mode-0600 environment file. The socket, spool, tokens, and provider selection all work the same way.Step 3: Configure Cost Integrations
Back on the Collector Setup screen in Vantage, under Configure Cost Integrations, select Configure Providers and choose which connected cost integrations should receive enrichment from this collector. Vantage enriches costs for the selected integrations on each provider’s next data refresh. Provider cost integrations you connect later are not enriched automatically; return to the source’s Manage Providers screen to enable them.Manage LiteLLM Enrichment Sources
Manage your collector from the LiteLLM integration page. The Connected Collectors table lists each collector with its Installation Token, Cost Providers, Status (for example, Pending before an import starts, Importing while one runs, Stable once all imports succeed, Warning if only some imports fail, Error if an import fails, or Paused if the collector is stopped), and creation date. Each row has an Edit button and an ellipses (…) menu with View import history and, for collectors that are not paused, Stop. The Collector Setup (Edit) screen also shows a Data Received indicator (Waiting, Received, or Unavailable) so you can confirm the collector is delivering usage independently of whether enrichment has run yet.Choose Which Integrations Are Enriched
Select Edit on the collector, then Manage Providers, to change which connected cost integrations receive enrichment. Remember that only one cost integration per provider can be enriched. Disabling an integration here stops going-forward enrichment for it but leaves its existing enriched history in place.View Import History
In the sources table, click the ellipses (…) next to a row and select View import history. This opens the collector’s Import History, where the Enrichment Runs table shows every run for this collector. The Enrichment Runs table lists one row per provider cost integration and billing period that ran enrichment, newest first, with columns for the Integration (and its account), Status, Tokens Kept, Log Lines Kept, Bill Match, Billing Period, and Last Enriched At. Lifecycle changes appear as their own rows: Added when enrichment is first enabled for an integration, and Paused or Resumed when you stop, disable, resume, or re-enable it. The three percentages look similar but measure different things at different stages, so a low number in one column means something very different from a low number in another:LLM Enrichment Import History
Stop a Collector
To stop a collector, click the ellipses (…) next to the row and select Stop, then confirm in the Stop enrichment source dialog. This is a soft deactivation, not a hard delete: Vantage stops enriching new cost data for that collector, but your existing enriched history is left unchanged, and the collector stays visible in the list so you can bring it back at any time.Resume a Collector
To resume a stopped collector, select Resume in its row. Resuming restarts going-forward enrichment for the integrations that were active when you stopped it (integrations you had already disabled individually stay disabled) and reopens the Edit screen so you can adjust the selection.View Enriched Costs on Cost Reports
Once enrichment runs, a single provider cost line is split into multiple rows, each carrying enrichment tags. You can filter and group by these tags anywhere tags are supported: Cost Reports, Virtual Tags, Budgets, and Cost Alerts. Because enrichment splits (allocates) your provider costs, enrichment tags behave like Vantage’s cost allocation tags: you can build a Virtual Tag on them, but a cost can be allocated only once, so an enrichment tag can belong to only one allocation chain.Enrichment Tag Reference
The callback builds each usage record’s tags from your LiteLLM request metadata. It merges themetadata.tags and metadata.spend_logs_metadata objects (when both set the same key, metadata.tags wins), and metadata.tags may also be a list of key:value strings. It also promotes these LiteLLM request attributes to tags automatically, under the same key names: organization_id, organization_alias, team_id, team_alias, project_id, project_alias, user_id, end_user_id, and key_alias. Tag keys must be strings, and values must be strings, booleans, integers, or floats (null, nested, and array values are dropped). Each record carries at most 32 tags, with keys up to 128 characters and string values up to 256 characters.
vntg:ai: prefix (for example, vntg:ai:model); this keeps them distinct from your keys and consistent with the AI tags Vantage applies to provider costs. Because a key like team_alias is not provider-namespaced, it lines up across OpenAI, Anthropic, and the other providers, so you can group your entire AI stack by one team_alias tag. Consistent key naming matters: team and Team are two different keys.
vntg:ai:* tags (such as vntg:ai:model); split rows also carry the tags from their usage slice. When a slice tag and an existing tag use the same key, the enrichment value wins. The leftover row (usage not covered by the collector) keeps the provider’s existing tags and the applicable vntg:ai:* tags but carries none of the slice tags. Request identifiers are never turned into tags. Vantage may also apply a tag denylist to drop noisy or sensitive keys before they become tags.- To group: open the Group By menu, select Tag, and choose the tag key, for example
vntg:ai:model, or one of your own keys such asteam. - To filter: open the Filters menu, click New Rule, select Tag, choose the Tag Key, then pick an operator and one or more values.
Troubleshooting
The Data Received status stays on Waiting
The Data Received status stays on Waiting
- Your application is sending traffic through the LiteLLM proxy, and LiteLLM was restarted after you added the callback.
- The callback is registered in
litellm_settings.callbacksasvantage_callback.callback_instanceand is importable in the LiteLLM process (installed withpip, or mounted so/appis on the import path). - The collector container is running, has outbound network access to Vantage, and shares the Unix socket with LiteLLM (
VANTAGE_COLLECTOR_SOCKET_PATHmatches on both). If the collector exits right after it starts, confirmVANTAGE_CORE_URLbegins withhttps://. VANTAGE_INTEGRATION_TOKENmatches the installation token from Step 1, andVANTAGE_API_TOKENis a valid Vantage API token with integration permissions.
My collector stays on Pending
My collector stays on Pending
- The Data Received indicator on the collector’s Collector Setup screen. Waiting means the collector has not uploaded usage yet: finish Step 2 and see the Waiting entry above.
- If Data Received shows Received, confirm that at least one cost integration is selected under Configure Providers (Step 3), with exactly one integration per provider.
A provider is not being enriched
A provider is not being enriched
- Confirm the provider has exactly one cost integration selected on the Manage Providers screen. A provider with more than one integration selected is not enriched because LiteLLM usage does not carry an account identifier.
- Confirm the provider is one of the supported providers and that its cost integration is active.
- For Azure, confirm your LiteLLM deployments use the
azure/orazure_ai/model prefix. Usage from anopenai/deployment that points at Azure is reported as OpenAI. See Prerequisites. - Allow a full data refresh for that provider to complete after connecting. Enrichment applies on the next refresh.
My costs are not being split
My costs are not being split
- There is no active collector for the account, or usage has not been ingested for that billing period.
- The cost row’s model or token type could not be matched to a logged request for that provider, date, and token kind.
- Records were skipped during ingestion because a required field was missing, the provider was blank or unsupported, the model was blank or could not be matched to your provider cost data, the request was not successful, or no usage value was a positive integer.
I can't tell which costs were enriched
I can't tell which costs were enriched
- Check coverage in Import History: Open the source’s Import History and review Tokens Kept (the share of your logged tokens Vantage attached to a cost row) and Bill Match (the share of eligible provider cost rows that received logged usage) for each integration and billing period.
- Check a Cost Report: Filter to the provider, then group by one of your own tag keys. Costs from requests that carried that key show its values. The leftover row (the portion of the bill your usage did not cover) and any usage sent without that key appear as Not tagged with the key name.
My totals changed after enabling enrichment
My totals changed after enabling enrichment
Use Cases
Each row below shows example dimensions to attach to your LiteLLM requests and what that attribution enables in Vantage. Attach custom dimensions through LiteLLM’smetadata.tags or metadata.spend_logs_metadata; LiteLLM’s organization, team, project, user, end-user, and key attributes are captured automatically as organization_id, organization_alias, team_id, team_alias, project_id, project_alias, user_id, end_user_id, and key_alias. The key names are examples; you choose your own. All rows assume the collector is running and delivering usage.
Frequently Asked Questions
Which providers are supported?
Which providers are supported?
How is this different from Custom LLM Enrichment?
How is this different from Custom LLM Enrichment?
Do I need to change my application code?
Do I need to change my application code?
Can I run multiple LiteLLM deployments or replicas with one token?
Can I run multiple LiteLLM deployments or replicas with one token?
VANTAGE_COLLECTOR_ID and its own persistent spool so Vantage can keep each batch stream distinct.What Vantage permissions do I need to enable this?
What Vantage permissions do I need to enable this?
Will this change my totals or break existing reports?
Will this change my totals or break existing reports?
Does Vantage store my prompt content?
Does Vantage store my prompt content?