Agentic Rate Card

Updated
August 14, 2026
What you want to accomplishTimePowerTokens in/outOpenAI / CodexAnthropic / ClaudeKimiGLM
Ask a question, draft, or rewriteExplain an idea, improve a paragraph, or make a quick first draft.Chat10 sec–2min2 / 51–10K / 0.2–2KGPT-5.6 Luna<$0.001–$0.005Claude Haiku 4.5$0.002–$0.02Kimi K2.5$0.001–$0.02GLM-4.5-Air$0.001–$0.001
Do one Slack actionAsk for a channel summary, find details in a thread, or draft a reply.Connected chat1–5min2 / 510–100K / 1–8KGPT-5.6 Luna$0.003–$0.03Claude Haiku 4.5$0.02–$0.14Kimi K2.5$0.01–$0.08GLM-4.5-Air$0.001–$0.02
Summarize a document or meetingTurn one transcript, deck, or report into themes, decisions, and next steps.Chat + file3–15min3 / 520–150K / 2–12KGPT-5.6 Terra$0.06–$0.44Claude Sonnet 5$0.06–$0.42Kimi K2.5$0.03–$0.35GLM-4.5$0.001–$0.09
Search Slack or Figma and synthesizeExplore many channels, frames, or comments and turn the evidence into themes.Connected agent10–45min4 / 50.1–1M / 5–50KGPT-5.6 Terra$0.09–$0.90Claude Sonnet 5$0.08–$0.80Kimi K2.5$0.10–$1.10GLM-4.5$0.03–$0.28
Research, analyze data, or create a shareable documentGather sources, compare evidence, develop a point of view, and polish a DOCX, HTML file, spreadsheet, slide deck, or PDF.Research agent15–90min4 / 50.5–3M / 20–150KGPT-5.6 Terra$0.39–$2.70Claude Sonnet 5$0.35–$2.40Kimi K2 Thinking$0.30–$3.50GLM-4.5$0.07–$0.88
Make small edits to a web appChange styles, adjust a component, fix a contained bug, or add one new page.Coding agent10–60min4 / 50.5–3M / 5–40KCodex · GPT-5.6 Terra$0.21–$1.38Claude Code · Sonnet 5$0.20–$1.30Kimi K2 Thinking$0.25–$3.50GLM-4.5$0.06–$0.88
Iterate heavily on the design of an appTake many screenshots, compare visual details, and go back and forth until it feels right.Visual coding agent1–4hr4 / 52–10M / 20–120KCodex · GPT-5.6 Terra$0.84–$4.44Claude Code · Sonnet 5$0.80–$4.20Kimi K2 Thinking$1–$12GLM-4.5$0.25–$3
Diagnose a difficult software problem or review codeTrace behavior across a codebase, reproduce the issue, test theories, and verify a fix.Reasoning agent30 min–3hr5 / 52–15M / 20–150KCodex · GPT-5.6 Sol$2.10–$15.75Claude Code · Opus 5$2–$15Kimi K2 Thinking$1.50–$18GLM-4.5$0.38–$4.5
Deep, decision-ready knowledge workWork across many sources, challenge assumptions, synthesize a position, and refine it.Long-running agent2–6hr5 / 53–20M / 50–300KCodex · GPT-5.6 Sol$3.75–$24Claude Code · Opus 5$3.50–$22.50Kimi K2 Thinking$3–$28GLM-4.5$0.75–$7
A heavy day of software developmentImplement several features, debug, run tests, review the whole system, and revise repeatedly.Coding agent4–10hr5 / 58–40M / 0.1–0.6MCodex · Terra → Sol$9–$48Claude Code · Sonnet → Opus$8.50–$45Kimi K2 Thinking$8–$65GLM-4.5$2–$16.25
Build a modest first version of an app from zeroPlan the structure, create the interface, connect data, test the flows, and make it shareable.Build agent8–24hr5 / 515–80M / 0.2–1.2MCodex · Terra → Sol$17–$96Claude Code · Sonnet → Opus$16–$90Kimi K2 Thinking$15–$130GLM-4.5$3.75–$32.5
Build and test an AI video pipelineResearch video models, compare renders, wire the pipeline, package model weights, deploy GPU workers, and monitor cloud tests.Model + infra stack1–3days5 / 550–250M / 0.3–10MCodex · Terra + Sol + video models$50–$500 + GPUClaude Code · Sonnet + Opus + video models$45–$450 + GPUKimi K2 Thinking$45–$480 + GPUGLM-4.5 + video models$11.25–$120 + GPU
Extreme: overnight team of 4–8 AI agentsSplit a large goal into parallel research, design, build, testing, and review workstreams.Multi-agent8–16hr5 / 540–250M / 0.5–4MCodex · Terra + Sol team$45–$308Claude Code · Sonnet + Opus$43–$288Kimi K2 Thinking$40–$330GLM-4.5 team$10–$82.5
Extreme: agent swarm across working treesRun 8–20 coding agents in parallel branches or worktrees, with continuous tests, reviews, merges, and retries.Agent swarm4–8hr5 / 560–250M / 0.3–20MCodex · Terra + Sol swarm$200–$800Claude Code · Sonnet + Opus swarm$170–$690Kimi K2 Thinking$160–$720GLM-4.5 swarm$40–$180
2/5 fast + economical4/5 strong synthesis5/5 maximum reasoning

Input includes repeated and cached reading; output includes what the model writes or reasons through. K = thousand tokens; M = million. An agent can use tools and complete a workstream.

Validated locally: Agentic Codex and Claude Code turns processed roughly 3–5M median input tokens versus 0.4–0.9M for no-tool turns. A swarm across working trees can reach hundreds of millions of processed tokens.

Cost assumption: Chat rows use standard list prices. Agentic rows assume repeated context is mostly cached—about 15% of standard input cost—while output is full price. The provider columns show a model stack, not one model. Kimi / GLM ranges are rough API planning estimates; plans and regional pricing can differ.

Agentic workflow calculator

1. Choose the primary outcome
2. Add scope modifiers
Estimated input1.3MEstimated output0.08M

Cost by model

API-equivalent estimate for this workflow. Cached input and GPU costs are not included.

Scale: $0–$100Bars grow as scope grows

One outcome establishes the base. Modifiers add the work that makes a workflow larger: project context, external sources, browser loops, visual iteration, verification, parallelism, and infrastructure.

Terminology

Input tokensEverything the model reads—your prompt, a pasted code file, a Slack thread, screenshots, or test output. Example: reopening a 200K-token repo adds that context again.
Output tokensEverything the model writes or reasons through. Example: a short Slack reply may be 300 tokens; a generated feature and test plan may be 20K.
Cached inputContext the provider has already seen and can reuse at a lower price. Example: an agent rereading the same repository map on its fifth test loop.
Cache writeThe first pass that stores reusable context. Example: the initial upload of a codebase or long meeting transcript before later reads become cheaper.
Tool callA concrete action outside the chat. Examples: search Figma, read Slack, edit a file, run tests, take a screenshot, or deploy to Vercel.
Working treeAn isolated checkout for one agent. Example: eight agents each develop in a separate worktree, then reviews and merges reconcile the changes.
Context windowHow much conversation and project material one model call can hold. Example: a large repo may need summaries or multiple passes when it does not fit at once.
Agentic Rate Card