the inference company

Inferencefor agents

Hosted open models, served for how agents actually work: long sessions, tool calls, and the waits in between. Change one URL and pay less for every task your agent finishes.

Drop-in for any OpenAI- or Anthropic‑compatible agent.

Animation: an agent, coding, answering a support ticket or checking contracts, sends a request across the world to our GPUs, the response streams back, and while the agent's tools run our GPUs keep the session's context ready for the next call.

Agents don't send requests. They run loops.

An agent calls the model, runs a tool, then calls again with almost the same context, dozens of times per task. Most serving stacks treat every call as new. The context can be evicted while your tools run, so the next call pays to rebuild it, and a background chat title waits in the same queue as a task that's blocked. We're building around the loop instead.

  1. Change one URL

    We speak the OpenAI and Anthropic APIs, streaming and tool calls included. Point your agent at us and nothing else changes: your own agents, any framework that lets you set a base URL, and tools like Claude Code. Switching back is the same one-line change.

  2. Context stays warm

    While your tools run, the session's context stays ready on the GPU, so the next call starts fast and doesn't pay to rebuild it. When the agent waits on a person, the context moves to memory and comes back when they reply.

  3. Blocked tasks go first

    Not every call matters equally. A call carrying a tool result is blocking a task in progress, so it runs first. Background calls, like writing a chat title, wait their turn. Your agents spend less time stalled.

  4. See cost per successful task

    Tokens tell you what you spent. We show what each finished task cost, for every agent you run, against the success signal you define: tests that pass, a ticket that's resolved, a report that's accepted.

  5. Every change is proven first

    Every change to how we serve, from caching to engine settings, is replayed on recorded agent sessions against the current setup. It ships only if tasks get cheaper or faster without failing more often.

Questions

What is The Inference Company?

An inference service for AI agents. We host an open model and serve it for the way agents work: long sessions, tool calls, and the waits in between. You get an OpenAI- and Anthropic-compatible endpoint, priced to make each finished task cheaper.

Which agents work with The Inference Company?

Any agent that talks to an OpenAI Chat Completions or Anthropic Messages endpoint, including streaming and tool calls. That covers agents you build yourself, frameworks that let you set a base URL, and ready-made agents like Claude Code, OpenCode and Cline.

Do I need to change my agent?

No. Change the base URL and API key. Your prompts, tools and harness stay as they are, and switching back is the same one-line change.

What is cost per successful task?

What you spent on an agent's tasks, divided by the number that succeeded. You decide what success means for each agent. Failed attempts still cost money, so their spend is shared across the tasks that worked. It is the number that tells you what an outcome costs, which a per-token price does not.

How is this different from other inference providers?

Most providers tune for tokens per second on single requests. We tune for what a finished task costs: we keep each session's context ready through tool calls, run the calls that block a task first, and report cost per successful task on every account.

Which model do you run?

We're starting with one open model, chosen for agent work and cost. We'll name it when early access opens.

How will pricing work?

Per token at launch, with cost per successful task shown on every account. For teams with a clear success signal, like passing tests or a resolved ticket, we plan to offer per-task pricing.

When can I use it?

We're onboarding a small group of teams first. Book a call and we'll get you set up.

Running agents on open models?

We're onboarding a small group of teams first. Book a call and walk us through your agents: what they do, what they cost today, and what counts as success. We'll show you how we'd measure cost per successful task against your current provider.