Zro icon

Zro

Zro is a private inference endpoint for coding agents on EU infrastructure, with OpenAI-compatible and Anthropic-compatible access and zero request retention.

Zro

Private inference for coding agents

Zro is a private inference endpoint for coding agents. It serves open-weight coding models from EU infrastructure and is positioned for teams and developers who want fast model access without request retention.

The product is built around developer workflows rather than a standalone chat app. It supports OpenAI-compatible chat completions, Anthropic-compatible Messages, and launcher-based setup for tools such as Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, Pi, and Codex App. The site also says prompts and completions are not retained by default and are not used for training, fine-tuning, evaluations, analytics, or dataset creation.

Zro’s technology page attributes its performance layer to HyperQuant compression, custom kernels, and hardware-aware deployment across AMD GPUs, NVIDIA GPUs, and Google TPUs. The pricing page shows individual plans, usage packs, and an enterprise option for shared access and procurement.

Core capabilities

Compatible API access

Zro exposes OpenAI-compatible chat completions, and it also supports Anthropic-compatible Messages requests at /v1/messages for tools that expect that API shape.

Launcher-based setup

Supported tools can be launched through the npm-based Zro launcher, which sets up temporary session configuration and disables vendor telemetry and analytics for launcher-supported apps.

Privacy-first inference

The service runs on EU infrastructure with zero request retention by default and no training on customer data, positioning it for private coding-agent inference.

Long-context coding workloads

Zro is tuned for long-context, multi-turn coding sessions and is described as supporting responsive, streaming inference for developer tools and production apps.

Systems-level performance stack

The technology page describes HyperQuant compression, custom inference kernels, and hardware-aware deployment beneath the endpoint to improve serving efficiency.

Current model options

The platform currently offers MiniMax M3 and GLM-5.2, with more open coding models listed as coming soon.

Common workflows

  • Coding agents with launcher-based setup

    Teams using Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, or Pi can launch a Zro-backed session with temporary configuration instead of wiring up a provider from scratch.

  • Drop-in API replacement

    Developers with existing OpenAI-compatible or Anthropic-compatible clients can point those tools at Zro’s base URL or Messages endpoint and keep their current workflow.

  • Solo development and experimentation

    Individual developers can use the Pro plan to run private coding-agent sessions with open-model inference, while usage packs cover occasional extra spend.

  • Shared team access and procurement

    Teams that need shared access, custom usage plans, and support can use the Enterprise option described on the pricing page.

  • Manual IDE integration

    Cursor and Cline users can configure Zro manually when they prefer those apps’ own provider settings rather than the launcher flow.

Pros and Cons

Pros

  • EU-based inference with zero request retention by default
  • No training on customer prompts or completions
  • Compatible with OpenAI-style and Anthropic-style request formats
  • Launcher support for several coding tools and CLI/IDE workflows
  • Usage packs are available for one-time spend without a subscription
  • Enterprise plan is available for shared access and procurement needs

Cons

  • The site does not publish a broad model catalog; it currently lists MiniMax M3 and GLM-5.2, with additional models coming soon.
  • Some integrations are launcher-managed while others are manual, so setup varies by tool.
  • Cursor and Cline are supported through manual provider configuration rather than the launcher flow.

FAQ

What is Zro used for?

Zro is a private inference endpoint for coding agents. It serves open-weight coding models from EU infrastructure and is designed to work with existing OpenAI-compatible clients, Anthropic-compatible clients, and supported coding tools.

Which tools and client types can connect to Zro?

Zro supports OpenAI-compatible chat completions and Anthropic-compatible Messages at /v1/messages. The integrations page also shows launcher-based setup for Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, Pi, and Codex App, plus manual configuration for Cline and Cursor.

How do I get started with Zro?

Use the Zro launcher by installing the npm package, logging in once, and launching a supported tool. For example, the integrations page shows zro launch commands for Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, Pi, and Codex App.

Does Zro retain or train on requests?

No. Zro says prompt and completion bodies are not retained by default after inference is processed, and customer prompts and completions are never used for training, fine-tuning, evaluations, analytics, or dataset creation.

Where does inference run, and is it optimized for speed?

Zro runs on privacy-forward EU infrastructure, with current regions listed as Finland and France. The site also says it is built for responsive, streaming inference for coding-agent workloads.

Quick Facts

Category
Developer Tool
Primary use
Private inference for coding agents
API compatibility
OpenAI-compatible and Anthropic-compatible Messages
Infrastructure
EU regions, including Finland and France
Pricing model
Paid plans with usage packs; Pro starts at $20/month
Website
zro.moonmath.ai