Compatible API access
Zro exposes OpenAI-compatible chat completions, and it also supports Anthropic-compatible Messages requests at /v1/messages for tools that expect that API shape.
Zro is a private inference endpoint for coding agents on EU infrastructure, with OpenAI-compatible and Anthropic-compatible access and zero request retention.
Zro is a private inference endpoint for coding agents. It serves open-weight coding models from EU infrastructure and is positioned for teams and developers who want fast model access without request retention.
The product is built around developer workflows rather than a standalone chat app. It supports OpenAI-compatible chat completions, Anthropic-compatible Messages, and launcher-based setup for tools such as Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, Pi, and Codex App. The site also says prompts and completions are not retained by default and are not used for training, fine-tuning, evaluations, analytics, or dataset creation.
Zro’s technology page attributes its performance layer to HyperQuant compression, custom kernels, and hardware-aware deployment across AMD GPUs, NVIDIA GPUs, and Google TPUs. The pricing page shows individual plans, usage packs, and an enterprise option for shared access and procurement.
Zro exposes OpenAI-compatible chat completions, and it also supports Anthropic-compatible Messages requests at /v1/messages for tools that expect that API shape.
Supported tools can be launched through the npm-based Zro launcher, which sets up temporary session configuration and disables vendor telemetry and analytics for launcher-supported apps.
The service runs on EU infrastructure with zero request retention by default and no training on customer data, positioning it for private coding-agent inference.
Zro is tuned for long-context, multi-turn coding sessions and is described as supporting responsive, streaming inference for developer tools and production apps.
The technology page describes HyperQuant compression, custom inference kernels, and hardware-aware deployment beneath the endpoint to improve serving efficiency.
The platform currently offers MiniMax M3 and GLM-5.2, with more open coding models listed as coming soon.
Teams using Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, or Pi can launch a Zro-backed session with temporary configuration instead of wiring up a provider from scratch.
Developers with existing OpenAI-compatible or Anthropic-compatible clients can point those tools at Zro’s base URL or Messages endpoint and keep their current workflow.
Individual developers can use the Pro plan to run private coding-agent sessions with open-model inference, while usage packs cover occasional extra spend.
Teams that need shared access, custom usage plans, and support can use the Enterprise option described on the pricing page.
Cursor and Cline users can configure Zro manually when they prefer those apps’ own provider settings rather than the launcher flow.
Zro is a private inference endpoint for coding agents. It serves open-weight coding models from EU infrastructure and is designed to work with existing OpenAI-compatible clients, Anthropic-compatible clients, and supported coding tools.
Zro supports OpenAI-compatible chat completions and Anthropic-compatible Messages at /v1/messages. The integrations page also shows launcher-based setup for Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, Pi, and Codex App, plus manual configuration for Cline and Cursor.
Use the Zro launcher by installing the npm package, logging in once, and launching a supported tool. For example, the integrations page shows zro launch commands for Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, Pi, and Codex App.
No. Zro says prompt and completion bodies are not retained by default after inference is processed, and customer prompts and completions are never used for training, fine-tuning, evaluations, analytics, or dataset creation.
Zro runs on privacy-forward EU infrastructure, with current regions listed as Finland and France. The site also says it is built for responsive, streaming inference for coding-agent workloads.
ByteAsk is a terminal-first AI coding agent for C and C++ that edits repos and verifies changes with compilers, debuggers, sanitizers, and tests.
Ghost is a terminal-based AI assistant for chatting, code generation, and CLI tasks. Includes free models, supports Linux, macOS, Windows, and is open source.
CreateOS Sandbox is an isolated compute environment for running code and agent workloads in Firecracker micro-VMs with private networking and SDK, CLI, or MCP control.
hob is an independent workspace for coding agents, with local control over sessions, terminals, history, routing, and follow-up work.
Manta AI is an autonomous web app testing tool that maps app behavior, catches regressions, and generates tests from a URL, no scripts or selectors needed.
SonOf connects to your repo and PM tool, audits your codebase, and turns approved work into shipped tickets with senior engineering review.