Compatible API access
Zro exposes OpenAI-compatible chat completions, and it also supports Anthropic-compatible Messages requests at /v1/messages for tools that expect that API shape.
Zro is a private inference endpoint for coding agents that runs on EU infrastructure and supports OpenAI-compatible and Anthropic-compatible access. It is designed for developers and teams that want open-model inference with zero request retention.
Zro is a private inference endpoint for coding agents. It serves open-weight coding models from EU infrastructure and is positioned for teams and developers who want fast model access without request retention.
The product is built around developer workflows rather than a standalone chat app. It supports OpenAI-compatible chat completions, Anthropic-compatible Messages, and launcher-based setup for tools such as Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, Pi, and Codex App. The site also says prompts and completions are not retained by default and are not used for training, fine-tuning, evaluations, analytics, or dataset creation.
Zro’s technology page attributes its performance layer to HyperQuant compression, custom kernels, and hardware-aware deployment across AMD GPUs, NVIDIA GPUs, and Google TPUs. The pricing page shows individual plans, usage packs, and an enterprise option for shared access and procurement.
Zro exposes OpenAI-compatible chat completions, and it also supports Anthropic-compatible Messages requests at /v1/messages for tools that expect that API shape.
Supported tools can be launched through the npm-based Zro launcher, which sets up temporary session configuration and disables vendor telemetry and analytics for launcher-supported apps.
The service runs on EU infrastructure with zero request retention by default and no training on customer data, positioning it for private coding-agent inference.
Zro is tuned for long-context, multi-turn coding sessions and is described as supporting responsive, streaming inference for developer tools and production apps.
The technology page describes HyperQuant compression, custom inference kernels, and hardware-aware deployment beneath the endpoint to improve serving efficiency.
The platform currently offers MiniMax M3 and GLM-5.2, with more open coding models listed as coming soon.
Teams using Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, or Pi can launch a Zro-backed session with temporary configuration instead of wiring up a provider from scratch.
Developers with existing OpenAI-compatible or Anthropic-compatible clients can point those tools at Zro’s base URL or Messages endpoint and keep their current workflow.
Individual developers can use the Pro plan to run private coding-agent sessions with open-model inference, while usage packs cover occasional extra spend.
Teams that need shared access, custom usage plans, and support can use the Enterprise option described on the pricing page.
Cursor and Cline users can configure Zro manually when they prefer those apps’ own provider settings rather than the launcher flow.
Zro is a private inference endpoint for coding agents. It serves open-weight coding models from EU infrastructure and is designed to work with existing OpenAI-compatible clients, Anthropic-compatible clients, and supported coding tools.
Zro supports OpenAI-compatible chat completions and Anthropic-compatible Messages at /v1/messages. The integrations page also shows launcher-based setup for Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, Pi, and Codex App, plus manual configuration for Cline and Cursor.
Use the Zro launcher by installing the npm package, logging in once, and launching a supported tool. For example, the integrations page shows zro launch commands for Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, Pi, and Codex App.
No. Zro says prompt and completion bodies are not retained by default after inference is processed, and customer prompts and completions are never used for training, fine-tuning, evaluations, analytics, or dataset creation.
Zro runs on privacy-forward EU infrastructure, with current regions listed as Finland and France. The site also says it is built for responsive, streaming inference for coding-agent workloads.
ByteAsk is a terminal-first AI coding agent for C and C++ that edits repositories and verifies changes with the real compiler, debugger, sanitizers, and tests before showing a diff. It offers a free tier plus paid plans, with editor connectors and zero-retention handling described in the source.
Ghost ist ein terminalbasierter KI-Assistent für Chats, Code-Generierung und Aufgaben direkt in der Kommandozeile. Mit kostenlosen Modellen, Linux, macOS und Windows. Open Source.
CreateOS Sandbox is an isolated compute environment for running code and agent workloads inside Firecracker micro-VMs. It is designed for workflows that need machine-level isolation, private networking between sandboxes, and programmatic control through SDK, CLI, or MCP.
hob is an independent workspace for coding agents that keeps agent sessions, terminals, history, and follow-up work organized around the tools and providers you already use. It is aimed at developers who want local control over routing, history, and workspace structure rather than a bundled model stack.
Manta AI is an autonomous web app testing tool for teams that want to map application behavior, catch regressions, and generate tests without writing scripts or maintaining selectors. It works from a URL and supports plain-English test flows, run results with screenshots, and scheduled or deployment-triggered checks.
SonOf connects to your repo and PM tool, audits the codebase and surrounding product context, and turns approved work into shipped tickets with senior engineering review. It is aimed at founders and engineering leaders who need backlog help without hiring a full team immediately.