Token compression
Reduces token usage at the edge with input and output compression. The source says this can trim tool-result payloads and reduce costs without changing application code.
Edgee is an AI gateway for coding agents and LLM apps. It compresses tokens, routes requests across models, and adds observability and team controls.
Edgee is an AI gateway for coding agents and LLM-powered apps. It sits between your clients and model providers, then compresses token traffic, routes requests across models, and records usage so teams can reduce cost and keep work moving when a provider fails or rate limits are reached.
The site positions Edgee around three core jobs: compress, route, and observe. For coding agents, it can be installed as a transparent proxy with no code changes; for applications, it offers an OpenAI-compatible API, SDK support, and bring-your-own-key options for direct billing control.
Reduces token usage at the edge with input and output compression. The source says this can trim tool-result payloads and reduce costs without changing application code.
Sends requests across multiple models and retries when providers fail or rate-limit. Edgee can also force routing to specific models for cost control or standardization.
Shows usage, cost, latency, errors, and compression savings in the console and logs. The observability docs also expose token usage fields in SDK responses.
Works as a transparent proxy for coding agents with a CLI install flow. The site says it supports Claude Code, Codex, Copilot, OpenCode, Cursor, and similar tools.
Supports an OpenAI-compatible API and SDK usage for apps and agents. The pricing page also notes that you can bring your own provider keys for billing control.
Provides team controls such as seat management, GitHub attribution, spending caps, and squad-level visibility on the Team plan. Enterprise adds SSO/SAML, private gateway options, and privacy controls.
Install the CLI and connect a coding assistant so requests pass through Edgee without changing application code. This is the main workflow for developers who want immediate token savings in existing agent sessions.
Set a priority-ordered model chain so Edgee can retry a request automatically when a provider returns a 429 or 5xx, or when a usage cap is reached. This keeps a Claude Code session running instead of stopping mid-task.
Use the gateway for an app or internal tool that calls models through an OpenAI-compatible API. Edgee can compress traffic, track usage, and let teams separate environments such as dev, staging, and production.
Track cost and usage by repo, PR, developer, squad, model, or environment in the console and logs. This helps teams attribute spend and review how compression or routing affects bills over time.
Use Team or Enterprise features to manage seats, spending caps, GitHub attribution, private gateways, or SSO/SAML. These controls are aimed at organizations that want more governance over how agents and models are used.
Edgee sits between your agent or app and the LLM provider. For coding agents, the source says you can install the CLI and connect it without changing application code; for apps and agents, it offers an OpenAI-compatible API and SDK support.
The source says Edgee compresses token traffic at the edge, routes requests across models, and can fall back automatically when a provider fails or a rate limit is reached. It also supports observability so you can track usage, latency, errors, and cost.
The pricing page states that coding-agent compression is free, with paid Team and Enterprise plans for added routing, observability, team controls, and enterprise features. It also shows a free tier for solo developers and a Team plan for developers who need shared controls.
The source says Edgee supports Claude Code, Codex, Copilot, OpenCode, and Cursor, and it also mentions apps and agents that call LLMs through its API. Fallback models are included on the Team plan, while some enterprise capabilities such as private gateways and SSO/SAML require Enterprise.
The observability docs say Edgee tracks requests, token usage, compression savings, latency, errors, and cost at session level and in the managed console. It also supports tags and dashboards for grouping by model, project, environment, user, or tenant.
AakarDev AI helps teams manage AI provider access, project setup, logs, and analytics in one dashboard. BYOK support included.
Benchspan is an AI agent security platform that discovers agents, blocks prompt injection and data exfiltration in real time, and supports pre-launch red teaming.
CreateOS Sandbox is an isolated compute environment for running code and agent workloads in Firecracker micro-VMs with private networking and SDK, CLI, or MCP control.
Codex Plugins bundle reusable skills, app integrations, and MCP servers into workflows you can install in the Codex app or use from Codex CLI.
Wallie is an open-source AI streamer that watches your screen, hears chat, and delivers live commentary in a configurable persona. Runs locally with your own keys.
Prompty Town turns a link into a building in a small internet city. Buy a tile, add a prompt, and publish it alongside other buildings.