Run models locally
Download Ollama and run models on your own machine, which the pricing page describes as unlimited for local hardware usage.
Ollama is a platform for running open models locally or in the cloud, with CLI, API, and desktop workflows for automation and app integration. It is designed for users who want to keep data private while using large language models.
Ollama is a platform for running open models and automating work through local and cloud-based model access. The site positions it as a way to get up and running with large language models while keeping data private.
Its core workflow spans downloads for macOS, Windows, and Linux, plus CLI, API, and desktop app access. The pricing page also adds cloud models, usage-based tiers, and support for community integrations and public models.
Download Ollama and run models on your own machine, which the pricing page describes as unlimited for local hardware usage.
Use the cloud model offering for larger models and managed compute when local hardware is not enough.
Work through the command line, API, or desktop apps, so the same product can fit scripting, application integration, and interactive use.
Connect with the wider ecosystem through the official Python and JavaScript libraries and 20+ community-supported libraries.
Use community integrations, with the pricing page calling out more than 40,000 community integrations and unlimited public models.
Build with agent workflows using cloud models that have been tested for tool calling before release.
Use the free or local workflow to chat with models, evaluate larger models, or keep AI-assisted work on your own hardware.
Use the cloud plans when you need larger models, sustained sessions, or more concurrent models for heavier tasks.
Build scripts, apps, or internal tools against Ollama’s API and official libraries in Python or JavaScript.
Use cloud models for coding automation, document analysis, or deep research workflows that benefit from more capacity than local hardware can provide.
Adopt community libraries and integrations when you want Ollama to fit into an existing tool stack without starting from scratch.
Ollama provides a local download for running open models on your own hardware, plus cloud models and an API for teams that want managed capacity. The docs also point to official Python and JavaScript libraries and community libraries for integration.
The download page shows Windows support, and the docs say Ollama can be downloaded on macOS, Windows, or Linux. The Windows page notes that Windows 10 or later is required.
The pricing page says Free is for light usage and running models on your own hardware is always unlimited. Pro and Max add cloud-model usage, higher concurrency, and additional features such as private model uploads and sharing.
Yes. The pricing page says cloud models that are trained to support tools are tested for tool calling and real agent workflows before they go live.
The pricing page states that prompt and response data is never logged or trained on, and that hosted models are served with no logging, no training, and zero data retention policies from its hosting partners.
AakarDev AI helps teams manage AI provider access, project-level setups, logs, and analytics from one dashboard. It supports BYOK workflows and lists providers including OpenAI, Google Gemini, Anthropic, Groq, Mistral AI, and Perplexity AI.
Benchspan is an AI agent security platform that discovers agents, blocks prompt injection and data exfiltration in real time, and supports pre-launch red teaming. It is aimed at teams running agents in production and includes Python and TypeScript SDKs.
Edgee is an AI gateway for coding agents and LLM-powered apps. It compresses token traffic, routes requests across models, and provides observability and team controls to help reduce cost and keep sessions running.
CreateOS Sandbox is an isolated compute environment for running code and agent workloads inside Firecracker micro-VMs. It is designed for workflows that need machine-level isolation, private networking between sandboxes, and programmatic control through SDK, CLI, or MCP.
Codex Plugins bundle reusable skills, app integrations, and MCP servers into workflows you can install in the Codex app or use from Codex CLI. They help extend Codex with connected-service tasks, reusable instructions, and shared team workflows.
Wallie is an open-source AI streamer that watches your screen, hears chat, and generates live commentary in a configurable persona. It runs locally on your machine with your own keys and is aimed at faceless content, autonomous streams, and real-time reactions.