Gemini 3.7 Flash is Google’s Flash-series AI model for coding, agent workflows, web development, and document-heavy knowledge work. It is available to developers, enterprises, and eligible Gemini app subscribers through multiple Google surfaces.
Statewave is an open-source memory runtime for AI agents and LLM apps. It helps teams store durable, structured context with provenance so systems can recall prior events across sessions.
Cohesor is an AI gateway and control plane for agents that routes requests, compresses tokens, brokers MCP tool calls, and enforces policy from one endpoint. It is designed for teams that want to lower model spend and improve visibility without rewriting existing Anthropic- or OpenAI-compatible clients.
Unsloth is a free, open-source desktop app for running and training AI models on local hardware. It supports macOS, Windows, Linux, and WSL, with tools for model downloads, agent connections, media generation, and private research.
Soup CLI is a command-line tool for LLM fine-tuning and post-training on your own hardware. It supports low-VRAM training via layer streaming, plus data checks, configuration generation, preference optimization, and checkpoint gating.
What's my local AI? is a browser-based tool that detects your hardware locally and recommends which AI models may run on it with Ollama or LM Studio. It is aimed at people who want to choose local models without uploading machine data.
ngrok.ai is a hosted AI gateway that routes, secures, and manages traffic to cloud or local LLMs through a single URL. It helps developers standardize model access, add scoped controls, and monitor usage without rebuilding their app’s infrastructure.
DeepSeek-V4-Flash-0731 is DeepSeek-AI’s official release of the DeepSeek-V4-Flash model on Hugging Face. It is positioned for agentic and long-context workflows, with speculative decoding support and deployment guidance for vLLM, SGLang, and local inference.
MiniMax is an AI platform for multimodal models, developer APIs, and AI-native products. Its highlighted model, MiniMax M3, focuses on coding, agentic workflows, and long-context multimodal tasks.
Claude Opus 5 is Anthropic’s Opus-tier model for coding, long-running agents, research, and other professional knowledge work. The pricing page places it within Claude’s plan lineup across web, mobile, desktop, and developer surfaces.
Grok 4.5 is SpaceXAI’s model for coding, agentic tasks, and knowledge work. It is available in Grok Build, Cursor, the SpaceXAI console, and the API, with token-based API pricing and limited-time free usage noted on the launch page.
Aymo AI is an all-in-one AI platform for teams that brings multiple models into a single collaborative workspace. It supports model switching, comparison, file analysis, web search, and shared team workflows.
Argmin AI helps teams turn their rules, docs, and examples into an AI evaluation they can run before release. It is positioned for product and engineering teams that need quality checks without building custom evaluation code or hiring an ML team.
BaseRT is an LLM runtime for Apple Silicon Macs that runs local models on your own device. It is positioned for on-device inference and a local coding-agent workflow, with public materials emphasizing speed and privacy by keeping execution on-machine.
Kimi K3 is Moonshot AI’s frontier model for coding, knowledge work, and reasoning. It is available across Kimi.com, Kimi Work, Kimi Code, and the Kimi API, with a 1-million-token context window and native vision support.
Zro is a private inference endpoint for coding agents that runs on EU infrastructure and supports OpenAI-compatible and Anthropic-compatible access. It is designed for developers and teams that want open-model inference with zero request retention.
derouter.ai is a web API for Claude and GPT models that aims to mirror the official Anthropic and OpenAI interfaces while offering fixed, discounted token pricing. It also supports CLI workflows such as Claude Code and Codex CLI, plus GPT Image 2 image generation.
SuperCompress is a query-aware context compressor for LLM applications that reduces input tokens before inference while preserving answer-critical evidence. It is open source, runs on CPU, and is available as a hosted API and Python package.
Muse Spark 1.1 is a multimodal reasoning model from Meta Superintelligence Labs for agentic tasks, coding, computer use, and multimodal understanding. It is available in public preview through the Meta Model API and in Thinking mode in the Meta AI app and on meta.ai.
Auriko is an LLM inference routing API that lets developers access multiple model providers through one integration. It focuses on cache-aware cost optimization, routing control, and reliability for AI applications.