DeepSeek-V4-Flash-0731, the official DeepSeek-V4-Flash release on Hugging Face, for agentic and long-context workflows with speculative decoding and vLLM/SGLang guidance.
MiniMax is an AI platform for multimodal models, developer APIs, and AI-native products. MiniMax M3 focuses on coding, agentic workflows, and long-context multimodal tasks.
Claude Opus 5 is Anthropic’s Opus-tier model for coding, long-running agents, research, and professional knowledge work across Claude plans.
Grok 4.5 is SpaceXAI’s model for coding, agentic tasks, and knowledge work. Available in Grok Build, Cursor, the SpaceXAI console, and the API.
Aymo AI is an all-in-one AI platform for teams with model switching, comparison, file analysis, web search, and shared workflows.
Argmin AI turns your rules, docs, and examples into AI evaluations you can run before release—without custom code or an ML team.
BaseRT is an LLM runtime for Apple Silicon Macs that runs local models on your own device for on-device inference and a local coding-agent workflow.
Kimi K3 is Moonshot AI’s frontier model for coding, knowledge work, and reasoning, with Kimi.com, Kimi Work, Kimi Code, Kimi API, a 1-million-token context window, and native vision.
Zro is a private inference endpoint for coding agents on EU infrastructure, with OpenAI-compatible and Anthropic-compatible access and zero request retention.
derouter.ai is a web API for Claude and GPT models with fixed, discounted token pricing, plus support for Claude Code, Codex CLI, and GPT Image 2.
SuperCompress is a query-aware context compressor for LLM apps that cuts input tokens before inference while preserving key evidence. Open source, CPU, API and Python package.
Muse Spark 1.1 is a multimodal reasoning model for agentic tasks, coding, computer use, and multimodal understanding. Available in public preview via the Meta Model API, Meta AI app, and meta.ai.
Auriko is an LLM inference routing API for one integration with multiple providers, offering cache-aware cost optimization, routing control, and reliability for AI apps.
Opper AI is an EU-hosted AI gateway for 300+ models via OpenAI SDK-compatible API, with routing, observability, guardrails and compliance.
Constellation Gate AI is a gateway for AI agents that screens requests, redacts sensitive data, records tamper-evident activity, and can reduce token usage. Supports desktop tools, CLI routing, and SDK setup without code changes.
LongCat-2.0 is a LongCat AI model announcement featuring a 1.6 trillion-parameter system trained entirely on domestic chips.
TuneLLM turns recurring Claude- or GPT-style workflows into smaller fine-tuned models inside your infrastructure for benchmarked quality at lower inference cost.
Alvoff Inference is an OpenAI-compatible API for speech-to-text, text-to-speech, embeddings, and chat/code generation.
RunInfra benchmarks GPUs, tunes runtime paths, and turns open-source models into production inference stacks with managed API or export for self-hosting.
ClinePass is a paid subscription for curated open weight models in Cline, with a first-month Product Hunt offer for developers using IDE and CLI workflows.