TuneLLM is an enterprise platform that distills recurring Claude- or GPT-style workflows into smaller fine-tuned models inside your infrastructure. It is aimed at teams that want benchmarked quality on narrow LLM tasks at lower inference cost.
Staff Engineer is a set of Claude Code skills that helps engineers scaffold apps, verify toolchains, and wire production workflows with repeatable scripts. The free skills cover app setup and ops checks, while the broader pack adds deployment, observability, real-time features, and background jobs.
TryCase provides disposable Linux environments for LLMs and coding agents to run apps, verify changes, and return screenshots, recordings, logs, and artifacts. It supports headless and desktop workflows through a CLI-driven process.
Archify is a Chrome browser extension that reveals components, APIs, libraries, and other runtime signals from a live web page. It is designed for frontend engineers, QA engineers, and technical founders who need to understand how an app works without leaving the browser.
LetMeCheck.ai is an AI-assisted code analysis platform for founders, startups, SMBs, and developers. It connects to a GitHub repository, scans for security, performance, scalability, and code quality issues, and returns a diagnostic report with recommendations.
SaaS Dummies is an AI website testing tool that runs named testers through a deployed URL and reports where the experience breaks down. It helps teams catch signup, accessibility, and page-flow issues quickly, with results delivered in about 60 seconds.
Retrace is an execution replay engine for AI agents that records runs, replays them, and lets developers fork from a failing step to debug behavior before shipping changes. It also supports prove-the-fix workflows, CI regression gates, and related evaluation controls on paid plans.
Slopdar scans a website URL and returns a 0–100 Slop Score with receipts explaining the result. It is aimed at people who want to inspect whether a public site looks hand-coded or assembled with AI or template tools.
Ferguson is an AI-assisted landing page auditor that reviews a public URL and returns a score, issue list, suggested fixes, and written copy changes. It is built for founders, marketers, and agencies who want page-specific feedback on landing page conversion.
Tinkerfont is a browser extension for inspecting and testing fonts on live websites in Chrome and Firefox. It lets you swap typefaces with Bunny Fonts or custom uploads, scope changes to parts of a page, and save rules per site.
schemafit is a local-first CLI that lints JSON Schema, structured-output specs, and tool definitions against major LLM provider constraints. It helps teams fail CI before provider-specific schema issues reach production.
Qabit is an embedded AI response evaluation tool that lets teams rate outputs with a structured rubric and store the results. It is built for developers and product teams that need repeatable feedback on chatbots, agents, RAG, and other AI features.
Kodwai is a CLI-led platform for developers to solve real coding challenges with an AI agent and get scored on how they directed the session. It supports local, on-machine workflows with tools such as Claude Code, Cursor, and Codex.
CoWork is QApilot’s mobile testing product for turning existing test cases into runnable automation on real devices. It combines AI-assisted planning with human approval to help teams execute more of their backlog before release.
Cloud World Model AI simulates cloud infrastructure for AWS, GCP, Azure, OCI, and DigitalOcean without provisioning real resources. It is built for Canvas Cloud AI learners and agents that need practice, RL training, chaos testing, and cost analysis.
VibeRaven scans AI-built repositories for production-readiness gaps across auth, billing, database, deployment, monitoring, and tests, then returns a focused prompt for the next coding-agent pass.
blop is a QA agent that turns plain-English browser journeys into Playwright-backed tests stored in your repository, runs them in GitHub Actions, and returns results in pull requests. It also clusters repeated failures and can schedule tests as recurring journey probes.
BrowserBash is a natural-language browser automation CLI for running browser tasks and tests from the command line. It supports local Chrome, CDP endpoints, and cloud browser providers, with a free open-source path and no API keys required to start.
BestDefense is a continuous security validation platform that tests applications, APIs, and networks on each deploy, validates exploitable findings, and generates remediation pull requests and proof records for audit use.
AgentX provides a production-focused framework for evaluating AI agents and LLMs with layered scoring, trace analysis, drift detection, and release gating. It supports real-data test sets, multi-step runs, and self-serve or enterprise purchasing options.