Parallel multi-model runs
Run the same prompt against multiple models in parallel and compare the outputs side by side in synced columns, with per-model timing and cost shown on the run.
Revalvo is a browser-based, local-first LLM eval workbench for running prompts across multiple models, scoring results, versioning prompts, and batch-testing datasets. It is designed for users who want to compare outputs and keep prompt work in the browser without creating an account.
Revalvo is a local-first LLM eval workbench for writing, running, versioning, and evaluating prompts across multiple models. The homepage positions it as a place to run the same prompt against every model at once, score results, and keep a traceable record of what was shipped.
It runs in the browser with no account and no hosted database. The site says you can connect providers with your own keys, use local runtimes such as Ollama or LM Studio, or work with custom endpoints, then move from single prompt tests to dataset-based batch evaluation and reporting.
Run the same prompt against multiple models in parallel and compare the outputs side by side in synced columns, with per-model timing and cost shown on the run.
Score responses in real time and review report outputs that surface pass rate, model ranking, per-case pass/fail, and debug details.
Save a run as an immutable snapshot, compare versions with model config diff and line-level word diff, and roll back to any earlier version.
Create datasets by hand, upload CSV files, start blank, or generate rows from a description; batch-run saved versions across the dataset.
Connect providers through API keys, OAuth, localhost runtimes, or custom endpoints, with workspace separation per provider.
Export saved versions or batch reports as Python, TypeScript, cURL, Node.js, Go, agent prompts, HTML, Markdown, JSON, or CSV.
Compare one prompt across several models at once, review the outputs side by side, and see which model scored best for the chosen criteria.
Treat prompts like code by saving immutable versions, comparing configuration and word-level changes, and rolling back if a later draft performs worse.
Build a dataset, run a saved prompt version across it, and use the report to check pass rate, latency, cost, and failures before shipping.
Connect a provider workspace with your own key or local runtime, then work in a browser without creating an account or using a hosted database.
Export prompt versions or batch results into code snippets, agent prompts, or common file formats for handoff to a coding agent or teammate.
Revalvo is a browser-based prompt engineering and LLM evaluation workbench. It lets you run prompts across multiple models, review scored outputs, version prompts, and export or roll back saved versions.
The site says you can connect providers by pasting a key, signing in with OAuth, or pointing to localhost. Supported options shown include OpenRouter, Ofox.AI, Vercel AI Gateway, Groq, OpenAI, Anthropic, Ollama, LM Studio, and custom OpenAI-compatible endpoints.
Revalvo is designed to work in the browser with no account and local storage for prompts, versions, datasets, eval configs, results, and API keys. The privacy policy says direct providers can be used without Revalvo servers seeing your content, while some providers are proxied through a worker in memory only.
The homepage shows a workflow for single runs in the playground and batch evaluation against a full dataset. Reports include pass rate, latency, cost, model ranking, per-case pass/fail, and export options in HTML, Markdown, JSON, and CSV.
No pricing page is available at /pricing; that URL returns a 404 on the site. The terms page states that Revalvo is free to use with no uptime SLA, but it does not describe paid tiers or plan details.
Edgee is an AI gateway for coding agents and LLM apps. It compresses tokens, routes requests across models, and adds observability and team controls.
ByteAsk is a terminal-first AI coding agent for C and C++ that edits repos and verifies changes with compilers, debuggers, sanitizers, and tests.
Prompty Town turns a link into a building in a small internet city. Buy a tile, add a prompt, and publish it alongside other buildings.
Manta AI is an autonomous web app testing tool that maps app behavior, catches regressions, and generates tests from a URL, no scripts or selectors needed.
Creativly is a web-based AI creative studio for fast visual concepts, mockups, and stylized images from short inputs.
AakarDev AI helps teams manage AI provider access, project setup, logs, and analytics in one dashboard. BYOK support included.