About EvalSprint
EvalSprint is an open-source tool for testing LLM prompts the way you test code: with repeatable cases, explicit assertions and a clear record of what changed.
Why it exists
Prompts change often, and a change that improves one behaviour can quietly break another. EvalSprint keeps a suite of test cases next to your prompts, so every version can be checked against the same expectations before it ships. Each failure comes with a reason, not just a number.
Current status
EvalSprint is an early release (v0.1). It includes deterministic assertions, a mock provider, prompt comparison, a CLI and an optional Anthropic provider. Known limitations and ideas for what comes next are listed in the README.
How it's built
- TypeScript throughout, with a shared evaluation engine used by the web app, the API and the CLI.
- A React interface and a small Node.js API, with no database. Suites are stored in your browser or in JSON files.
- Automated tests, linting and type checks run in GitHub Actions on every change.
The public demo
The demo at evalsprint.vercel.app runs the mock provider only, so anyone can try the full workflow without an API key. To use real models, run EvalSprint yourself. See the quick start.
Open source
EvalSprint is released under the MIT License and maintained by bhargavthaparbusiness. Issues and pull requests are welcome on GitHub.